AI security

Models as insider risks in the superintelligence era

As models gain access to sensitive data, tools, and business workflows, the useful security question is not whether a model is malicious. It is whether an organization has placed independent controls around a powerful, non-deterministic actor.

By Munis Badar · Founder / Visionary · ·

Separate intelligence from authority

A model can be capable without being trustworthy, and a trustworthy model can still make an unsafe decision in an unfamiliar context. The safer architecture keeps authority outside the model: identity, policy, inspection, routing, approval, containment, and audit should not depend on the model describing its own behavior accurately.

Separate intelligence from authority

User or agent
→
Identity and policy
→
Inspection and containment
→
AI model
→
Enterprise system
  • Treat agents and model-backed workflows as privileged actors, not as anonymous software clients
  • Give each caller only the model, tool, and action scope it needs
  • Keep the decision and evidence path independent from the model's response

Where AI-FW fits

AI-FW is the governance layer between agents, applications, model providers, and the enterprise systems those workflows can reach. It provides decision points around the AI path, while customers retain responsibility for the surrounding identity, endpoint, network, and workforce controls.

AI-FW control plane

Agents and copilots
→
AI-FW
→
Public, private, or local models
→
Enterprise tools and systems
  • AI agents, copilots, custom applications, Claude Code, Cursor, and MCP clients
  • Public, private, local, and OpenAI-compatible model endpoints
  • Enterprise systems reached by model-generated or agent-executed actions

Establish identity and limit privilege

AI-FW can authenticate and attribute supported traffic through API keys, JWT, mTLS, Kerberos, OIDC, and agent identity integrations. Agent Registry, Agent Trust, role-based access, policy scopes, and risk signals provide the operating context for decisions. Chain of Command adds consent-based swarms, delegated skills, and signed command authorization for supported agent workflows.

  • Identify the user, application, agent, or workflow behind a request
  • Apply model, tool, routing, and action permissions outside the model
  • Revoke credentials, consent, grants, or delegated capabilities when risk changes

Inspect prompts, responses, and agent effects

AI-FW applies prompt and response guardrails, prompt-injection and jailbreak detection, PII masking, AI DLP, semantic inspection, and response checks. For supported Cursor events, Cursor Agent Hooks extend visibility to shell commands, MCP calls, file reads, prompts, and agent effects. Capture is off by default, and only configured events and integrations are observed.

From request to evidence

Prompt or agent event
→
Inspect and classify
→
Allow, mask, deny, or route
→
Model or workstation
→
Audit and evidence
  • Use model routing for traffic that reaches the gateway
  • Use supported pre-action hooks when a workstation event must be decided before execution
  • Use IDE Artifacts and Skills & Tools to review observed effects and tool usage

Containment is a policy choice, not a model promise

AI-FW can apply allow, deny, mask, route, and risk-based decisions where the configured surface supports them. Some events are notification-only, observe-only mode does not block, and fail-open or fail-closed behavior depends on the event configuration. A hook or gateway cannot protect activity that bypasses the configured integration or runs on an unmanaged endpoint.

  • Enable observe-first capture before selecting a blocking posture
  • Use per-event failure posture when an outage must not permit an action
  • Keep model traffic enforcement in the gateway even when workstation hooks are deployed

Build an independent audit trail

AI-FW records transaction metadata, policy decisions, identity context, risk signals, routing outcomes, tool calls, audit events, and supported IDE artifacts. Conversation correlation can connect prompts, tools, and effects, but telemetry is evidence of observed activity, not proof that no other activity occurred. Configure retention, access, masking, no-log rules, and regional processing for your environment.

Compliance evidence has a boundary

The Compliance module can support an assessment of selected technical controls using the configuration and telemetry available in the assessed scope. It does not establish legal conformity, certification, operating effectiveness, or compliance across unobserved systems. Generated evidence and customer-provided attestations require review by the appropriate security, privacy, compliance, legal, auditor, or certification teams.

Privacy and workforce considerations

Prompts, file contents, paths, digests, identities, session identifiers, and tool metadata can be sensitive even when raw content is not stored. Before enabling workstation capture, define the purpose, scope, retention, access controls, masking and no-log rules, residency requirements, regulated-data handling, and workforce notice or consultation required in your jurisdiction. Digest-first storage reduces content exposure but is not automatically privacy-safe.