AI firewall & governance gateway

Every AI Agent, Identified. Authorized. Inspected. Governed.

AI-FW sits between your AI agents and the models they call, blocking jailbreaks, masking PII, enforcing policy, and auditing everything. One gateway for all of it.

Fail-closed by design Zero content retention OpenAI + Anthropic native
AI-FW gateway, live
POST /v1/chat/completions
{ "model": "gpt-4o", "messages": [ … ], "stream": false }
Prompt inspected
Jailbreaks blocked · PII masked · policy rules applied
Policy matched
Routing rule #12 · risk profile within bounds
Routed to backend
gpt-4o → configured provider
Response scanned
Toxicity + data-exfiltration checks
Audited
Metadata-only log · zero content retention
200 OK· streamed to client

Works with the models and tools your teams already use

OpenAIDeepSeekAnthropic ClaudeCursorM365 CopilotMCP toolsOllamavLLMLM StudioMistralGroqOpenRouterOpenAIDeepSeekAnthropic ClaudeCursorM365 CopilotMCP toolsOllamavLLMLM StudioMistralGroqOpenRouter
The control plane

Your whole AI estate, on one dashboard

Live traffic, guardrail decisions, identities, and audit - all in a single operational view built from metadata, never raw content.

AI-FW dashboard showing live traffic, guardrail decisions, identities, and audit
Capabilities

Security for AI, without slowing it down

Six capabilities that turn a model API into a governed, auditable, enterprise-grade AI pipeline.

Prompt & response guardrails

Jailbreak and prompt-injection detection, PII masking, toxicity filtering, and data-exfiltration blocking on every request, inbound and outbound.

Semantic intent analysis

Judge the meaning of traffic, not just its text. Score requests against natural-language policies with embeddings, a chat classifier, or a fully offline engine.

Intelligent model routing

Route each request to the right model and backend, with per-model keys, failover, and weighted distribution across providers.

Identity & access

Authenticate agents with API keys, JWT, mTLS, or Kerberos. Enforce roles with RBAC, and connect SSO with OIDC and SCIM provisioning.

Risk & audit

Rolling risk scores per user, agent, and IP, with automatic blocking. Every transaction lands in a metadata-only audit log.

Reliability & caching

Per-model retries, failover, and distribute modes. An exact and semantic completion cache keeps latency down and bills in check.

How it works

One pipeline. Four steps.

No agent-side changes. Point your clients at the gateway and the pipeline takes over.

01

Connect

Point any OpenAI-compatible or Anthropic client at the AI-FW gateway, OpenAI SDKs, Claude Code, Cursor, MCP tools, or plain HTTP.

02

Inspect

Every prompt and response is scanned in real time: jailbreaks, PII, toxicity, exfiltration, and your own custom rules.

03

Govern

Semantic intent scoring, per-model routing policies, and risk profiles decide what flows, and what gets blocked before it leaves.

04

Audit

Metadata-only transaction logs, live dashboards, and audit events give your security team full visibility without storing content.

0+
Guardrail combinations
8
Log types
100 audit actions
0+
Models Supported
0%
Fail-closed, by design
Cost & security

Pay less per prompt. Expose fewer keys.

Three reasons teams move from direct-to-LLM integrations to a gateway.

  • Caching cuts spend. Exact and semantic cache hits re-serve repeated prompts without another upstream call: a first request returns 200 in 0.64 s, an identical second request in 0.0075 s, about 80x faster, with byte-identical bodies. Fewer upstream calls means lower token bills.
  • API keys stay server-side. Provider keys are stored in the model registry and injected on forward. Endpoint agents that talk directly to the LLM never need to hold them, shrinking the blast radius if an agent is compromised.
  • Compression trims tokens. Semantic-gated prompt compression strips filler before forwarding, meaning never changes, and a rejected gate sends the original. Explore prompt compression.
Explore caching
0.64 s
First request (uncached)
0.0075 s
Identical second request (exact cache hit)
~80x
Faster response on cache hit
0
Provider keys on endpoint agents
Governance

Know what your AI is doing, and stop what it shouldn't

Shadow AI, risky agents, leaked credentials, policy drift, AI-FW surfaces them all on live dashboards and risk profiles, and can block automatically before anything sensitive leaves the building.

  • Metadata-only transaction log, content is never persisted
  • Risk auto-block: N violations in a rolling window → cooldown
  • Shadow AI discovery for unregistered models
  • Audit events for every configuration change
  • M365 Copilot admin bridge with agent sync
live

Transaction feed

09:41:02user@corpgpt-4o200 · allowed
09:41:05agent-7gpt-4o400 · jailbreak blocked
09:41:11agent-7claude-3.5200 · PII masked
09:41:17user@corpunknown-model400 · shadow AI
09:41:23ops-botgpt-4o200 · semantic flag

Put a firewall in front of your AI

Get the gateway running in minutes. Inspect, route, and audit your first prompt today.