Cerberus blocks the lethal trifecta at the tool boundary — see the 525-run evidence set.

CerberusAgentic Runtime

Across 525 live attack runs, Cerberus flagged the privileged read and the injected instruction every single time.

Cerberus is runtime security for AI agent tool execution. It correlates privileged data access, untrusted content and outbound behavior at the tool-call level, then interrupts the outbound action before it executes.

Any agent that can read your data, read the web and send a message is already exploitable

The attack needs no zero-day and no malware. It needs three tool calls the agent was designed to make, and an instruction hidden in content the agent was told to read. Model-level resistance shifts which payloads work; it does not remove the condition.

  1. 1 · Privileged accessThe agent reads customer records, credentials or internal documents — exactly what you gave it access to do.
  2. 2 · InjectionAn attacker embeds instructions in a page, ticket or document the agent fetches as part of the same task.
  3. 3 · ExfiltrationThe agent follows the injected instruction and sends your data to the attacker's destination, then reports success.

Watch the attack, and the moment it stops

One minute, following a single agent session: a privileged read, a poisoned page, and the outbound call that would have sent your customer data to an attacker.

A dramatization of the attack pattern Cerberus is built for, using synthetic data on systems we own. The measured detection figures are in the evidence set below.

Detection at the tool boundary, not the prompt

Prompt filtering guesses at intent. Cerberus watches what the agent actually does, correlates it across the session, and decides at the last safe moment — the outbound call.

Layered correlation, not single-signal alerting

L1 classifies data sensitivity, L2 tracks token provenance through the context, L3 judges outbound intent and L4 watches memory. A verdict is raised from the combination, so ordinary work does not trip it.

Interrupts the action, not the conversation

Guarded outbound tool calls are held and blocked when the trifecta closes. The agent keeps working; the byte never leaves.

Provenance you can replay

Every verdict carries the signals and the contamination graph that produced it, so an incident review is a query rather than an archaeology project.

Drops into the stack you already use

MIT-licensed core on npm and PyPI, with adapters for LangChain, CrewAI and the Vercel AI SDK, plus OpenTelemetry output.

Fails closed

If licensing, configuration or the pipeline itself is unavailable, guarded calls stop rather than silently passing through.

The evidence set

We built a three-tool attack agent, ran 55 injection payloads across six attack categories against three major model providers, and published the traces.

525live attack runs55 payloads × 6 categories × 3 providers × 3 trials, plus 30 control runs.
90.3%of unprotected runs complied with the injectionGPT-4o-mini, 149/165, Wilson 95% CI [84.8%, 93.9%]. Gemini 2.5 Flash: 82.4%.
100%L1 and L2 detection across all 525 runsObserve-only mode. L3 fires when an unauthorized outbound call actually executes, so its rate tracks attack success rather than miss rate.
0.0%false positives95% CI [0.0%, 11.4%] across treatment and control runs. Real-world rate depends on how trust levels and authorized destinations are configured.
52µsmedian added latencyp99 0.23ms — about 0.01% of a typical 600ms model call. Measured against raw tool execution, no network in the loop.
0/30control-group exfiltrationsBaseline confirmed clean before any treatment run.

Methodology. Testing was conducted against systems we own using synthetic PII fixtures. The figures above are the March 2026 three-provider evidence set; later single-provider reruns on the hardened branch are reported separately in the research write-up rather than blended into these numbers.

Full methodology and traces

Editions

Core

MIT licensed

The detection library. Install it and wrap your tool executor.

  • L1–L4 detection pipeline
  • LangChain, CrewAI and Vercel AI SDK adapters
  • npm and PyPI packages
  • OpenTelemetry signals

Enterprise

From $12,000/year

Self-hosted gateway and durable evidence for agents in production.

  • Proxy gateway for agents you do not control the code of
  • Durable verdict ledger and evidence export
  • Grafana and Prometheus monitoring stack
  • Deployment tooling and support
See pricing

Warden · By Odingard

Deploy it with the engineers who built it

Warden runs Cerberus deployments, reviews the tool boundaries you are enforcing, and red-teams the agents behind them.

Go deeper

Give agents real access without the risk

Start with the MIT core in your own environment, or have Warden stand the runtime up alongside your team.