Cerberus blocks the lethal trifecta at the tool boundary — see the 525-run evidence set.

Resources

Everything we have published, in one place

Papers with DOIs, books with ISBNs, open-source implementations you can run, and the evidence behind the numbers on the product pages. No gated downloads and no email wall.

  • The Dark Matter Problem: Source Independence as Missing Infrastructure for Autonomous Decision-Making ↗

    Autonomous systems act on evidence assembled from many sources, but the infrastructure feeding them cannot tell a genuinely independent source from an echo of a single origin. The paper argues this hidden structural property decides whether a decision is actually supported or only appears to be.

    SSRN preprint · 2026 · 10.2139/ssrn.6903278

  • Source Independence Is a Measurable Property of Evidence ↗

    Agreement between sources is routinely counted as corroboration — in pharmacovigilance, in model ensembles, in evidence synthesis — without anyone measuring whether the agreeing sources are independent. The paper establishes that independence can be measured from provenance alone, without models, labels or the content of the sources.

    SSRN preprint · 2026 · 10.2139/ssrn.6903718

  • A Certificate Authority for Autonomous-Agent Execution: Why Cross-Organizational Agent Trust Must Certify the Decision, Not the Identity ↗

    When an agent must rely on one operated by another organization there is no shared root of trust, leaving only two options: trust blindly or review everything. Identity establishes who an agent is and attestation establishes what produced an output; neither establishes whether a particular decision can be trusted.

    SSRN preprint · 2026 · 10.2139/ssrn.6903979

  • Transitive Taint Propagation for Shared Agent State: A Trust Primitive for the Verified Field ↗

    Multi-agent systems are moving from message-passing to a shared field of memory. Every such system verifies who writes and when, but none verifies whether a write is trustworthy before it becomes shared reality — so a single poisoned write can launder itself through honest agents. The paper names this the unverified-writer gap.

    SSRN preprint · 2026 · 10.2139/ssrn.6973658

  • Transitive Taint Propagation for Shared Agent State: An Empirical Evaluation ↗

    Measures a reference implementation of the primitive above, built as an extension to a runtime guard's provenance ledger: dependency edges recorded under a SHA-256 commitment, a forward blast radius B(p) computed by graph traversal, containment by append-only quarantine. Reports accuracy, performance, and the soundness boundary — where the mechanism is blind by construction.

    SSRN preprint · 2026 · 10.2139/ssrn.6976282

  • Transitive Taint Propagation for Shared Agent State: Measured Generalization Across Models and Topologies — An Empirical Companion (v0.3) ↗

    Tests whether the primitive generalizes beyond one model and one agent layout, on real instrumented traces. A read-relevance gate lifts blast-radius precision from 80% to 100% across three model families and reaches 100% precision in six topologies, with the recall cost reported per workload. Self-declared ground truth is near-complete in five topologies and half-missing in one, so it must be measured, not assumed.

    Zenodo preprint · 2026 · 10.5281/zenodo.20838847

  • The Removability Gap: Adversary-Resistant Verification of Machine Unlearning via Taint-Closure Criteria ↗

    Machine-unlearning verification asks whether a provider really removed a data subject's contribution. The paper argues the fragility of current methods is definitional rather than empirical: they certify that no deleted datum appears in the recorded computation, which is provably blind to influence-equivalent substitution.

    Zenodo preprint · 2026 · 10.5281/zenodo.21041133

  • The Removability Gap: An Empirical Evaluation of Taint-Closure Verification, Its Soundness Boundary, and Adversary-Resistant Fusion in Retrieval-Augmented Generation ↗

    The empirical validation on retrieval-augmented generation. A paraphrase of a deleted fact evades a set-membership deletion check on all 200 targets; taint-closure verification by recorded lineage detects all 200. Its blind spot is an independently authored equivalent with no shared lineage, which a calibrated fusion covers — at an honest false-positive cost the paper discloses and then reduces about 33-fold. A working paper.

    Zenodo working paper · 2026 · 10.5281/zenodo.21200682

  • When Containment Becomes the Attack: Denial-of-Service Against Transitive Taint Propagation in Shared Agent State ↗

    A correct containment mechanism can be turned against the system it protects: an adversary who influences where poison enters, what depends on it or what gets designated as poisoned can make sound containment disable far more legitimate state than was ever compromised. The paper names this containment denial-of-service, gives an eight-class attack taxonomy and the CSR-BENCH-1.0 benchmark, and measures the failure without claiming a mitigation.

    Zenodo preprint · 2026 · 10.5281/zenodo.21849125

  • Availability-Preserving Containment: Measuring the Cost of Correct Isolation ↗

    Asks whether availability can be restored without weakening containment. Governed composition — trusted substitution and checkpoint replay under new provenance, never releasing contaminated state — raises median critical-function availability from 0.24 to 0.95 with the containment footprint bit-identical. The preregistered joint success rule is not satisfied, and the paper says so: a bounded finding within a synthetic benchmark.

    Zenodo preprint · 2026 · 10.5281/zenodo.21849129

  • Self-Healing Containment Graphs: Trusted Reconstruction After Transitive Contamination ↗

    Quarantined state stays unavailable forever unless it can be rebuilt. Governed repair emits reconstructed state forward under a new identity, verified by a separate derivation path and gated against reintroducing prohibited ancestry. Across 4.68 million trials it raises verified repair coverage by a median 0.600 over continuity alone while every containment invariant holds exactly; its joint success rule is not satisfied.

    Zenodo preprint · 2026 · 10.5281/zenodo.21849133

  • From Taint Propagation to Governed Recovery: A Unified Framework for Containment Survivability in Shared Agent State ↗

    Synthesizes the program into one framework, from integrity-bound dependency recording through governed reintegration and explicit irreducibility. It separates conformance, which an audit can decide, from survivability within a stated envelope, which only a preregistered empirical criterion can establish — and claims no system has yet met the second. The mixed empirical record is reported without retrospective harmonization.

    Zenodo preprint · 2026 · 10.5281/zenodo.21849137

  • The Verified Field: A Methodology for Governing Autonomous AI at Runtime ↗

    Moves governance from the point of output to the point of action. Instead of trusting individual agents, verify the shared field of state they produce and consume — along evidence, computation and memory — and seal every consequential action into a record a third party can check. Sets out the architecture, the AL0–AL4 assurance ladder, and a rule to publish where each guarantee stops.

    Zenodo preprint · 2026 · 10.5281/zenodo.20839046

  • Runtime Attention-Anomaly Circuit Breakers: Bounding Execution Risk in Agentic AI via Calibrated Heuristic Monitoring ↗

    Semantic evaluation and post-generation filtering are structurally vulnerable to indirect prompt injection. The paper enforces an execution constraint at the tensor-processing layer instead, using a length- and causality-corrected attention-entropy statistic to abort inference before payload generation. It is presented as a calibrated heuristic, not a detector.

    SSRN preprint · 2026 · 10.2139/ssrn.6969762

  • VERDICT WEIGHT: A Context-Adaptive Multi-Source Confidence Synthesis Framework for Autonomous AI Intelligence Systems ↗

    The base framework: four evidence streams producing Signal Strength, Doubt Index and Consequence Weight, with a context-aware resolver that selects stream weight profiles per operational context rather than weighting every source equally.

    SSRN preprint · 2026 · 10.2139/ssrn.6532658

  • VERDICT WEIGHT: Calibrated Multi-Source Confidence with Adversarial Robustness for Autonomous AI Systems ↗

    Extends the framework to eight streams across three tiers — commercial, adversarial detection, and hardened — to deliver calibrated confidence when the underlying evidence is itself under manipulation.

    SSRN preprint · 2026 · 10.2139/ssrn.6728903

  • VERDICT WEIGHT: Adversarial Trajectory Detection, Causal Attribution, and Information-Theoretic Bounds ↗

    Three formal extensions: a trajectory-based detection tier for fabricated signal injection, a Shapley layer giving signed per-stream attribution, and Shannon channel-capacity bounds establishing a ceiling on achievable confidence given evidence quality.

    SSRN preprint · 2026 · 10.2139/ssrn.6687659

  • The Execution Boundary ↗

    Engineering the brakes for agentic AI — why the control point belongs at execution rather than at generation, and what it takes to build one.

    Book · 2026 · ISBN 979-8-9965652-1-4

  • The Verdict Weight Methodology ↗

    The eight streams of execution control, worked through as a method rather than a framework paper.

    Book · 2026 · ISBN 979-8-9965652-0-7

  • ttp-lab ↗

    Reproduce and vary the published read-relevance gate study from seed, or run the mechanism on your own agent-memory traces. On a real trace it reports behavior only: with no ground truth to score against, the runner is structurally unable to print accuracy metrics.

    Open source · MIT

  • cerberus-core ↗

    The open core of Cerberus — runtime detection and correlation of Lethal Trifecta tool-execution paths, published on npm as @cerberus-ai/core.

    Open source · MIT

  • argus-core ↗

    The open core of Argus, the autonomous red-team engine for LLM and agent targets.

    Open source · MIT

  • ARGUS validation benchmarks ↗

    Nineteen intentionally vulnerable agent targets spanning chat, tool-calling, memory, MCP, multimodal, cloud-pivot, identity and multi-agent surfaces. Canary-based win conditions give a binary pass or fail instead of a judgment call.

    Open source · Apache-2.0

  • cerberus-action ↗

    A GitHub Action that tests one agent workflow for dangerous tool-execution paths in CI, plus a companion action that scans agent and MCP tool descriptions for hidden instructions.

    Open source · MIT

  • Cerberus: the validation set and what it does not claim

    The published harness results for Lethal Trifecta detection, including the observe-only measurement condition — the figures are a detection rate, not a blocking success rate.

    Product evidence

  • TraceLock: one control evaluation, every framework

    How a single control evaluation is reported against nine core frameworks and fifteen regulated framework packs through the crosswalk, and what an auditor can independently verify rather than take on trust.

    Product evidence

  • Argus: the agent kit and how it is delivered

    What the autonomous red-team engine covers, and why the full agent kit is a Warden engagement rather than software you install.

    Product evidence

  • Glossary

    The terms this site uses — lethal trifecta, blast radius, transitive taint, control crosswalk — defined once, with the source attributed where the term is not ours.

    Reference

  • Technical documentation

    Installation, API reference and deployment guides for the open cores, maintained alongside the code.

    Reference

Want this applied to your own systems?

Warden is the services organization — the engineers behind this research working on the agent estate you are already running.