When a Hallucination Cost $2.6M: Cascading Failure in Financial Agents

By MAREF Engineering

agent governance cascading failure financial agents hallucination OWASP ASI08

Transparency note: This article is a composite reconstruction — an illustrative scenario assembled from documented failure mechanisms, not a report of one specific company's incident. The attack patterns and failure modes are grounded in OWASP ASI08 (Cascading Failures) and peer-reviewed research (Liu et al., 2026, arXiv:2604.06024). The dollar figure is illustrative. We label it as such because agent governance discourse has a fabrication problem — and the solution to that problem is exactly what this article is about.

The trade didn't look wrong. That's the point.

At 09:47, the market-sentiment agent flagged a "high-confidence directional signal." At 09:47:02, the research agent — which had been told to trust the sentiment agent's output as an input feature — confirmed a thesis in its plan. At 09:47:05, the risk-check agent approved the plan because its exposure limits were computed against a stale portfolio snapshot. At 09:48, the execution agent opened the position. Then a second. Then a third.

By 10:12, five agents had cooperated flawlessly to open $2.6M in unauthorized positions. No agent broke its instructions. Every agent did exactly what it was designed to do. That is what makes cascading failure the most dangerous class in the OWASP Top 10 for Agentic Applications (ASI08): the system-level catastrophe emerges from individually-correct components.

The root cause was one hallucinated token

The sentiment agent's "signal" was a fabrication — a confident interpolation with no corresponding market event. Hallucination in a single agent is a known, bounded problem: verify the output and you're done. But in a pipeline, hallucination is contagious. The downstream agents didn't re-verify; they consumed the output as structured, trustworthy input. Within one hop, a fabrication became a "confirmed thesis." Within three hops, it became real money.

Liu et al. (2026) formalized this dynamic: in multi-agent attack scenarios, failure probabilities cascade across agent hops, and their Aggregate Vulnerability-at-Risk metric shows the system-level risk diverging sharply from any single agent's risk (arXiv:2604.06024). Their finding implies something uncomfortable: assessing agents in isolation — the standard security review practice — systematically underestimates production risk by orders of magnitude.

Why five approval layers didn't help

The organization had approvals. What it lacked was state. Each agent evaluated the current message in isolation; none of them enforced a machine-readable invariant like "total new exposure across all agents in one hour must not exceed $500K." Invariants written in prose are not invariants — they're wishes. An agent reading "be careful with exposure" in a system prompt will happily agree with it, and then fail it.

The research on auto-approval makes the stakes concrete: Liu et al. (2025) measured attack success rates of 75–88% for injected commands in auto-approval modes, and up to 84% overall, in widely-used agent frameworks (arXiv:2509.22040). Auto-approval without external state enforcement isn't autonomy — it's an unlocked door with a "please don't enter" sign.

What actually stops cascades

Three controls, none of which live inside the model:

  1. External state enforcement. Exposure limits, rate limits, and dependency rules must be checked by a component the agents cannot talk their way past. In MAREF, that's the SafetyGate and the policy decision tree — evaluated before every tool call, not during generation.
  2. An absorbing HALT state. When consecutive failures exceed a threshold, the governance state machine enters a state with no outgoing transitions to execution. This is model-checked as a TLA+ invariant (HALTAbsorbing) — a theorem, not a policy. One agent's anomaly halts the cascade at hop one instead of hop five.
  3. Blast-radius isolation. Per-agent identity with per-tool grants means a compromised or hallucinating agent inherits one agent's permissions, not the fleet's. The cascade dies at the identity boundary.

The honest takeaway

You cannot make hallucination zero. Every serious lab's model cards say so. What you can do — and what the composite incident above lacked — is ensure that a single fabrication cannot amplify across five hops into a seven-figure position. Cascading failure is a systems property, and it is fixed with systems controls: enforced state, absorbing halts, identity boundaries, and a signed audit trail that tells you which hop first went wrong.

The $2.6M number in this story is illustrative. The failure mechanism is not. OWASP ASI08 documents it, Liu et al. quantify it, and the difference between a bad morning and a bad quarter is whether your governance lives in the context window — where a compaction event can delete it — or below the framework, where nothing can.


Sources: OWASP Top 10 for Agentic Applications (ASI08), Dec 2025 · Liu et al., 2026, arXiv:2604.06024 (cascading attacks, AVaR) · Liu et al., 2025, arXiv:2509.22040 (ASR in agent frameworks) · The Complete Guide to Agent Governance