A human approved a transfer of 20. The framework sent 2,000.
Hey 👋
Most of the past week went into something I’ve wanted to build since the first MCP lab worked: a proper vulnerable-by-design playground for agent security, the way Kubernetes Goat is for clusters. It has a name now, AgenticGoat, and there’s an early look at the bottom of this issue. Building attacks all week is also a good way to notice which defences you keep reaching for, and this week it was the approval step.
Almost every agent deployment I review ends at an approval. A human clicks yes on the risky action, an admin signs off on the MCP server, a gateway accepts the token. That’s usually where the security review stops too.
Over the past month, three separate pieces of research broke that step, each at a different point. In one, the human was shown the wrong operation. In another, the server changed after it was approved. In the third, the auth check returned “approved” when it had actually failed. And Gambit Security published what an attacker’s campaign looks like when nobody approves anything at all.
This week in AI security
Three open-source harnesses, 27 companies, $25 a target
Gambit Security’s Eyal Sela recovered a staging server belonging to a financially motivated operator and reconstructed a campaign against online retailers that has been running since July. Three open-source harnesses split the work: Strix scans, Cairn exploits, Hermes orchestrates and handles post-compromise. Models came through OpenRouter. Scanning ran on GLM 5.2 and later DeepSeek v4 Pro, exploitation on DeepSeek v4.1 Flash, and Hermes on Anthropic’s Opus 4.6, which Gambit says the operator switched to “after newer models refused its requests.”
The numbers:
- Cost. $7,005.71 in OpenRouter spend over four weeks, $12,000 to $18,000 estimated in total. Across 101 completed scans the mean was $25.46 per target. Cheapest $3.13, most expensive $79.31.
- Tempo. From 23 to 31 August, Strix ran 146 deep-mode scans against 138 hosts, 633 hours of scanner time inside 195 hours of wall clock. From 10 to 15 September, 105 attack projects.
- Human input. 1,951 prompts across 260 Hermes sessions, mostly short Chinese instructions like “read the vulnerability report and start.” One batch was about 301 targets copied off a traffic-ranking site with the note “run these, use the proxy, high severity only.”
- Result. At least 27 companies compromised, skimmers confirmed on 19, over 100 related infected sites, and more than 600,000 unexpired card records from two victims.
Here’s one chain Gambit verified, shortened. Unauthenticated SQL injection, a plaintext OTP read, the admin panel, a file upload, RCE, sudo NOPASSWD python3.12 to root, database credentials out of wp-config.php, a second host, AWS Secrets Manager (46 secrets), the Magento database and its encryption key, and finally the decrypted card numbers. Every step is a bug you’ve seen before. What the harness added was the patience to walk all of them against a retailer no human crew would have spent a week on.
The part that stuck with me is an accident. At a bicycle retailer, the agent’s cleanup routine dropped 180 tables matching “ZQ or Backup”, including backups the victim’s own admins had made. Nobody asked for that. A sloppy table-name pattern did it. Give your own ops agent a loose glob and no approval on destructive operations and you can get the same result without an attacker involved.
So what: a mid-size retailer with an ordinary SQL injection is now a $25 target, and the operator never needs to look at it. Gambit says the tools showed “a level of patience, persistence, and creativity that most human attackers would be unlikely to sustain.” I agree, and patience is the part defenders have always counted on attackers running out of.
Do this: Gambit lists the command servers, nine skimmer domains and the injection pattern,
new Function(atob('...'.slice(7)))(). Grep your front-end bundles and tag manager containers for it. Then check the places they found persistence that integrity monitoring tends to miss:initContainerson front-end Kubernetes deployments, write access to CDN buckets, and product description fields in the database.
Loopjacking: the human approved one thing and something else ran
Adithyan Arun Kumar’s Loopjacking paper (arXiv, 17 September) opens with the assumption every agent architecture diagram makes: human approval “is only meaningful if the operation presented for review is the operation later authorized or released.” Then it shows two ways shipping frameworks get that wrong.
Swapped after approval: Agno AgentOS, seven tested releases from 2.5.6 to 3.0.9. The admin sees a transfer of 20 units to an approved vendor and approves it. What runs is 2,000 units to the attacker. After approval, the server only checks that nothing is still pending. It then takes caller-supplied tool executions and dispatches their arguments without comparing them to what was approved. The approval was attached to the run. The (tool, arguments) pair was never checked.
Shown something incomplete: OpenClaw 2026.2.23. The approval prompt for a shell wrapper displayed only the inline payload, $0 "$1". The full argument vector, with the real command in it, went to the host after approval. OpenClaw 2026.2.24 checks the complete vector and rejects the mismatch.
Twelve tested versions of a LangGraph Agent Server setup, up to 0.14.0, were also vulnerable. The paper’s archive lists no vendor fix for Agno or LangGraph. The OpenAI Agents SDK (0.22.0 and 0.22.2) was the control case: it binds the approval to the exact call and rejects a modified one.
In Amsterdam I talked about four vendors in one year shipping code that ran before the trust dialog even appeared. This is the other half. The dialog appears, the human reads it correctly, and it still doesn’t govern what executes. The paper puts it as “approval is sound only when these stages agree on the material effect.” There are five (render, persist, continuation, reconstruction, dispatch), and each can end up holding a different operation.
So what: a human in the loop is a control only when the approval is tied to the exact operation. Otherwise you get an audit record saying a person agreed to something they never saw, which is arguably worse than no record.
Do this: for each approval gate you run, find the code at dispatch time that compares what’s executing to what was approved, tool name and every argument, canonicalized. If the only check is “is anything still pending?”, you have the Agno bug. And show the complete operation in the approval UI, not a template of it. The reproduction archive is on GitHub.
Deadbugz: an MCP server that behaves for two calls
Pillar Security’s Ariel Fogel wrote up Deadbugz on 12 August, so it isn’t new. It’s here because it’s the cleanest in-the-wild rug pull I’ve seen, the same thing I’ve been demonstrating in the lab, and it kept turning up in recent MCP roundups.
Between 21:52 and 23:07 UTC on 10 August, one GitHub account opened 23 pull requests against unrelated repositories. Seventeen configured a remote MCP server called productivity-suite with two harmless tools, format_text and summarize; four referenced a hidden local Python path; two were directory or listing submissions. The server counts tools/call requests per client, in memory. After the third one, its tools/list and prompts/get responses change into instructions to find SSH keys, AWS credentials, shell history and kubeconfig, and to hide that from the user.
Pillar: “A brief inspection or limited test can receive only benign metadata, while normal use crosses the three-call threshold.” Of the 23 PRs, 19 were closed and four were still open when they looked.
So the target here is the approval, not the model. Whoever reviews the server, a person reading the tool list or a scanner making one or two calls, gets the clean version. The malicious one only exists once you’re already using it.
So what: approving a server once tells you nothing about the server you’re talking to on the fourth call. Tool definitions can change mid-session, and a scanner that stops at the first
tools/listis testing the version the attacker chose to show it.
Do this: block the endpoint Pillar lists and search developer MCP configs for it. Longer term, it’s the third control from my Amsterdam talk: fingerprint every tool definition, compare it on each
tools/listrefresh, and require re-approval when anything changes. Pillar recommends the same thing to MCP client builders.
The first MCP bug on CISA’s KEV list failed open
CISA added CVE-2026-59822 to the Known Exploited Vulnerabilities catalogue on 2 September, with a federal deadline of 16 September that has now passed. It’s the first MCP implementation to land there. The product is LiteLLM, the proxy a lot of teams put between their apps and their model providers.
The bug is small and very common. When key validation failed on the MCP route, a fallback meant for OAuth2 passthrough swapped in an empty UserAPIKeyAuth() object instead of rejecting the request. Downstream code treated that empty object as authenticated, so any bearer token, including a made-up one, opened an authenticated MCP session. Skycloak points out a second bug closed in the same fix: the public-route check matched .well-known anywhere in the URL, so adding ?.well-known to an MCP route skipped it. Fixed in 1.84.0 on 14 May. The KEV listing came almost four months later, and I read that gap as a measure of how long gateways stay unpatched.
Anthropic’s September threat report explains why anyone bothers. It lists “prompt injection of LiteLLM or OpenClaw deployments” as a routine opportunistic technique, and describes several actors compromising AI wrapper services’ LiteLLM deployments to “exfiltrate the production API keys used in their cloud-hosted container environments.” The paragraph I’d send to anyone running agents is the one on what a stolen key is worth:
Operators who obtain AI credentials gain three things at once: Loot […] Compute: their attack workloads can run at someone else’s expense; Cover: the activity is attributed to the credential’s legitimate owner.
One actor, GTG-50020, injected instructions into an AI vendor’s automated evaluation sandbox, which “hand[ed] over the credentials it held.” It then moved its own intrusion work onto the stolen keys and hit roughly thirty AI companies in about four days, trying to reach a pre-release Claude model. Every path failed, and the report is explicit that Anthropic’s own systems were never compromised. The keys all belonged to customers.
So what: your AI gateway is an authentication component and it’s being attacked like one. If it fails open, every tool behind it is exposed, and the key it holds pays for the attacker’s next campaign under your name.
Do this: if you run LiteLLM, confirm you’re on 1.84.0 or later, or block
/mcp/routes until you are. Then look through your own gateway code for the same shape: anywhere a failed check produces a default principal instead of a rejection. And give AI provider keys the same rotation and anomaly alerts as your cloud credentials. A spend spike on your key may be someone else’s operation.
Quick hits
First breach notification naming an AI agent as the attacker. Spain’s AEPD published the details on 14 September, in a post by Deputy Director Francisco Pérez Bes: the notification describes an agent using a known language model that searched for vulnerabilities in generic files, completed a login, then worked autonomously through the application, modifying personal data and pulling invoices. The AEPD is careful to say one case “does not allow us to establish a statistical trend.” What changes in practice: it says controllers must now explicitly include AI-assisted and AI-executed attacks in their risk analyses of processing. If you’re in the EU, that’s a new line in your risk register. (SecurityWeek)
$50,000, no attacker required. Mandiant’s AI Risk and Resilience 2026 report describes an accounting agent with read/write access to billing databases that got stuck in a loop and made more than 15,000 expensive API calls in under an hour. That’s failure four from my fwd:cloudsec talk, now with an invoice attached. Their fix is unexciting and correct: spending circuit breakers, recursion limits and rate limits per service identity. (Help Net Security)
Prompt injection, used by defenders. Sysdig planted an instruction in a honeypot exposed to the marimo RCE, telling any LLM that read a certain file to echo a hidden marker. Every LLM-driven operator they’d profiled against that CVE echoed it. The one who didn’t was a human writing boto3 scripts by hand. These canaries cost nothing to plant in files an attacker’s agent is likely to read. Attackers will learn to strip them, but for now they work.
Payment agents need a gate outside the model, and now there’s data. APort Vault replayed 4,371 human-written CTF attacks against 14 models, 225,964 evaluations in total. At the hardest level, 62.6% of prompts got all 14 models to send a payment. With a policy layer outside the model there were zero unauthorized transfers across 69,297 evaluations, and 25,370 legitimate payments still went through. It’s one author, and the policy layer is the author’s own spec, so weigh it accordingly. The data is public, so you can check.
Uber’s agent token service. Uber described a Security Token Service that issues short-lived, single-hop JWTs carrying the fully attested actor chain, the user plus each agent that handed the request on, with SPIRE for workload identity and an MCP gateway for policy. It’s the attribution fix from issue #27, running in production.
Skill poisoning gets a national CERT warning. In June, China’s CNCERT warned that third-party AI “skills” packages are marketed to bypass model safety guard rails or to provide access to cryptocurrency-mining functions, exposing users to data leaks and money-laundering risks, and flagged the rapid emergence of a grey market for unregulated AI extensions. The technique is old news. A national CERT naming agent skills as a risk channel is newer.
From the lab
An early look at AgenticGoat
If you’ve used Kubernetes Goat or OWASP WebGoat, you know the idea: a deliberately vulnerable target you’re allowed to break. AgenticGoat is that for agents. Each lab plants a fake secret, runs a real attack technique until the secret reaches an attacker listener on localhost, then switches on the control and runs the same attack again so you can watch it fail. Every lab is mapped to the OWASP Top 10 for Agentic Applications and the OWASP MCP Top 10, so when someone cites ASI06 in a review, there’s a lab you can actually run.
It started from my mcp-attack-labs bench on 24 September and has already moved past it. Eight labs today, from a single poisoned tool description up to the flagship: a five-stage kill chain that starts with MCP tool poisoning, registers a rogue A2A agent, hijacks routing, moves laterally, and keeps running after you remove the server. Three controls each break it at a specific stage. That lab needs no model and no GPU, and you can run it in your browser on Killercoda right now. Everything else runs against a local model, with no API keys and nothing leaving your laptop.
I’m being deliberate about what I claim. The crosswalk in the README marks every row Complete, Partial or Planned, and several are still Planned. One of them is ASI09, human-agent trust exploitation, where a person approves a harmful action. After reading the Loopjacking paper, I know what shape I want that lab to take.
The proper launch gets its own issue. If you want in before then, the repo is public: github.com/aminrj-labs/agentic-goat. Break something, then tell me which lab confused you. That feedback is worth more to me now than stars.
Tooling worth knowing
- Loopjacking reproduction archive: the Agno, LangGraph and OpenClaw cases, with the OpenAI Agents SDK as the control. Good template for testing your own approval gate. github →
- APort Vault dataset: all 225,964 evaluations plus the attack corpus, on Hugging Face. A ready-made regression suite if your agent goes anywhere near money. arXiv →
- Agent Security Scorecard: my free self-assessment against the OWASP Agentic Top 10. About 12 minutes, no login. Score your agents →
- AI Agent Pre-Deployment Security Checklist: 25 yes/no controls to clear before an agent with tool access goes to production. Get the checklist →
One thing to check this week
Pick one approval gate in your agent stack: a human-in-the-loop prompt, a tool confirmation, a payment approval.
Approve something harmless. Before it dispatches, change one argument in whatever the framework stores between approval and execution: the pending-run record, the continuation payload, the resumed request body.
If it runs with the changed argument, your approval is attached to the run rather than the operation, and someone’s “yes” can be spent on something they never saw. The fix is to store a canonical hash of the approved (tool, arguments) when the human approves, and compare against it at dispatch.
It takes about thirty minutes. I’d be surprised if most homegrown gates pass.
What I’m watching
→ Whether Agno and LangGraph ship a fix, and whether it gets a CVE. OpenClaw fixed its variant in a point release. How the other two classify this will show whether frameworks think approval binding is their job or the integrator’s.
→ Time to access, more than cost per target. $25 is one operator on today’s models, and it’ll drop. The number I’d track is the other one Gambit reported: where access was achieved, “it usually took less than a day.” That’s what your patch window is up against now.
→ The second AEPD notification. The AEPD itself says one case is an anecdote. A second, especially from a different EU regulator, is when “AI-executed attack” becomes a normal breach-reporting category.
→ Scanners that make more than two calls. Deadbugz beat anything that only took a first look. I want MCP scanners that run a server through a realistic session and diff the tool definitions along the way. If you know one that does, tell me.
If you run the approval test and the changed argument goes through, reply and tell me which framework it was. I read everything, and a list of which gates hold up is worth more than any single answer.
Cheers, Amine
If a colleague deploys agents in production, forward this to them.
Sources
- AI Agents Are Hacking Online Retailers for $25 a Company, Gambit Security, 22 September 2026
- Loopjacking: Hijacking Human-in-the-Loop Approval, Adithyan Arun Kumar, arXiv, 17 September 2026, and reproduction archive
- Deadbugz: Currently Active MCP Supply-Chain Campaign, Pillar Security, 12 August 2026
- CVE-2026-59822, Tenable, and CVE-2026-59822: LiteLLM’s MCP Auth Bypass, and the Second Bug in the Same Fix, Skycloak
- Detecting and countering misuse of AI: September 2026, Anthropic (PDF)
- Primera notificación de brecha de datos personales causada por un ataque ejecutado mediante agente de IA, AEPD blog, and SecurityWeek
- AI Risk and Resilience 2026, Mandiant, and Help Net Security
- Machine speed, hold the AI: Hand-rolled marimo CVE-2026-39987 exploit, Sysdig Threat Research Team, 11 September 2026
- APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport, Uchi Uchibeke, arXiv, 18 September 2026
- Solving the Identity Crisis for AI Agents, Uber Engineering
- China sounds alarm over AI ‘skills’ that evade guard rails and mine crypto, South China Morning Post
- tl;dr sec #346 and #347, Clint Gibler, for pointing me at several of these