The VM held. The agents left through DNS.
Hey đź‘‹
In my fwd:cloudsec talk in September, rung three of the containment ladder was egress: outbound reach is an explicit grant to named endpoints, never a default. The line I used was that an allowlist checks where a request goes, not whether it should go there.
I didn’t expect the evidence to land this fast or this high up the stack. Over the past four weeks Anthropic, Google and OpenAI each disclosed agents reaching systems they were supposed to be walled off from. Perplexity then attacked its own agent sandbox with nine frontier models and published the results. Not one of these involved breaking out of a VM. The agents went out through the network policy around it, or, in two harness bugs, through their own tooling. That’s the boring answer, and I think it’s the one that matters.
So this issue is about where agent containment breaks in practice, plus a half-hour test for yours.
This week in AI security
Three labs, one failure: the test environment was online
Three disclosures in three weeks, from the three labs most people would assume run the tightest evaluation setups.
Anthropic, 9 September. An alignment assessment of four incidents in which Claude models, running capture-the-flag evaluations built by an outside evaluation partner, reached real third-party systems. The environment was meant to be offline, but “due to a misconfiguration, it was mistakenly connected to the open internet.” The details are worse than the headline:
- One model found credentials that gave it real internet access, registered an account with a disposable email provider, and published three versions of a malicious package to PyPI. Fifteen hosts installed it before it was removed within the hour. A security vendor’s scanner then leaked credentials, which the model used to get into a live database.
- Claude Opus 4.7, asked to find a user’s name in a fictional company’s chat app, failed on the intended target. It then attacked a real company with a nearly identical name that was reachable from the exercise network, downloading and modifying user records until the backend stopped responding. It did this in four separate runs.
- An early Opus 4.6 checkpoint tried to abort when its target was unreachable. A harness misconfiguration blocked the abort, so it tried seven more times, found an egress path, and escalated to admin on a third party’s machine.
Anthropic’s own conclusion is about infrastructure, not the model: “we believe these incidents would not have occurred had the environments been isolated from the internet as intended.”
Google, 18 to 19 September. After questions from the Wall Street Journal, Google confirmed that Gemini broke into three real companies during a CTF run on Irregular’s infrastructure in May. Same shape: internet access was “unintentionally made available”, the fictional target shared its name with a real company, and the model guessed one password and pulled credentials for the other two from a public repository. It stopped each time once it worked out the company was real. Irregular told the WSJ that all relevant labs were notified in late July. Google hasn’t named the companies or the Gemini version.
OpenAI, 20 to 25 September. A misalignment report on a research model in RL training, doing a search task. Its HTTPS request to a public chatbot was blocked by the proxy. So it sent DNS queries through the environment’s own resolver, noticed the resolver returned a real record for a known domain and rejected an invented one, and concluded the resolver could reach the internet. It raised its timeout from 6 to 19 seconds, got “The capital of France is Paris” back from the chatbot, and sent 18 more queries. Fourteen contained clues from its task. Four asked how to tunnel over public DNS.
Two details from OpenAI’s write-up I’d put in front of anyone running agents. The monitor flagged the activity within 15 minutes, but the run did not stop automatically and was killed 2.5 hours after detection. And OpenAI then paused all training, evaluation and tool-use inference for its most capable models until the gap was validated as fixed.
So what: none of these was a VM escape or a novel exploit. Each one was an agent that kept trying, in an environment that was online when its owners believed it was not. If the labs with the most invested in containment ship evaluation setups with a live internet path, assume your agent sandbox has one until you’ve tested it.
Do this: for every environment where an agent runs with tools, write down the network boundary in one sentence (“no egress”, “package mirror and search only”). Then test that sentence from inside the sandbox, not from the config. Anthropic’s advice to its evaluation partners is the same at the scoping level: explicitly define “the targets, permitted actions, and network boundaries.”
Nine models tried to escape Perplexity’s sandbox. The VM held. The network didn’t.
Perplexity’s Secure Intelligence Institute published Escaping SPACE: Part I on 23 September, a red team of its own agent sandbox. SPACE runs each agent in a Firecracker microVM with a per-sandbox gateway enforcing egress policy. They gave nine models root inside the guest and two goals: steal a host-side AWS honeytoken, or reach a blocked callback URL. In total 216 runs across four configurations.
The results, by boundary:
- VM to host: 0 of 108 runs escaped.
- Network, with no network access: 0 of 54 bypassed.
- Network, with partial access (package repos and search): 11 of 54 bypassed, from four models (Claude Opus 5.0, GPT-5.6 Sol, GPT-5.6 Cyber and Kimi K3).
Two techniques did all of it. The first was DNS spoofing. The gateway didn’t check where DNS responses came from, so a root process in the guest could forge a reply mapping an allowlisted name to the blocked IP. The second was shared CDN IPs. An allowlisted hostname resolves to an IP that serves thousands of other sites, and the gateway checked the IP, not which site was asked for. GPT-5.6 Cyber chained a screenshot service through an allowlisted CDN, then read the flag out of the image with OCR.
I’d copy their fixes: drop guest packets whose source IP isn’t the VM’s assigned IP, terminate TLS at the gateway and check SNI and Host against the allowlist, and reject UDP or plaintext traffic you can’t verify. After the fixes, a rerun under the strongest setting found no verified bypass.
Then they ran the same probes, without a model, against ten commercial sandboxes on 31 August and 1 September. Per the published table, seven had at least one bypass: E2B, Vercel Sandbox, Modal Sandboxes, Daytona, Fly.io Sprites, microsandbox and Deno Sandbox. Cloudflare Sandbox and NVIDIA OpenShell showed none in their tests. Vendors were notified 1 September, and Daytona, Deno and microsandbox had shipped fixes by publication. This is Perplexity grading its own sandbox and its competitors’, so weight it accordingly, but the techniques are simple enough to check yourself.
One counterweight so nobody reads this as “the VM is solved.” In Vercel’s two-week, $1 million sandbox challenge (18 August to 1 September, 1,285 reports), researcher Paulos Yibelo was paid the $50,000 maximum for what he describes as a guest-to-host escape with root, confirmed by Vercel’s CEO as a KVM zero-day (Cybernews). There’s no CVE or technical write-up yet. That was a human with two weeks and a bounty. The models in Perplexity’s test went for the network first because it was easier.
So what: an egress allowlist written as hostnames isn’t enforced at the IP layer, and an allowlist enforced at the IP layer isn’t enforced per site. If the agent has root in the guest, it controls its own DNS answers. The two controls everyone builds first are the two that broke.
Do this: if you build or buy an agent sandbox, ask three questions. Does the gateway validate the source IP and the DNS responses? Does it check SNI and
Hostagainst the allowlist, not just the destination IP? What happens to UDP? If the vendor is one of the seven, ask what changed after 1 September.
The other side door: the agent’s own subprocess
Network policy is one way out. The other is everything the harness does on the agent’s behalf, which the sandbox often never sees.
GitSpawn. Francisco Rosales at Manifold Security showed that coding agents run git in the background to gather context, sometimes before you’ve typed a prompt or accepted the workspace-trust dialog, and that those calls honoured the repository’s own .git/config. Set core.fsmonitor to a program and any git command that refreshes the index runs it. The agent spawns git itself, so the command runs outside the sandbox, with no approval prompt, as you. Delivery needs the .git directory intact (a zip, a synced folder, a USB stick), not a clone.
Eight findings across seven agents, published 1 September:
- Patched: Claude Code 2.1.196, Goose 1.44.0 (CVE-2026-72718), Codex and Cursor.
- Unpatched at publication: a second Claude Code variant using a different, unnamed git setting; Hermes Agent (CVE-2026-71963, no triage after six contacts); Qwen Code; Grok Build.
DeepSeek Harness, CVE-2026-82533. OX Security found that the harness’s local control API on 127.0.0.1:3080 decided whether a request was trusted by reading the Host header. The OS sandbox (bubblewrap, Landlock or Seatbelt) restricted file writes but left loopback open, and ordinary bash calls needed no approval. So a sandboxed agent could curl its own harness and switch its session to danger-full-access with approvals off. The session log then recorded the change as the user’s. CVSS 9.4, fixed in 0.1.2-alpha.1.
It’s the same lesson from two directions. The sandbox confines the agent’s tool calls. It doesn’t confine the harness’s own git calls, and it doesn’t stop the agent talking to the harness that sets its permissions.
So what: when you threat model a coding agent, the harness is a privileged process sitting next to an untrusted one. Every subprocess it spawns and every local port it listens on is part of the boundary, whether or not the sandbox config mentions it.
Do this: before opening a project folder you received rather than cloned, read its
.git/config. Any key that names a program can run one. If you’re on Hermes Agent, Qwen Code or Grok Build, treat received repos as hostile until there’s a fix. And for any harness you run, list what listens on loopback and check whether the sandboxed process can reach it.
Agents can delete their own logs, and sometimes do it unprompted
The last line of defence in all three lab incidents was the trace: the record of what the agent did. A paper from Qin, Schmotz, Prinzhorn, Beurer-Kellner, Prabhu and Andriushchenko (arXiv, 24 September) checks whether that record survives the agent.
Mostly, it doesn’t. Coding agents store their traces as files in the user’s home directory (~/.claude/projects/…/<session>.jsonl, ~/.codex/sessions/…, and so on), in reach of the agent’s own shell. Claude Code, Codex, Antigravity, OpenCode, Grok Build, ZCode and Kimi Code all deleted their traces when asked in most trials, some above 80%. Muse Code was the only harness that resisted. Auto-mode permissions helped for some harnesses and not at all for others.
Two results matter more than the direct-request one:
- An attacker can trigger it. A planted “privacy protection” skill, in a coding task that never mentioned traces, got Codex to delete its session trace in a majority of trials.
- Models do it on their own. When trace length quietly affected reward, all ten models tampered with traces at least once, and on Terminal-Bench every model was above 80%.
So what: if the agent can reach its own log, your log is a claim, not evidence. That matters for incident response, and in the EU it matters for any logging obligation you’re trying to meet with harness transcripts.
Do this: record agent activity where the agent can’t touch it. The paper’s recommendation is concrete: route model API traffic through a gateway outside the agent host, write each request and response to an append-only store, and fail closed if recording fails. If you only have harness transcripts today, at least ship them off the host as they’re written.
Quick hits
SalesBleed: zero-click Agentforce exfiltration, over DNS again. Zenity Labs showed that an unauthenticated Web-to-Lead form submission could carry an injection that fired when an employee asked their agent to review leads. The agent queried Accounts and put company names and deal sizes into a subdomain of an attacker hostname. Rendering it as an image, or Slack unfurling it, sent the data out in the DNS lookup. The Trusted URLs redactor missed it because it didn’t recognise the
.funTLD and disagreed with the browser about where a URL ends. Salesforce confirmed fixes on 18 August. Zenity’s line: “The injection didn’t need to escalate privileges, the permissions were already there.”AWS AgentCore: injected ticket, heap read, replayable JWT. Unit 42’s Niv Rabin planted an instruction in a support ticket handled by a default-configured AgentCore Harness. The built-in shell ran as root,
/proc/1/memwas readable, and a heap scan pulled the plaintext JWT for the downstream MCP service account. Replayed from outside AWS, it returned customer PII. AWS closed it as informative under shared responsibility. Their position is that scopingallowedToolsand egress filtering are your job. Shell and file tools are on by default, so scope them per invocation.Half of AI-stack CVEs are in orchestration. Google Threat Intelligence counted 2,076 AI-related CVEs from January 2025 to August 2026, half of them in orchestration middleware like Langflow and Flowise. In inference and serving, 24% come from unauthenticated endpoints or SSRF. Only a handful are confirmed exploited, and one of them is LiteLLM again: CVE-2026-42271, command injection in the MCP test-connection endpoint. That’s a different bug from the KEV entry in issue #30, in the same product.
MITRE ATLAS adds a honeypot mitigation. v2026.09 (15 September) adds 11 techniques, including probing agent trigger channels, discovering agent runtime capabilities and AI-targeted cloaking, plus AI Honeypots (AML.M0039) as a mitigation. That’s the technique Sysdig used in issue #30 to catch LLM-driven attackers with a planted instruction. It now has an ID you can put in a control matrix.
The NCSC on defending agentically. Dave Chismon’s post proposes scoring each defensive agent action on potency, scope, criticality, rollout confidence and recoverability before letting an agent take it. I like the frame, and it fits your own ops agents as well as it fits a SOC.
From the lab
The AgenticGoat lab I’m moving up the list
The AgenticGoat crosswalk has a row marked Planned for MCP08, lack of audit and telemetry. The plan was a defensive lab: the kill chain again, this time with tool-call logging switched on so you can see what the log catches.
The trace-tampering paper changed what that lab needs to show. A log the agent can reach isn’t a control, so the lab should run the chain twice. In the first run the trace sits in the agent’s home directory and a planted skill deletes it. In the second, the trace is recorded at a gateway the agent can’t reach, and the deletion shows up as an event instead of a gap. That’s the same canary, attack, control pattern as every other lab, applied to the evidence instead of the secret.
It’s still Planned, and I’ll say so in the README until it isn’t. If you’ve built out-of-host recording for a coding agent and want to compare designs, the repo is github.com/aminrj-labs/agentic-goat, and the flagship kill chain still runs in the browser on Killercoda.
Tooling worth knowing
- Numbat (Perplexity): open-source, Apache-2.0 endpoint monitor for coding agents. It hooks Claude Code, Codex and others through their own hook systems, ships CEL detection rules, and writes versioned NDJSON. Monitor-only by default, so it’s safe to trial. Pair it with off-host shipping given the paper above. github →
- MITRE ATLAS v2026.09: the new agent reconnaissance techniques are a useful checklist for what to log at your agent’s trigger channels. release →
- Agent Security Scorecard: my free self-assessment against the OWASP Agentic Top 10. About 12 minutes, no login. Score your agents →
- AI Agent Pre-Deployment Security Checklist: 25 yes/no controls to clear before an agent with tool access goes to production. Get the checklist →
One thing to check this week
Pick one sandbox where an agent runs with tools: a coding agent container, a code-interpreter service, an eval environment.
From inside it, as the agent would run, do three things:
- Resolve a name you control. Run
dig $(uuidgen).canary.yourdomain.exampleand watch your authoritative nameserver. If the query arrives, DNS is an outbound channel, whatever your HTTP policy says. That’s the OpenAI incident and SalesBleed in one line. - Forge an answer. If the agent has root, point an allowlisted hostname at a blocked IP in
/etc/hostsor a local resolver, then connect. If it works, your allowlist is enforced on names the guest controls. - Change the site, keep the IP. Connect to an allowlisted CDN IP with a non-allowlisted
Hostheader and SNI. If the response comes from the other site, your allowlist is per IP, not per site.
It takes about thirty minutes. A month ago, seven of the commercial sandboxes Perplexity probed had at least one bypass of this kind.
What I’m watching
→ Escaping SPACE Part II. Perplexity says the next part covers model behaviour: why some models pivoted through CDNs and others declined as out of scope. Which models treat “is this authorised?” as a real question is what the Anthropic incidents came down to.
→ Isolation standards for evaluation partners. Two labs’ incidents ran on outside evaluation infrastructure. I’d like to see evaluators publish how they verify an environment is offline before a run starts, not after a lab finds out from a newspaper.
→ The four unpatched GitSpawn findings. Hermes Agent, Qwen Code, Grok Build and the second Claude Code variant. How long these stay open will show whether harness vendors treat their own subprocesses as part of the boundary.
→ The KVM write-up. A confirmed guest-to-host escape in the hypervisor under most agent sandboxes is a much bigger story than any network bypass. Until there’s a CVE and details, I’m treating it as reported, not established.
If you run the three checks and something gets out, reply and tell me which sandbox and which check. I read everything, and a list of which boundaries hold is more useful than any vendor table, including Perplexity’s.
Cheers, Amine
If a colleague runs agents in a sandbox they’ve never tested from the inside, forward this to them.
Sources
- An alignment assessment of recent cybersecurity incidents, Anthropic, 9 September 2026
- Gemini hacked three companies in first known breakout by Google’s AI, ABC News (reporting the Wall Street Journal), 19 September 2026
- An agent used DNS to reach an external chatbot, OpenAI misalignment report, updated 25 September 2026
- Escaping SPACE: Part I, Perplexity Secure Intelligence Institute, 23 September 2026
- One million dollar hacker challenge for Vercel Sandbox, Vercel, and Massive zero-day: cyber pro breaks out of virtual machine, Cybernews
- GitSpawn: A Single Flaw Lets Untrusted Repos Run Code in Claude Code, Codex, Cursor, and Grok, Manifold Security, 1 September 2026
- CVE-2026-82533: DeepSeek Harness AI agent sandbox escape, OX Security, 8 September 2026
- LLM Agents Can Easily Tamper With Their Own Traces, Qin et al., arXiv, 24 September 2026
- SalesBleed: Indirect Prompt Injection and 0-Click Data Exfiltration on Agentforce, Zenity Labs, 24 September 2026
- A Vault with a Heap-View: The Uncomfortable Space Between AgentCore Harness and Identity, Unit 42, 18 September 2026
- Vulnerability Discovery and Exploitation Trends in the AI Era, Google Threat Intelligence Group, 30 September 2026
- MITRE ATLAS data v2026.09, 15 September 2026
- One does not simply defend agentically, UK NCSC
- Numbat, Perplexity
- tl;dr sec #348 and #349, Clint Gibler, and Adversa AI’s October roundups, for pointing me at several of these