Post

Why Your CVE Program Misses Most AI Agent Vulnerabilities

CVEs still name code bugs well. For AI agents and MCP servers, most real risk sits in defaults, prompt injection, tool composition and supply chain, where a CVE says little or nothing. A map of where it helps, where it misleads, and the triage, inventory and controls that replace it.

Why Your CVE Program Misses Most AI Agent Vulnerabilities

When we gave an AI agent read access to our whole Kubernetes fleet, the vulnerability scanner had nothing to say about it. Every package was current. No open CVE touched the deployment.

The review found four ways it could still hurt us, and I walked through them at fwd:cloudsec in September. Its read scope reached Secrets. It shared a service account, so nobody could attribute what it read. The model filled in the target cluster as a tool parameter, which made it a confused deputy. Nothing capped how much it could pull into one response. None of the four has a CVE, and none ever will, because no component is broken. The risk is in how the pieces were put together.

That is the gap this post is about. I work in the OWASP Agentic Top 10 and MCP Top 10 working groups, and the question I hear most from security leaders is some version of: “Our vulnerability program is built on CVEs. Does it cover our agents?” Partially, and less than you think. Here is the map I use to decide where the CVE helps, where it misleads, and what has to do the rest of the job.

TL;DR

  1. Agent and MCP flaws live in six places, and a CVE fits only two of them well. The further a flaw sits from a single line of code, the less the record tells you and the more of the fix falls on you.
  2. A CVE closes an instance, not a class. In the MCP git server I study, the CVE fix guarded two tools. Four more were guarded three months later, with no CVE.
  3. Severity belongs to your deployment. The same CVE is a footnote or a breach depending on the agent’s tool grants, identity and egress.
  4. The pipeline behind the CVE is saturated, so expect slower records, thinner enrichment and exploitation before the patch.
  5. In the EU, disclosure is now a legal clock: 24 hours for an early warning on an actively exploited vulnerability, since 11 September 2026.

Six places agent flaws live

I sort every agent or MCP finding by where the flaw sits, because that decides whether a CVE will ever describe it, who can fix it, and what actually closes it.

Six places agent flaws live, from implementation bugs where a CVE fits cleanly down to toxic composition and malicious packages where it does not fit at all, with who can fix each and what closes it

#Where the flaw sitsReal exampleCVE?Who fixes itWhat closes it
1Implementation bug in an MCP server or clientmcp-server-git path traversal and argument injection (CVE-2025-68143, -68144, -68145). nginx-ui’s /mcp_message endpoint shipped without auth (CVE-2026-33032, CVSS 9.8, exploited).YesVendorPatch
2Trust and approval logic in an agent hostCursor did not re-prompt when an approved MCP config changed (MCPoison, CVE-2025-54136). Copilot could be injected into turning on its own auto-approve (CVE-2025-53773).Yes, per hostVendorPatch, then the same class in the next host
3Unsafe default in a protocol or SDKThe STDIO transport in the official MCP SDKs hands configuration to OS process creation. OX Security traced ten downstream CVEs to it, LiteLLM and Windsurf among them.Downstream onlyEvery integratorSafe default upstream, sandbox downstream
4Prompt injectionShareLeak (CVE-2026-21520): a comment on a public SharePoint form made a Copilot Studio agent email customer records outside the company, even when Microsoft’s safety layer flagged the request.One path at a timeNobody, fullyLimits on what a steered agent can do
5Toxic composition across toolsInvariant’s GitHub MCP finding: a public issue steers an agent with a broad token into leaking private repos. “This is not a flaw in the GitHub MCP server code itself.”NoYouRule of Two, scoped tokens, isolation
6Malicious package or rug pullpostmark-mcp shipped 15 clean versions, then added a hidden BCC to every email in 1.0.16.Malware advisoryYou and the registryPinning, provenance, re-approval on change

My fleet agent sat entirely in row 5, with a little of row 4. That is typical. Rows 3 to 6 are where agent incidents come from, and they are where a CVE-driven program is blind. No advisory will tell you that your agent holds a token that reads private repositories and writes public ones.

Row 1 deserves one more sentence: the worst MCP bugs are not AI bugs. When Equixly tested popular MCP servers in March 2025, 43% had command injection flaws. Your AppSec program already knows how to find those. Point it at your MCP servers.

What the CVE record leaves out: one server, five commits

I keep a private corpus of MCP CVEs that I reproduce by hand. The method is simple and humbling: read the NVD one-liner, write down where I think the bug is, then check against the patch. When I did this for the filesystem server’s containment bug (CVE-2025-53110), the flaw was startsWith(allowedDir) with no path separator, in three places. With /tmp/ws/data allowed, a request for /tmp/ws/data-evil/flag.txt passed every check. The record says CWE-22, path traversal. It says nothing about which of your own servers contain the same three lines.

The mcp-server-git history shows the bigger problem. These are public commits in modelcontextprotocol/servers:

DateCommitWhat changed
17 Dec 20259e5d5b8, a37158bThe CVE fixes. CVE-2025-68144 names git_diff and git_checkout, which passed flag-like values straight to the git CLI.
29 Dec 2025db96050Better path validation in git_add
15 Mar 2026ae40ec2“add missing argument injection guards to git_show, git_create_branch, git_log, and git_branch”
14 Jun 20260588ec0“harden git_add (defense-in-depth)”

Your scanner marked you clean the day you upgraded past the December fix. Four more tools in the same server took a value from the model and lacked the guard for another three months, and the fix for them carries no CVE. The CVE described the instance in the report. The class was the whole server.

Argument injection is also a good reminder of why reviews miss this class. No shell is involved, so every “never use shell=True” check passes it.

If you report findings, the lesson runs the other way: report the class, not the instance. If the root cause is an SDK default or a host’s approval model, say so and show the variants. Ten downstream CVEs for one upstream default means the reports stopped a level too early.

Why the instruments misfire

The scores disagree, and none of them is yours. The same bug, CVE-2025-68144, is 7.1 from NVD under CVSS 3.1, 6.3 from GitHub under CVSS 4.0, and 8.1 in some press coverage. Any of those numbers describes the component alone. In practice:

  • An agent that runs this server against your own repos, with no outside content and no egress, carries little real risk.
  • An agent that triages pull requests from external contributors, with a token that can push and a web fetch tool, turns the same CVE into a path to file overwrite, code execution and exfiltration.

Only you know your tool grants, so environmental scoring is no longer optional. The CSA’s draft CNA manual for agentic AI moves the same way: it scores impact “against the systems accessible to the agent through its authorized tool grants,” and adds metrics for authorization footprint and delegation chain depth.

The product field does not fit. The same draft notes that “the affected component is often a behavior class rather than a discrete software package version,” which breaks CPE matching. A record for row 3 or 4 tells you which product got patched, not whether your integration has the same shape.

The taxonomy is thin. CWE has CWE-1427 for prompt injection, and CWE 4.20 added an AI/ML view with no new weakness entries. Nothing covers delegation chain escalation, memory poisoning or approval bypass. For triage I use the OWASP Agentic Top 10 and MCP Top 10 as the working taxonomy and keep the CWE for the record.

The pipeline is saturated. Four numbers set the background:

The bottleneck is no longer finding bugs. It is fixing them and deciding which ones matter to you.

The operating model

Five workstreams, each with something you can run or copy.

1. Inventory: find the MCP servers you did not install

Start with the developer laptops. This lists every MCP config file under a home directory and the servers in each:

1
2
3
4
5
6
7
8
9
find ~ -maxdepth 8 \( -name node_modules -o -name .git -o -name .venv -o -name venv \) -prune -o \
  \( -name mcp.json -o -name .mcp.json -o -name claude_desktop_config.json \
     -o -name mcp_config.json -o -name .claude.json \) -type f -print 2>/dev/null |
while read -r f; do
  echo "== $f"
  jq -r '[.. | objects | (.mcpServers? // .servers? // empty) | objects | keys[]] | unique[]' "$f" 2>/dev/null
done
# Codex keeps servers in TOML:
grep -h '^\[mcp_servers\.' ~/.codex/config.toml 2>/dev/null

I ran it on my own workstation before writing this. It found seven server definitions. None had a secret pasted inline, which was a relief. Five were catalog entries for plugins I had never turned on. The other two were in configs I never wrote: they shipped inside third-party repositories I had cloned, one of them an unpinned npx server that an IDE would offer to start the moment I opened the folder. That is row 6 arriving through git clone, and it is the part of shadow MCP (MCP09) nobody counts.

Then record each agent with the four fields that decide severity. If you cannot fill in egress, that agent fails the review on its own:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
- agent: pr-triage-agent
  owner: platform-security
  identity: sa-pr-triage (own service account, not shared)
  mcp_servers:
    - name: git
      package: pkg:pypi/[email protected]
      transport: stdio
      tool_hash_baseline: baselines/git.json
    - name: fetch
      package: pkg:pypi/[email protected]
      transport: stdio
  tool_grants: [git_diff, git_log, git_show, fetch]
  data_reach: [public repos, private repo "payments-api" read]
  untrusted_inputs: [pull request bodies, issue comments, fetched web pages]
  egress: [api.github.com, any URL via fetch]   # this line is the finding

2. Intake: stop treating NVD as the source of truth

Most MCP servers and agent frameworks are npm, PyPI or Go packages. GitHub Security Advisories is usually first and package-native. OSV.dev adds malicious package entries (the MAL- IDs), which covers row 6. Run it against every MCP server you deploy:

1
osv-scanner -r ./path/to/mcp-server

For exploitation signal, use CISA KEV and VulnCheck KEV together. nginx-ui’s MCP flaw reached VulnCheck KEV, and LiteLLM’s MCP auth bypass became the first MCP implementation in CISA KEV on 2 September. Then add the source that never shows up in a feed: research on classes. Invariant’s toxic flow write-up has no feed entry and matters more to your agents than most CVEs that do.

3. Triage: three questions, every combination answered

I triage per deployment, not per CVE. The CVSS number doesn’t enter into it:

1
2
3
4
5
6
7
8
9
10
11
12
13
Q1. Does attacker-influenced content reach the component?
    (email, tickets, PRs, web pages, documents, other agents)
Q2. Can the agent that uses it write, execute, or send data out?
    Check the hidden paths: rendered images, link previews, DNS, shipped logs.
Q3. Is it exploited (KEV), or is there a public PoC?

              Q3: KEV         Q3: public PoC        Q3: neither
Q1 yes, Q2 yes  Act now        Out of cycle          Next cycle, high priority
Q1 yes, Q2 no   Out of cycle   Next cycle            Next cycle
Q1 no           Out of cycle   Next cycle            Backlog

Act now = patch or disable the tool within hours, and check whether your CRA clock started.
Then, for every cell: is this one instance of a class you also have? If yes, go to step 4.

The last line matters most. A prompt injection CVE in someone else’s product is low priority as an instance and high priority as evidence that your own agents share the class.

4. Class controls: what bounds the risk when no patch exists

These are deterministic controls that sit outside the model. Most of them are rungs on the containment ladder I presented at fwd:cloudsec.

ClassControl that bounds itOWASP
Prompt injectionTreat every model-produced tool argument as attacker-influenced. Meta’s Rule of Two: no session combines untrusted input, sensitive data and the ability to change state or communicate out. Human approval on irreversible actions.ASI01, MCP06
Exfiltration through allowed actionsEgress allowlists, recipient restrictions on send tools, DLP that knows external content triggered the actionASI02, MCP10
Unsafe STDIO launchNever build launch config from untrusted input. Allowlist commands. Run local servers in a sandbox with no ambient credentials.ASI05, MCP05
Over-privileged tokensOne identity per agent. Short-lived, narrow tokens bound to their audience (RFC 8707), no passthrough, as the MCP authorization spec requiresASI03, MCP01, MCP02, MCP07
Tool poisoning and rug pullsPin versions, hash tool definitions, re-approve on any changeASI04, MCP03, MCP04
Memory poisoningProvenance on every memory write, separate memory per trust domain, expiryASI06

Two of these are worth showing in full.

Read-only is not a safe default. It was the word that got my fleet agent approved, and I wrote up what it actually reaches. A read-only agent with an egress path is an exfiltration tool. Classify by reach, not by write permission.

Hash tool definitions, and compare across sessions. In my rug-pull lab, a server behaves for five sessions and then adds a hidden instruction to one tool description. Static scanners miss it, because the server’s code never changes. A hash taken at every session start catches it. Here is the minimal version, on the official Python SDK:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
"""Hash an MCP server's tool definitions and flag drift from the approved baseline.

Usage: python tool_hash.py baseline.json -- <server command> [args...]
"""
import asyncio, hashlib, json, pathlib, sys

from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client


async def tool_hashes(cmd, args):
    params = StdioServerParameters(command=cmd, args=args)
    async with stdio_client(params) as (read, write):
        async with ClientSession(read, write) as session:
            await session.initialize()
            tools = (await session.list_tools()).tools
    return {
        t.name: hashlib.sha256(
            json.dumps(t.model_dump(mode="json", by_alias=True, exclude_none=True),
                       sort_keys=True).encode()
        ).hexdigest()
        for t in tools
    }


def main():
    baseline_path = pathlib.Path(sys.argv[1])
    cmd, *args = sys.argv[sys.argv.index("--") + 1:]
    current = asyncio.run(tool_hashes(cmd, args))
    if not baseline_path.exists():
        baseline_path.write_text(json.dumps(current, indent=2, sort_keys=True))
        print(f"baseline written: {len(current)} tools")
        return
    baseline = json.loads(baseline_path.read_text())
    drift = sorted(n for n in current.keys() | baseline.keys()
                   if current.get(n) != baseline.get(n))
    for name in drift:
        print(f"DRIFT {name}: {'added' if name not in baseline else 'removed' if name not in current else 'changed'}")
    sys.exit(1 if drift else 0)


if __name__ == "__main__":
    main()
1
2
pip install mcp
python tool_hash.py baselines/everything.json -- npx -y @modelcontextprotocol/server-everything

The first run writes the baseline. Every later run exits non-zero on drift, so it can gate a CI job or an agent’s startup. I tested it against the reference server: 13 tools recorded, a clean re-run, and a flagged change when I altered one hash. It hashes the whole tool definition, so a changed title or annotation counts too.

5. Detection and disclosure

Log what a patch cannot close. For every tool call, record the agent identity, the delegated user, the tool, the full arguments, the result size, and where the content that preceded the call came from: the user, a retrieved document, or another agent. That last field is how you tell “the user asked for this” from “a web page asked for this”. Ship it off the agent host as it is written, because an agent that can reach its own trace can delete it.

Decide your CRA position before an incident. Since 11 September 2026, manufacturers selling products with digital elements in the EU report actively exploited vulnerabilities through ENISA’s Single Reporting Platform: early warning in 24 hours, notification in 72, final report within 14 days of a fix. Open-source stewards follow on 11 December 2027. The question every agent vendor will face is whether an exploited prompt injection path counts. Write the answer down now. A starting sentence you can adapt:

We treat a prompt injection path in our product as an actively exploited vulnerability when we have reliable evidence that an attacker used it, in any deployment, to make the product take an action or disclose data that its operator did not authorize.

Say what you accept in your disclosure policy. Researchers need to know whether prompt injection is in scope. A paragraph to adapt:

Prompt injection is in scope when a specific, documented input causes the product to take an action, call a tool, or disclose data outside what the operator configured. Please include the input, the configuration, and the tools and data the agent could reach. Model outputs that are merely incorrect or offensive, with no action or disclosure, are out of scope. We assign CVEs for in-scope findings through [our CNA / the relevant CNA].

The first sentence follows the CSA draft’s test: a behavior qualifies when “a specific, documented input pattern” triggers it and a reasonable developer would consider it exploitable. For dependencies that are not reachable in your product, publish VEX (CSAF or OpenVEX). With CVE volume where it is, a “not affected” statement saves your customers more time than anything else you can publish.

Three things to do this week

  1. Run the inventory command on three developer laptops, including your own. Count the servers nobody approved and the configs that arrived inside cloned repos.
  2. Take your last ten agent or MCP findings and sort them into the six rows. If they all sit in rows 1 and 2, your process is only finding what a CVE can describe.
  3. Fill in the inventory entry for one production agent. If the egress line says “any”, you have your first class control to build, and no CVE will ever tell you about it.

What I am watching

  • 28 October: CWE 5.0, with custom lifecycle phases for AI. It’s the first real chance for agentic weakness entries.
  • The CVE program’s AI researcher CNA pilot. Its requirements include human validation, a proposed fix where feasible, and an assessment of “whether the issue is likely to matter in real-world configurations.” That’s the right bar for every report, AI-generated or not.
  • OWASP AIVSS v1.0, promised by the end of the year, for agent-amplified severity.
  • 11 December 2027, when CRA reporting reaches open-source stewards. Most MCP servers are open source.

The takeaway

The CVE is a good instrument for one job: naming a specific flaw in a specific version so everyone can patch it. Keep using it for that. For agents and MCP, the risk lives between components: what the agent can reach, as whom, and what happens when it is steered. My fleet agent had a clean scan and four real problems. No advisory feed will describe that space for you. Your inventory, your triage and your boundary controls have to.

If you want a quick read on your own agents against the OWASP Agentic Top 10, the Agent Security Scorecard takes about twelve minutes. If you want help building the inventory and triage above, get in touch.


References

Examples and history

Pipeline

Scoring, taxonomy and disclosure

My own work referenced

The Security Lab Newsletter

This post is the article. The newsletter is the lab.

Subscribers get what doesn't fit in a post: the full attack code with annotated results, the measurement methodology behind the numbers, and the week's thread — where I work through a technique or incident across several days of testing rather than a single draft. The RAG poisoning work, the MCP CVE analysis, the red-teaming patterns — all of it started as a newsletter thread before it became a post. One email per week. No sponsored content. Unsubscribe any time.

Join the lab — it's free

Already subscribed? Browse the back-issues →

This post is licensed under CC BY 4.0 by the author.