- Ten disclosures, four classes by trigger point: before trust, interaction layer, plumbing, authority.
- In seven of ten, the model's judgment was never the deciding factor. In the other three the model was manipulated, but permissions and configuration decided how much damage followed.
- The loudest signal is GitSpawn: one bug class found independently in seven CLI agents, the third disclosure in five months on the same surface.
- Prompt injection is not solved: the research of the same weeks shows the model layer is porous. But the bugs found, reported and patched this summer cluster at trust boundaries set by configuration, startup and tool plumbing.
- Section [07] turns the ten cases into eight checks you can run on Monday, whether you use coding agents, run MCP servers or give agents authority.
- Part 1 of 5: over the next four Tuesdays we measure those boundaries in a home lab, ending in a preprint in early November.
[00] Result first
Between late June and late September 2026, researchers and vendors disclosed a steady stream of vulnerabilities in AI agents: coding assistants, MCP servers, IDE integrations, CI workflows. Public discussion files most of them under “prompt injection”, as if the problem were always a model that obeyed the wrong text.
We took ten of those disclosures and asked one question of each: where does the flaw trigger? Before the model is involved, in the interface around it, in the plumbing that connects it to tools, or in the authority it was given once it acted.
In seven of ten, the model's judgment was never the deciding factor. Code ran before the trust prompt, an approval dialog showed something other than what would run, or an MCP endpoint simply had no authentication. In the remaining three the model was indeed manipulated, but how much damage followed was decided by permissions and configuration, not by the model.
This is not a claim that prompt injection is solved. The research of the same weeks shows the opposite (section [06]). It is a claim about where defenders should look: the bugs that were actually found, reported and patched this summer cluster at trust boundaries set by configuration, startup and tool plumbing.
This post opens a five-week series that measures those boundaries in a home lab, ending in a preprint in early November.
[01] Scope and method
- Window: 26 June – 24 September 2026. Disclosure date, not exposure date (Kiro was fixed in January and disclosed in August).
- Selection: public disclosures with a primary source (researcher write-up, vendor advisory, CVE record) that affect a tool-using agent or its connectors. Research papers without a product finding are kept separate, in [06].
- Verification: every CVE checked on CVE.org; every write-up opened. Where sources disagree, both are reported.
- Evidence labels: demonstrated (public write-up or vendor-confirmed fix), reported (claimed, not independently checkable).
-
Classification: by trigger point, four classes:
- Before trust — runs at startup or before the user's trust decision.
- Interaction layer — the approval or confirmation step misrepresents or skips what happens.
- Plumbing — classic authentication, authorization or origin flaws in MCP servers and SDKs.
- Authority — the model is manipulated; the impact is set by the privileges it holds.
[02] The ten disclosures
Swipe the table sideways to see all six columns →
| Date | Case | Class | Affected | Status | Source |
|---|---|---|---|---|---|
| 29 Jun | CVE-2026-55607 — sandbox escape via git worktree path confusion | 1 Before trust | Claude Code | Demonstrated (advisory) | CVE.org |
| 1 Sep | GitSpawn — 8 findings in 7 CLI agents: repo git config runs outside the sandbox, on some agents before the trust prompt | 1 Before trust | Claude Code, Goose (CVE-2026-72718), Hermes Agent (CVE-2026-71963), Qwen Code, Grok Build, Codex, Cursor | Demonstrated; Hermes, Qwen Code, Grok Build and one of two Claude Code paths unpatched at Manifold's publication (1 Sep). The Hermes CVE record (3 Sep) references a fix commit | Manifold Security |
| 15 Jul | PromptFiction — a crafted link submits a hidden prompt without user confirmation | 2 Interaction | Claude Desktop (fixed ≥ 1.1.2321) | Demonstrated (vendor fix) | Oasis Security |
| 15 Jul | DeepJack — deeplinks hide an MCP install command from the approval dialog | 2 Interaction | Cursor 3.4.20, 3.9.8 (Windows) | Demonstrated by researcher; vendor status reported | Adversa AI |
| 9 Jul | CVE-2026-59726 “RufRoot” — unauthenticated MCP endpoints expose command execution and agent memory | 3 Plumbing | Ruflo < 3.16.3 | Demonstrated (advisory, CVSS 10.0) | CVE.org, GHSA; “RufRoot” name from Noma Labs |
| 15 Jul | CVE-2026-59950 — WebSocket transport without Host/Origin validation | 3 Plumbing | MCP Python SDK < 1.28.1 | Reported with patch | CVE.org, GHSA |
| 23 Jul | CVE-2026-15015 — OAuth dynamic registration yields an administrator-bound token | 3 Plumbing | MountDev AI MCP Connector for WordPress ≤ 1.6.1 | Reported by CNA (CVSS 9.8) | CVE.org, Wordfence |
| 6 Jul | GitLost — a public issue drives an over-privileged agentic workflow to publish private repository content | 4 Authority | GitHub Agentic Workflows (specific configuration) | Demonstrated | Noma Labs |
| 21 Jul | Confused deputy via hidden PR content; content-isolation defence applied to some tools but not the PR tool | 4 Authority | Azure DevOps MCP server (tested with Copilot CLI and Claude Code) | Demonstrated | Manifold Security |
| 27 Aug | Kiro — workspace content leads to data exfiltration through IDE features | 4 Authority | Kiro IDE 0.7.45 (fixed 0.8.140, January) | Demonstrated | Mindgard |
[03] Case by case
The table compresses each case into a line. This is what each one involved, grouped by class and taken from the primary sources listed at the end.
Class 1 — before trust
CVE-2026-55607: Claude Code, 29 June
Anthropic's advisory (GHSA-7835-87q9-rgvv) describes a sandbox escape: confusion over git worktree paths let code run outside Claude Code's sandbox. The failure is in how the tool resolves paths at the edge of its own sandbox, a piece of machinery the user never sees and never approves. Unusually for this list, it came with a CVE and an advisory from the vendor itself.
GitSpawn: one bug, seven agents, 1 September
Manifold Security reported eight findings across seven command-line agents: Claude Code (two separate paths), Goose, Hermes Agent, Qwen Code, Grok Build, OpenAI Codex and Cursor. The mechanism is the same everywhere. While gathering context in the background, the agent runs git without neutralising the repository's own git configuration. A repository received as files, with its .git directory included, can therefore run code with the user's privileges, outside the sandbox, with no approval prompt and, on some agents, before the workspace-trust prompt has even appeared.
The disclosure timeline is as telling as the bug. Between 26 June and 20 July, five of the eight reports were closed as duplicates: someone else had found the same thing first. One earlier report on Grok Build had been closed by xAI as "informative". Hermes Agent was never triaged despite six contact attempts across five channels; its CVE (CVE-2026-71963) was assigned by VulnCheck, not by the vendor. At publication, Goose was fixed in 1.44.0 (CVE-2026-72718), Codex and Cursor were patched, Claude Code had fixed its startup path in 2.1.196 but not the second one, and Hermes, Qwen Code and Grok Build were still exposed. The Hermes CVE record, published two days later, references a fix commit.
GitSpawn is also not the first time this surface gave way. In May, GitHub Copilot CLI had its own case of a repository's git metadata leading to command execution (CVE-2026-45033). Three disclosures in five months on the same surface, across unrelated vendors, point to a design assumption rather than an implementation slip: that a repository is data, when for a tool that shells out to git it is partly configuration.
Class 2 — interaction layer
PromptFiction: Claude Desktop, 15 July
Oasis Security showed that a crafted app-launch link could deliver a prompt that Claude Desktop submitted without the user seeing or confirming it, with the harmful part hidden below harmless-looking text. The researchers name two impacts: exfiltration of conversation history on a standard install, and possible code execution where the app has file access. Anthropic fixed it after a report through its disclosure programme. A prompt arriving through such a link now waits, pre-filled, until the user reads it and presses send; the fix is in 1.1.2321 and later. The remedy tells you where the flaw was: nothing changed in the model, the human was put back in the loop.
DeepJack: Cursor, 15 July
On the same day, Adversa AI described two variants of a cursor:// deeplink that open Cursor's MCP-install confirmation dialog with the server command partly out of view: the dialog shows it in a single-line field, so the user approves a command they cannot fully read. The researchers present it as a bypass of the fix for an older deeplink bug (CVE-2025-54133). They first observed it on Cursor 3.4.20 on Windows 11 and found 3.9.8 still affected on publication day. Cursor closed both reports as duplicates of an April report; we found no vendor advisory or new CVE. Approval dialogs are a security boundary, and this one regressed.
Class 3 — plumbing
CVE-2026-59726 "RufRoot": Ruflo, 9 July
In Ruflo's default docker-compose deployment, the MCP bridge endpoints were reachable without authentication. The CVE record lists what followed: command execution in the bridge container, reading of the provider API keys, and poisoning of the agent's learning store. GitHub, acting as CNA, scored it 10.0; the fix is in 3.16.3. The name comes from Noma Labs. No model is involved at any step: this is an unauthenticated service that happens to be able to act.
CVE-2026-59950: MCP Python SDK, 15 July
The deprecated WebSocket server transport in the official MCP Python SDK accepted handshakes without validating the Host or Origin headers, the classic opening for cross-site WebSocket abuse. The patched version is 1.28.1. One detail matters for anyone who upgrades: according to the GitHub advisory, the new validation setting defaults to off. Upgrading alone does not turn the protection on; it has to be configured, or the server moved to the Streamable HTTP transport.
CVE-2026-15015: WordPress MCP connector, 23 July
In the MountDev AI MCP Connector for WordPress, up to and including 1.6.1, an unauthenticated attacker could obtain an administrator-bound OAuth bearer token by registering their own client and using an authorization endpoint that did not check who was asking. Wordfence, the CNA, scored it 9.8. The vendor was notified on 8 July and the flaw disclosed on 22 July. Five days later the MCP 2026-07-28 revision deprecated Dynamic Client Registration, the mechanism at the start of this chain, in favour of Client ID Metadata Documents.
Class 4 — authority
GitLost: GitHub Agentic Workflows, 6 July
Noma Labs planted text in an issue on a public repository. An agentic workflow that could read the organisation's repositories and write public comments followed it, and published private repository content in a public comment. The research was disclosed to GitHub. Here the model was manipulated, but the outcome was set by a configuration that combined two powers that should never meet in one job: reading private data and writing in public.
Azure DevOps MCP server, 21 July
Manifold Security hid an instruction in an invisible comment inside a pull-request description. Microsoft's official Azure DevOps MCP server had a content-isolation defence, spotlighting, applied to some tools such as pipelines and wiki, but not to the tool that returns pull-request descriptions. The researchers validated the result with two different agents, Copilot CLI and Claude Code, and MSRC acknowledged and triaged the report. The defence existed; it was simply not applied to every door.
Amazon Kiro, 27 August
Mindgard showed that attacker-controlled workspace content, delivered through workspace settings and Kiro's Powers feature, could steer the agent into exfiltrating sensitive local data, and that it worked in both trusted and untrusted workspaces. Testing was on Kiro IDE 0.7.45 on Windows. Amazon fixed it in 0.8.140 in January 2026; no CVE had been assigned when the research was published in August, seven months later. Mindgard framed the post around what that delay says about disclosure for AI tools.
[04] Where the lines are hard to draw
Any classification like this one involves judgment calls, so here is the rule we applied and the places where it was hard to apply.
The rule. If the attack only works because the model follows instructions it should not have trusted, the case is class 4, whatever the delivery channel. Otherwise, the case goes to the earliest step that fails without any model decision: before the trust decision (1), in what the user is shown or asked to approve (2), or in the authentication and authorization of the services around the agent (3).
Kiro. The payload travels through workspace settings, which argues for class 1. But the exfiltration only happens because the agent acts on injected instructions, so by our rule it is class 4. Moved to class 1, the headline would read eight of ten.
PromptFiction. The model does act on text it should not have received. We kept it in class 2 because the text arrives as if the user had typed it: the model is not fooled into obeying a document, it obeys what looks like its user, and the fix was a confirmation step, not a model change. Moved to class 4, the headline would read six of ten.
Counting. GitSpawn counts once, although it covers eight findings. Counting findings instead of disclosures would put nine entries in class 1 and make our point look stronger than a fair count allows, so we did not.
Whichever way the two borderline cases go, the result holds between six and eight out of ten: in most of this summer's disclosures, the model's judgment was not what decided the outcome. For the preprint in November we will repeat the classification with criteria written in advance and a second, independent classifier, and report how often the two agree.
[05] What the pattern says
Class 1 is the loudest signal. GitSpawn is one bug class found independently in seven agents: background context gathering runs git without neutralising the repository's own configuration. It is the third disclosure in five months on the same surface — after GitHub Copilot CLI in May (CVE-2026-45033) and Claude Code in June (CVE-2026-55607). The response pattern matters as much as the bug: several reports were closed as duplicates, one as “informative”, one was never triaged after six contact attempts. A class that many teams hit at once is a design gap, not an implementation slip.
Class 2 moves the attack into the interface. In both cases the user is asked to approve something, and what they see is not what runs. Human-in-the-loop only works if the loop shows the truth.
Class 3 is web security, at agent scale. Missing authentication, missing origin checks, an over-permissive OAuth flow: none of this is new. What is new is that the endpoint behind it can execute commands or write into an agent's memory. The MCP 2026-07-28 revision hardens authorization and deprecates dynamic client registration, but it states plainly that the protocol cannot enforce consent or access control by itself — that stays with implementers.
Class 4 is where prompt injection actually lives — and even there, the deciding variable was authority. GitLost needed a workflow that could both read private repositories and write public comments. The Azure DevOps server protected some tools and not the one that mattered.
[06] Meanwhile, the model layer is not getting safer
The research of the same weeks shows why “the model will refuse” is not a boundary:
- Splitting a payload across two MCP channels took some models from 0% compliance to 100% exfiltration, and seven MCP security scanners missed the fragments (arXiv 2609.18217, 16 Sep).
- Delayed, conditional injections succeeded 43–83% of the time on nine production agents, against ≤3% for direct instructions (arXiv 2609.22510, 18 Sep).
- Defences that score well on attack success can still leave unnecessary privileges open, and the two metrics do not predict each other (Ajar, arXiv 2609.26900, 22 Sep).
Put together with the table: the model layer is porous, and the bugs being found are around it. Both push the same way — measure the boundaries.
[07] What to check on Monday
None of this needs a lab to act on. If you use or run AI agents, these are the checks the ten cases point to. They are defensive measures; none of them requires knowing how the attacks work in detail.
If you use coding agents
- Treat a repository as executable input. Open code you did not write, such as archives, forks, take-home tests or bug-report attachments, in a disposable VM or container before any agent touches it. After GitSpawn, a repository with its
.gitdirectory is not just data. - Check your versions against the fixes. Goose 1.44.0 or later, Claude Code 2.1.196 or later for the startup path (the second path was still open on 2.1.252 when GitSpawn was published, so check Anthropic's release notes), Claude Desktop 1.1.2321 or later. For Hermes Agent, Qwen Code and Grok Build, check the vendor's notes: no fix was public when GitSpawn was published.
- Read the whole thing before you approve. Expand the full command before approving an MCP server install, and do not approve anything that arrives from a link you did not expect. An approval dialog is only a control if it shows you everything.
- Review agent configuration like CI configuration. Files that declare MCP servers or agent settings inside a project, for example
.mcp.json,.vscode/mcp.json,.cursor/,.gemini/or.claude/, decide what runs on a developer's machine. They deserve the same code review as a build pipeline.
If you run MCP servers
- No MCP endpoint without authentication, including the ones that are "only" reachable inside a docker-compose network. If you run Ruflo, update to 3.16.3 or later.
- Upgrading is not always enough. On the MCP Python SDK's WebSocket transport, 1.28.1 adds Host and Origin validation but leaves it off by default: configure it explicitly, or move to Streamable HTTP.
- Plan the move away from Dynamic Client Registration. The 2026-07-28 revision deprecates it in favour of Client ID Metadata Documents, with at least twelve months before removal. If you run the MountDev WordPress connector, move past 1.6.1.
- Apply your content defences to every tool. If your server has a helper that isolates third-party content, check that every tool returning third-party content goes through it, and add a test that fails when a new tool does not.
If you give agents authority
- Never combine private read and public write in one agent job. GitLost needed both. Split them, and give each workflow its own least-privilege token.
- Log what agents start, not only what they say. Several of these cases leave no trace in the conversation, because nothing in them went through the model. Process and network logs on developer machines are where they show up. We come back to this on 27 October.
[08] What comes next
Over the next four Tuesdays we measure those boundaries on current versions, in a lab made only of our own repositories, servers and virtual machines, with harmless canaries:
- 6 Oct — does trusting a repository now mean trusting the MCP servers it declares?
- 13 Oct — when an MCP server changes a tool after you approved it, does your client tell you?
- 20 Oct — do official MCP servers apply their own content defences to every tool, or only some?
- 27 Oct — what can a defender actually see of all this in default logs?
If a test finds something new in a product, the post reports its class and disclosure status only; details follow the fix or 90 days.
Sources
- Manifold Security, “GitSpawn”, 1 Sep 2026 — manifold.security
- Manifold Security, Azure DevOps MCP confused deputy, 21 Jul 2026 — manifold.security
- Oasis Security, PromptFiction — oasis.security
- Adversa AI, DeepJack — adversa.ai
- Noma Labs, GitLost — noma.security
- Mindgard, Amazon Kiro — mindgard.ai
- Noma Labs, RufRoot (CVE-2026-59726) — noma.security
- GitHub Security Advisories: GHSA-7835-87q9-rgvv (Claude Code, CVE-2026-55607), GHSA-c4hm-4h84-2cf3 (Ruflo, CVE-2026-59726), GHSA-vj7q-gjh5-988w (MCP Python SDK, CVE-2026-59950)
- Wordfence, MountDev AI MCP Connector (CVE-2026-15015) — wordfence.com
- CVE.org: CVE-2026-55607, CVE-2026-71963, CVE-2026-72718, CVE-2026-45033, CVE-2026-59726, CVE-2026-59950, CVE-2026-15015
- MCP specification 2026-07-28 and changelog — specification · changelog
- arXiv 2609.18217, 2609.22510, 2609.26900