// TL;DR

[00] Result first

Between late June and late September 2026, researchers and vendors disclosed a steady stream of vulnerabilities in AI agents: coding assistants, MCP servers, IDE integrations, CI workflows. Public discussion files most of them under “prompt injection”, as if the problem were always a model that obeyed the wrong text.

We took ten of those disclosures and asked one question of each: where does the flaw trigger? Before the model is involved, in the interface around it, in the plumbing that connects it to tools, or in the authority it was given once it acted.

In seven of ten, the model's judgment was never the deciding factor. Code ran before the trust prompt, an approval dialog showed something other than what would run, or an MCP endpoint simply had no authentication. In the remaining three the model was indeed manipulated, but how much damage followed was decided by permissions and configuration, not by the model.

This is not a claim that prompt injection is solved. The research of the same weeks shows the opposite (section [06]). It is a claim about where defenders should look: the bugs that were actually found, reported and patched this summer cluster at trust boundaries set by configuration, startup and tool plumbing.

This post opens a five-week series that measures those boundaries in a home lab, ending in a preprint in early November.

[01] Scope and method

[02] The ten disclosures

Swipe the table sideways to see all six columns →

Date Case Class Affected Status Source
29 Jun CVE-2026-55607 — sandbox escape via git worktree path confusion 1 Before trust Claude Code Demonstrated (advisory) CVE.org
1 Sep GitSpawn — 8 findings in 7 CLI agents: repo git config runs outside the sandbox, on some agents before the trust prompt 1 Before trust Claude Code, Goose (CVE-2026-72718), Hermes Agent (CVE-2026-71963), Qwen Code, Grok Build, Codex, Cursor Demonstrated; Hermes, Qwen Code, Grok Build and one of two Claude Code paths unpatched at Manifold's publication (1 Sep). The Hermes CVE record (3 Sep) references a fix commit Manifold Security
15 Jul PromptFiction — a crafted link submits a hidden prompt without user confirmation 2 Interaction Claude Desktop (fixed ≥ 1.1.2321) Demonstrated (vendor fix) Oasis Security
15 Jul DeepJack — deeplinks hide an MCP install command from the approval dialog 2 Interaction Cursor 3.4.20, 3.9.8 (Windows) Demonstrated by researcher; vendor status reported Adversa AI
9 Jul CVE-2026-59726 “RufRoot” — unauthenticated MCP endpoints expose command execution and agent memory 3 Plumbing Ruflo < 3.16.3 Demonstrated (advisory, CVSS 10.0) CVE.org, GHSA; “RufRoot” name from Noma Labs
15 Jul CVE-2026-59950 — WebSocket transport without Host/Origin validation 3 Plumbing MCP Python SDK < 1.28.1 Reported with patch CVE.org, GHSA
23 Jul CVE-2026-15015 — OAuth dynamic registration yields an administrator-bound token 3 Plumbing MountDev AI MCP Connector for WordPress ≤ 1.6.1 Reported by CNA (CVSS 9.8) CVE.org, Wordfence
6 Jul GitLost — a public issue drives an over-privileged agentic workflow to publish private repository content 4 Authority GitHub Agentic Workflows (specific configuration) Demonstrated Noma Labs
21 Jul Confused deputy via hidden PR content; content-isolation defence applied to some tools but not the PR tool 4 Authority Azure DevOps MCP server (tested with Copilot CLI and Claude Code) Demonstrated Manifold Security
27 Aug Kiro — workspace content leads to data exfiltration through IDE features 4 Authority Kiro IDE 0.7.45 (fixed 0.8.140, January) Demonstrated Mindgard
Timeline of the ten disclosures, June to September 2026, by class Jul Aug Sep MCP 2026-07-28 1 Before trust 2 Interaction 3 Plumbing 4 Authority CVE-2026-55607 GitSpawn PromptFiction DeepJack RufRoot MCP Py SDK WP connector GitLost Azure DevOps Kiro
The ten disclosures between 26 June and 24 September 2026, by class. The dashed line marks the MCP 2026-07-28 revision. Three cases landed on 15 July.

[03] Case by case

The table compresses each case into a line. This is what each one involved, grouped by class and taken from the primary sources listed at the end.

Class 1 — before trust

CVE-2026-55607: Claude Code, 29 June

Anthropic's advisory (GHSA-7835-87q9-rgvv) describes a sandbox escape: confusion over git worktree paths let code run outside Claude Code's sandbox. The failure is in how the tool resolves paths at the edge of its own sandbox, a piece of machinery the user never sees and never approves. Unusually for this list, it came with a CVE and an advisory from the vendor itself.

GitSpawn: one bug, seven agents, 1 September

Manifold Security reported eight findings across seven command-line agents: Claude Code (two separate paths), Goose, Hermes Agent, Qwen Code, Grok Build, OpenAI Codex and Cursor. The mechanism is the same everywhere. While gathering context in the background, the agent runs git without neutralising the repository's own git configuration. A repository received as files, with its .git directory included, can therefore run code with the user's privileges, outside the sandbox, with no approval prompt and, on some agents, before the workspace-trust prompt has even appeared.

The disclosure timeline is as telling as the bug. Between 26 June and 20 July, five of the eight reports were closed as duplicates: someone else had found the same thing first. One earlier report on Grok Build had been closed by xAI as "informative". Hermes Agent was never triaged despite six contact attempts across five channels; its CVE (CVE-2026-71963) was assigned by VulnCheck, not by the vendor. At publication, Goose was fixed in 1.44.0 (CVE-2026-72718), Codex and Cursor were patched, Claude Code had fixed its startup path in 2.1.196 but not the second one, and Hermes, Qwen Code and Grok Build were still exposed. The Hermes CVE record, published two days later, references a fix commit.

GitSpawn is also not the first time this surface gave way. In May, GitHub Copilot CLI had its own case of a repository's git metadata leading to command execution (CVE-2026-45033). Three disclosures in five months on the same surface, across unrelated vendors, point to a design assumption rather than an implementation slip: that a repository is data, when for a tool that shells out to git it is partly configuration.

Class 2 — interaction layer

PromptFiction: Claude Desktop, 15 July

Oasis Security showed that a crafted app-launch link could deliver a prompt that Claude Desktop submitted without the user seeing or confirming it, with the harmful part hidden below harmless-looking text. The researchers name two impacts: exfiltration of conversation history on a standard install, and possible code execution where the app has file access. Anthropic fixed it after a report through its disclosure programme. A prompt arriving through such a link now waits, pre-filled, until the user reads it and presses send; the fix is in 1.1.2321 and later. The remedy tells you where the flaw was: nothing changed in the model, the human was put back in the loop.

DeepJack: Cursor, 15 July

On the same day, Adversa AI described two variants of a cursor:// deeplink that open Cursor's MCP-install confirmation dialog with the server command partly out of view: the dialog shows it in a single-line field, so the user approves a command they cannot fully read. The researchers present it as a bypass of the fix for an older deeplink bug (CVE-2025-54133). They first observed it on Cursor 3.4.20 on Windows 11 and found 3.9.8 still affected on publication day. Cursor closed both reports as duplicates of an April report; we found no vendor advisory or new CVE. Approval dialogs are a security boundary, and this one regressed.

Class 3 — plumbing

CVE-2026-59726 "RufRoot": Ruflo, 9 July

In Ruflo's default docker-compose deployment, the MCP bridge endpoints were reachable without authentication. The CVE record lists what followed: command execution in the bridge container, reading of the provider API keys, and poisoning of the agent's learning store. GitHub, acting as CNA, scored it 10.0; the fix is in 3.16.3. The name comes from Noma Labs. No model is involved at any step: this is an unauthenticated service that happens to be able to act.

CVE-2026-59950: MCP Python SDK, 15 July

The deprecated WebSocket server transport in the official MCP Python SDK accepted handshakes without validating the Host or Origin headers, the classic opening for cross-site WebSocket abuse. The patched version is 1.28.1. One detail matters for anyone who upgrades: according to the GitHub advisory, the new validation setting defaults to off. Upgrading alone does not turn the protection on; it has to be configured, or the server moved to the Streamable HTTP transport.

CVE-2026-15015: WordPress MCP connector, 23 July

In the MountDev AI MCP Connector for WordPress, up to and including 1.6.1, an unauthenticated attacker could obtain an administrator-bound OAuth bearer token by registering their own client and using an authorization endpoint that did not check who was asking. Wordfence, the CNA, scored it 9.8. The vendor was notified on 8 July and the flaw disclosed on 22 July. Five days later the MCP 2026-07-28 revision deprecated Dynamic Client Registration, the mechanism at the start of this chain, in favour of Client ID Metadata Documents.

Class 4 — authority

GitLost: GitHub Agentic Workflows, 6 July

Noma Labs planted text in an issue on a public repository. An agentic workflow that could read the organisation's repositories and write public comments followed it, and published private repository content in a public comment. The research was disclosed to GitHub. Here the model was manipulated, but the outcome was set by a configuration that combined two powers that should never meet in one job: reading private data and writing in public.

Azure DevOps MCP server, 21 July

Manifold Security hid an instruction in an invisible comment inside a pull-request description. Microsoft's official Azure DevOps MCP server had a content-isolation defence, spotlighting, applied to some tools such as pipelines and wiki, but not to the tool that returns pull-request descriptions. The researchers validated the result with two different agents, Copilot CLI and Claude Code, and MSRC acknowledged and triaged the report. The defence existed; it was simply not applied to every door.

Amazon Kiro, 27 August

Mindgard showed that attacker-controlled workspace content, delivered through workspace settings and Kiro's Powers feature, could steer the agent into exfiltrating sensitive local data, and that it worked in both trusted and untrusted workspaces. Testing was on Kiro IDE 0.7.45 on Windows. Amazon fixed it in 0.8.140 in January 2026; no CVE had been assigned when the research was published in August, seven months later. Mindgard framed the post around what that delay says about disclosure for AI tools.

[04] Where the lines are hard to draw

Any classification like this one involves judgment calls, so here is the rule we applied and the places where it was hard to apply.

The rule. If the attack only works because the model follows instructions it should not have trusted, the case is class 4, whatever the delivery channel. Otherwise, the case goes to the earliest step that fails without any model decision: before the trust decision (1), in what the user is shown or asked to approve (2), or in the authentication and authorization of the services around the agent (3).

Kiro. The payload travels through workspace settings, which argues for class 1. But the exfiltration only happens because the agent acts on injected instructions, so by our rule it is class 4. Moved to class 1, the headline would read eight of ten.

PromptFiction. The model does act on text it should not have received. We kept it in class 2 because the text arrives as if the user had typed it: the model is not fooled into obeying a document, it obeys what looks like its user, and the fix was a confirmation step, not a model change. Moved to class 4, the headline would read six of ten.

Counting. GitSpawn counts once, although it covers eight findings. Counting findings instead of disclosures would put nine entries in class 1 and make our point look stronger than a fair count allows, so we did not.

Whichever way the two borderline cases go, the result holds between six and eight out of ten: in most of this summer's disclosures, the model's judgment was not what decided the outcome. For the preprint in November we will repeat the classification with criteria written in advance and a second, independent classifier, and report how often the two agree.

[05] What the pattern says

Class 1 is the loudest signal. GitSpawn is one bug class found independently in seven agents: background context gathering runs git without neutralising the repository's own configuration. It is the third disclosure in five months on the same surface — after GitHub Copilot CLI in May (CVE-2026-45033) and Claude Code in June (CVE-2026-55607). The response pattern matters as much as the bug: several reports were closed as duplicates, one as “informative”, one was never triaged after six contact attempts. A class that many teams hit at once is a design gap, not an implementation slip.

Class 2 moves the attack into the interface. In both cases the user is asked to approve something, and what they see is not what runs. Human-in-the-loop only works if the loop shows the truth.

Class 3 is web security, at agent scale. Missing authentication, missing origin checks, an over-permissive OAuth flow: none of this is new. What is new is that the endpoint behind it can execute commands or write into an agent's memory. The MCP 2026-07-28 revision hardens authorization and deprecates dynamic client registration, but it states plainly that the protocol cannot enforce consent or access control by itself — that stays with implementers.

Class 4 is where prompt injection actually lives — and even there, the deciding variable was authority. GitLost needed a workflow that could both read private repositories and write public comments. The Azure DevOps server protected some tools and not the one that mattered.

[06] Meanwhile, the model layer is not getting safer

The research of the same weeks shows why “the model will refuse” is not a boundary:

Put together with the table: the model layer is porous, and the bugs being found are around it. Both push the same way — measure the boundaries.

[07] What to check on Monday

None of this needs a lab to act on. If you use or run AI agents, these are the checks the ten cases point to. They are defensive measures; none of them requires knowing how the attacks work in detail.

If you use coding agents

  1. Treat a repository as executable input. Open code you did not write, such as archives, forks, take-home tests or bug-report attachments, in a disposable VM or container before any agent touches it. After GitSpawn, a repository with its .git directory is not just data.
  2. Check your versions against the fixes. Goose 1.44.0 or later, Claude Code 2.1.196 or later for the startup path (the second path was still open on 2.1.252 when GitSpawn was published, so check Anthropic's release notes), Claude Desktop 1.1.2321 or later. For Hermes Agent, Qwen Code and Grok Build, check the vendor's notes: no fix was public when GitSpawn was published.
  3. Read the whole thing before you approve. Expand the full command before approving an MCP server install, and do not approve anything that arrives from a link you did not expect. An approval dialog is only a control if it shows you everything.
  4. Review agent configuration like CI configuration. Files that declare MCP servers or agent settings inside a project, for example .mcp.json, .vscode/mcp.json, .cursor/, .gemini/ or .claude/, decide what runs on a developer's machine. They deserve the same code review as a build pipeline.

If you run MCP servers

  1. No MCP endpoint without authentication, including the ones that are "only" reachable inside a docker-compose network. If you run Ruflo, update to 3.16.3 or later.
  2. Upgrading is not always enough. On the MCP Python SDK's WebSocket transport, 1.28.1 adds Host and Origin validation but leaves it off by default: configure it explicitly, or move to Streamable HTTP.
  3. Plan the move away from Dynamic Client Registration. The 2026-07-28 revision deprecates it in favour of Client ID Metadata Documents, with at least twelve months before removal. If you run the MountDev WordPress connector, move past 1.6.1.
  4. Apply your content defences to every tool. If your server has a helper that isolates third-party content, check that every tool returning third-party content goes through it, and add a test that fails when a new tool does not.

If you give agents authority

  1. Never combine private read and public write in one agent job. GitLost needed both. Split them, and give each workflow its own least-privilege token.
  2. Log what agents start, not only what they say. Several of these cases leave no trace in the conversation, because nothing in them went through the model. Process and network logs on developer machines are where they show up. We come back to this on 27 October.

[08] What comes next

Over the next four Tuesdays we measure those boundaries on current versions, in a lab made only of our own repositories, servers and virtual machines, with harmless canaries:

If a test finds something new in a product, the post reports its class and disclosure status only; details follow the fix or 90 days.

Sources

  1. Manifold Security, “GitSpawn”, 1 Sep 2026 — manifold.security
  2. Manifold Security, Azure DevOps MCP confused deputy, 21 Jul 2026 — manifold.security
  3. Oasis Security, PromptFiction — oasis.security
  4. Adversa AI, DeepJack — adversa.ai
  5. Noma Labs, GitLost — noma.security
  6. Mindgard, Amazon Kiro — mindgard.ai
  7. Noma Labs, RufRoot (CVE-2026-59726) — noma.security
  8. GitHub Security Advisories: GHSA-7835-87q9-rgvv (Claude Code, CVE-2026-55607), GHSA-c4hm-4h84-2cf3 (Ruflo, CVE-2026-59726), GHSA-vj7q-gjh5-988w (MCP Python SDK, CVE-2026-59950)
  9. Wordfence, MountDev AI MCP Connector (CVE-2026-15015) — wordfence.com
  10. CVE.org: CVE-2026-55607, CVE-2026-71963, CVE-2026-72718, CVE-2026-45033, CVE-2026-59726, CVE-2026-59950, CVE-2026-15015
  11. MCP specification 2026-07-28 and changelog — specification · changelog
  12. arXiv 2609.18217, 2609.22510, 2609.26900