Prompt injection becomes code execution when an agent can call a dangerous tool and the code behind that tool trusts the arguments the model fills in. Two vulnerabilities that Microsoft researchers found in its own Semantic Kernel agent framework show exactly how. In one, a search filter built from a model-supplied value ended up in Python’s eval(). In the other, a file-download helper exposed to the model let it choose where to write on the host machine. In both cases the model did what it was designed to do. The boundary failed in the software around it. Microsoft’s conclusion is the one to remember: “your LLM is not a security boundary. The tools you expose define your attacker’s affected scope.”
From bad answers to bad actions
In our explainer on prompt injection, the risk scaled with what an agent can do: a chatbot gives a bad answer, while an agent with tools can take a bad action. The Semantic Kernel cases show the step after that, where a tool call turns into code running on the machine that hosts the agent.
Microsoft published the research on May 7, 2026, in a post titled “When prompts become shells”. Both flaws had already been fixed and disclosed in February through GitHub security advisories, each rated 9.9 out of 10 (critical) on the CVSS scale.
Case 1: a search filter that became code (CVE-2026-26030)
Microsoft’s researchers built a demo “hotel finder” agent with Semantic Kernel’s Python SDK. The model could call a search function with a parameter such as city="Paris". Behind the scenes, the framework’s default filter for its In-Memory Vector Store turned that value into a small piece of Python and ran it with eval().
The chain, at a high level:
- Untrusted input reaches the model. The attacker needs a way to get text into the agent’s context, which is ordinary prompt injection.
- The model picks a legitimate tool. Searching hotels is exactly what the agent is for.
- The model fills in a parameter the attacker shaped. Instead of a city name, the value carries code.
- The framework splices that value into code and runs it. Model-controlled text became part of an expression passed to
eval(). - A safety check that could be bypassed. The developers had anticipated the risk and added a validator that blocked dangerous names. Microsoft found a way around that blocklist. As it put it, “blocklists in dynamic languages like Python are inherently fragile.”
The result was code execution on the host. In Microsoft’s demo, a prompt launched the Windows Calculator. The fix added four layers, led by allowlists: only safe kinds of expressions, only approved function calls, only the filter’s own variable as a name, plus a block on the attributes used to climb Python’s class hierarchy.
Affected: the Python semantic-kernel package before 1.39.4, when the agent’s search relies on the In-Memory Vector Store’s filter in its default configuration. Fixed in 1.39.4.
Case 2: a helper that should never have been a tool (CVE-2026-25592)
The second flaw is more instructive, because the sandbox existed and worked. Semantic Kernel’s SessionsPythonPlugin runs model-written Python inside Azure Container Apps dynamic sessions, an isolated cloud sandbox. Helper functions move files between that sandbox and the host.
In the .NET SDK, one of those helpers, DownloadFileAsync, was accidentally marked as a function the model could call. Its localFilePath parameter decides where on the host the file is written, and it had no path validation. So the chain became:
- Injected instructions get the agent to create a file inside the sandbox. So far, it’s contained.
- Further instructions get the agent to call the download helper with a host path of the attacker’s choosing.
- The file lands in a location where Windows runs it at the next sign-in.
Microsoft’s summary: exposing that function “provided a direct file write primitive on the host filesystem, effectively negating the container isolation.” A related flaw ran the other way. The upload helper, in both the Python and .NET SDKs, accepted any local path, so an injection could copy sensitive host files, such as SSH keys, into the sandbox.
The fix removed the function from the model’s reach entirely and added path validation (canonicalizing the path, then checking it against allowed directories) for developers who call it from their own code.
Affected: the Semantic Kernel .NET SDK before 1.71.0 (the GitHub advisory names the Microsoft.SemanticKernel.Plugins.Core package). It also lists the Python package before 1.39.3 for the upload side. Fixed in .NET 1.71.0 and Python 1.39.3. At the time of writing, the main .NET package, Microsoft.SemanticKernel.Core, is at 1.81.0 on NuGet, and the Python package is at 1.45.0 on PyPI. If you build on Semantic Kernel, check your version against the project’s security advisories.
Model failure vs. agent architecture failure
It’s tempting to file these under “the AI got tricked.” Microsoft argues the opposite: “The AI model itself isn’t the issue as it’s behaving exactly as designed by parsing language into tool schemas. The vulnerability lies in how the framework and tools trust the parsed data.”
| Model failure | Agent architecture failure | |
|---|---|---|
| What goes wrong | The model follows injected instructions or misreads intent | Code acts on the model’s output without enforcing limits |
| Typical result on its own | A wrong answer or an unwanted tool call | Whatever the tool can reach: files, commands, data |
| Can you fully prevent it? | No. Models are probabilistic, and filters miss attacks | Largely, with ordinary engineering: validation, allowlists, isolation |
| In these cases | The model was induced to pass hostile parameters | eval() on model text; a host-writing helper exposed as a tool |
You can’t count on the left column never failing. The right column decides what happens when it does. That’s an old lesson: Microsoft compares it to the early days of web security, when untrusted input went straight into SQL queries and filesystem calls.
What actually limits the blast radius
- Treat every model-filled argument as attacker input. Validate it the way you’d validate a form field from the internet: type, length, format, allowed values. Prefer enumerations and IDs to free text.
- Never pass model output to an interpreter. No
eval(), shell strings or dynamic queries built from model text. If you must evaluate something, use allowlists of what’s permitted, not blocklists of what isn’t. - Expose tools deliberately. Keep an explicit allowlist of what the model can call, and review it when frameworks update. Case 2 was one attribute on a helper that should have been internal.
- Constrain paths and resources. Canonicalize file paths, then check them against allowed directories. Do the same for URLs and hosts.
- Least privilege for the agent process. Run it with an account and credentials that can reach only what the task needs.
- Sandbox, and guard the bridges. Isolation is only as strong as the functions that cross it. Every upload, download or callback between sandbox and host is part of the boundary.
- Require human approval for consequential actions. Writing outside a workspace, running commands, sending data out, deleting things.
- Watch the host, not just the prompt. Microsoft recommends correlating model-level signals with endpoint telemetry, like an agent process suddenly spawning command lines or dropping scripts in startup folders.
OWASP’s Top 10 for Agentic Applications, published in December 2025, has separate entries for both halves of this story: “Tool Misuse” and “Unexpected Code Execution.”
Why it matters beyond Semantic Kernel
Microsoft says it found “structurally similar execution vulnerabilities” in other widely used agent frameworks, and framed Semantic Kernel as the first case in a series. The pattern applies to any agent with tools:
- Coding agents run commands, edit files and install packages by design, so the parameters they pass to a shell are the attack surface. Sandboxes, approved command lists and reviewing diffs before running them matter more than the model’s judgment.
- MCP-connected agents call tools published by servers that someone else may have written. Each tool’s parameter handling is part of your security, as we noted in MCP vs APIs vs computer use.
- File and email agents read the untrusted content an injection needs and often hold tools that can send or save data. File paths and recipients are model-controlled parameters.
- Browser and computer-use agents read the whole web, and their “tools” are clicks and keystrokes in a live session. Isolated environments and confirmation before purchases or data entry are the equivalent controls.
The practical test for any agent you build or deploy is simple. For each tool, assume the model will one day call it with the worst arguments an attacker can write. If the result is unacceptable, the fix belongs in the tool, not the prompt.
Advisories, package versions and Microsoft’s research checked on October 9, 2026.
