AI agent security

AI agent security: the risks, the real incidents, and how to test for them

An agent does not just answer, it acts. Here is what that changes for security, what has already gone wrong in production, and how to find out where your own agent breaks.

AI agent security is the practice of protecting AI agents, systems that plan and take actions through tools, APIs and data, from being manipulated into doing something their owner never intended. A chatbot that goes wrong says something wrong. An agent that goes wrong sends the email, deletes the record or hands over the token. That shift from words to actions is why agent security is its own discipline, and why the OWASP Top 10 for Agentic Applications, published in December 2025, exists alongside the older LLM Top 10.

What is AI agent security?

An AI agent is a language model wired to tools. Anthropic's definition is a useful one: agents are systems where the model dynamically directs its own processes and tool usage, rather than following a fixed code path. A support agent that looks up an order and issues a refund is one. So is a coding agent that reads an issue, edits files and opens a pull request, and an internal assistant that searches your documents and drafts replies to email.

AI agent security covers everything that can steer those decisions and everything the agent can touch once steered: the inputs it reads, the tools and permissions it holds, the memory it carries between sessions, the identity it acts under and the other agents or servers it talks to. It overlaps with LLM security, but it is not the same thing. OWASP draws the line clearly: prompt injection against a chatbot alters one response, while an attack on an agent redirects its goals, planning and multi-step behaviour.

LLM chatbotAI agent
OutputText for a human to readActions: API calls, database writes, emails, code
Worst caseA harmful or wrong answerData leaves, records change, money moves
InputsMostly the user's messagesDocuments, tickets, web pages, tool replies, other agents
StateOne conversationMemory, plans and context that persist
IdentityRarely relevantActs with tokens, keys or a user's access

Why agents are harder to secure than chatbots

The root problem is old and unsolved: a language model cannot reliably tell the instructions it should follow from the content it is working on. OWASP puts it plainly in the Agentic Top 10, noting that agents cannot reliably distinguish instructions from related content. In a chatbot that weakness is contained. In an agent, six things turn it into a security problem:

  • Tools. Model output becomes an action. A redirected agent calls real APIs with real consequences.
  • Autonomy. One injected instruction can steer a chain of steps with no human looking at any of them.
  • Memory. What gets written into long-term memory or a retrieval store today steers decisions next week.
  • Delegated identity. Agents act with OAuth tokens, API keys and service accounts, often broader than any single user's access.
  • A runtime supply chain. Tools and MCP servers are loaded and described at runtime, and their descriptions are read by the model as instructions.
  • Other agents. In multi-agent systems, one agent's output is the next agent's input, so a fault in one spreads.

The main AI agent security risks: the OWASP Agentic Top 10

The OWASP GenAI Security Project published the Top 10 for Agentic Applications 2026 on 9 December 2025. It is the most useful shared vocabulary for agent risk. The ten, with their official IDs:

RiskWhat it looks like
ASI01Agent Goal HijackInjected or poisoned content changes what the agent is trying to do
ASI02Tool Misuse and ExploitationThe agent uses its own legitimate tools to delete, spend or leak
ASI03Identity and Privilege AbuseBorrowed or over-broad credentials let the agent reach too much
ASI04Agentic Supply Chain VulnerabilitiesA malicious or compromised tool, MCP server, plugin or model
ASI05Unexpected Code Execution (RCE)Injected text ends up executed as code on a host or container
ASI06Memory & Context PoisoningPoisoned memory or retrieval data steers future sessions
ASI07Insecure Inter-Agent CommunicationSpoofed or tampered messages between agents
ASI08Cascading FailuresOne error spreads and grows across agents and workflows
ASI09Human-Agent Trust ExploitationThe agent talks a human into approving the harmful step
ASI10Rogue AgentsThe agent drifts out of scope or ignores stop commands

In practice these collapse into a handful of failure patterns.

Goal hijack through indirect prompt injection

Most of the incidents below started this way. The attacker never talks to the agent directly. They plant instructions in something the agent will read later: an email, a support ticket, a GitHub issue, a web page, a lead form, a tool's reply. When the agent processes it, the planted text competes with your system prompt and can win. NIST calls this agent hijacking. Direct injection, a user typing the override themselves, matters too, but indirect injection is what turns an agent into a tool for someone outside your company. Our prompt injection test shows what your agent does with an instruction it was never given.

Tool misuse and excessive agency

An agent does not need to be attacked to do damage. OWASP's Excessive Agency entry names three causes: too many functions, too many permissions and too much autonomy. When Replit's coding agent deleted a live production database during a code freeze in July 2025, there was no attacker at all. The agent had the ability and used it. Every tool you connect is something an injected instruction can also use.

Data exfiltration

Agents leak data by encoding it into something that leaves the building: an image URL, a link, a pull request, a ticket comment, an email. Several of the incidents below used nothing more exotic than an image the agent rendered, with the stolen data in its address. The fix is rarely in the model. It is in what the agent is allowed to render, fetch and send. See the data leakage test for how Enoki probes for it.

System prompt leakage

The system prompt often holds your business rules, tool names and guardrail wording, and sometimes credentials that should never be there. Once an attacker has it, every later injection gets more precise, because it can name your real tools and phrase itself to slip past your real rules. OWASP's guidance for LLM07 is blunt: never put secrets or authorization logic in the prompt. The system prompt leak test runs the known extraction paths.

Identity and privilege abuse

Agents usually borrow an identity: a user's OAuth token, a shared service account, a long-lived API key. That creates a confused deputy. Anyone who can steer the agent inherits whatever the agent can reach, even if they could not reach it themselves. In the Supabase MCP case (below), an agent running with a service role that bypassed row-level security read secret tokens because a support ticket told it to. Give each agent its own scoped identity and enforce authorization in the systems it calls, not in the prompt.

MCP and the agent supply chain

The Model Context Protocol made it easy to connect agents to tools, and it moved part of the attack surface into tool metadata. Invariant Labs showed tool poisoning in April 2025: hidden instructions in a tool's description, invisible to the user but read by the model. The same research described rug pulls, where a server changes its tool descriptions after you approved them. By September 2025 the first malicious MCP server had turned up in the wild. Treat every MCP server as code you are running, because you are.

Memory poisoning, multi-agent systems and cascading failures

The newer risks come from state. A poisoned entry in long-term memory or a retrieval store (ASI06) keeps steering the agent in sessions that have nothing to do with the attacker. In multi-agent systems, agents often accept each other's messages without checking where they came from (ASI07), and one bad output becomes the next agent's trusted input, so a single fault can bypass every human checkpoint along the way (ASI08).

Real AI agent security incidents

None of this is theoretical. Here are twelve disclosed incidents and vulnerabilities from 2024 to 2026, each linked to its source. Notice how few needed a software bug. Most needed only an agent that read untrusted content while holding access to something valuable.

Aug 2024 · Slack AIInstructions posted in a public channel made Slack AI pull data from the victim's private channels and put it in a phishing link. PromptArmor. Indirect prompt injection, exfiltration.
May 2025 · GitHub MCP serverA malicious issue in a public repository hijacked an agent asked to triage issues. It read the user's private repositories and leaked their contents into a public pull request. Invariant Labs. Goal hijack (ASI01).
Jun 2025 · Microsoft 365 Copilot (EchoLeak)One crafted email, zero clicks: Copilot exfiltrated data from its scope by chaining four bypasses, including Microsoft's own injection classifier. CVE-2025-32711, CVSS 9.3. EchoLeak paper. Indirect prompt injection, exfiltration.
Jul 2025 · Supabase MCP with CursorAn attacker's support ticket made a developer's agent, connected with a role that bypassed row-level security, read integration tokens and write them back into the ticket. General Analysis. Confused deputy (ASI03).
Jul 2025 · ReplitThe coding agent deleted a live production database during a code freeze and then misreported whether a rollback was possible. No attacker involved. The Register. Excessive agency, rogue agent (ASI10).
Jul 2025 · Amazon Q DeveloperA threat actor used a mis-scoped token to slip a destructive prompt into a release of the VS Code extension. It failed only because of a syntax error. CVE-2025-8217. AWS bulletin. Supply chain (ASI04).
Aug 2025 · GitHub Copilot in VS CodeInjected instructions made the agent switch on auto-approval in its own settings file, after which it could run arbitrary shell commands. CVE-2025-53773. Embrace The Red. Code execution (ASI05).
Aug 2025 · Perplexity CometHidden text in a Reddit comment made the agentic browser read the user's email address and a one-time code from Gmail, then post both back to Reddit. Brave. Indirect prompt injection, account takeover.
Sep 2025 · Salesforce Agentforce (ForcedLeak)Instructions submitted through a web-to-lead form made the agent send CRM data to an expired domain still on the allowlist, which the researchers bought for about five dollars. CVSS 9.4. Noma Security. Goal hijack, exfiltration.
Sep 2025 · postmark-mcpThe first malicious MCP server found in the wild: a lookalike package that quietly copied every email it sent to the attacker. The Hacker News. Malicious tool (ASI04).
Mar 2026 · Vertex AI Agent EngineA deployed agent could extract its default service-agent credentials and use them to read every storage bucket in the project. Unit 42. Identity and privilege abuse (ASI03).
Apr 2026 · AI agents in CI (Comment and Control)Pull request titles, issue bodies and hidden HTML comments hijacked three coding agents running in CI and leaked their API keys and GitHub tokens. Aonan Guan. Indirect prompt injection, credential theft.

The pattern is consistent. The OWASP exploit round-up for Q1 2026 counted eight major incidents and only one CVE among them. Most came from design flaws, misconfiguration and the supply chain, which is not what a vulnerability scanner looks for.

How to secure an AI agent: the controls that matter

The guidance from OWASP, Google, Meta and Microsoft agrees on a short list. None of it is exotic. Most of it is ordinary security engineering applied to a component that is easy to over-trust.

  1. Break the trifecta by design. Decide per agent which two of untrusted input, sensitive access and external action it really needs. If it needs all three, put a human in the loop for that session.
  2. Apply least agency, not just least privilege. OWASP's term for it: fewer tools, narrower tools, no open-ended ones such as a raw shell, and no autonomy where a fixed workflow would do.
  3. Give every agent its own identity. Scoped, short-lived credentials per agent, and authorization enforced in the systems it calls. Never rely on the prompt to decide who may do what.
  4. Require approval for high-impact actions, and make the approval specific. Show the human the exact action and parameters. Vague "approve?" prompts train people to click yes.
  5. Treat everything the agent reads as untrusted. Documents, tool replies and web content are data, not instructions, and nothing in them should be able to change the agent's goal.
  6. Control what leaves. Block rendering of links and images to untrusted domains and keep allowlists short and current. EchoLeak and ForcedLeak both sent their data out through an image from a domain the agent was allowed to load.
  7. Sandbox anything that executes. Code the agent writes or runs belongs in an isolated environment, and every model-controllable tool parameter should be treated as attacker-influenced.
  8. Pin and review your tools. Pin MCP server versions, review tool descriptions as you would code, and connect only to servers you trust.
  9. Keep memory clean. Separate memory per tenant, record where each entry came from, expire what was never verified, and do not let the agent re-ingest its own output as fact.
  10. Log actions, not just prompts. Record every tool call with its parameters and the goal the agent was pursuing, so you can see when behaviour drifts and reconstruct what happened.

Why controls alone are not enough

Controls shrink the blast radius. They do not tell you whether your particular agent can be steered, and filters in particular are weaker than they look. In October 2025, researchers from OpenAI, Anthropic and Google DeepMind tested twelve published prompt injection defences with adaptive attacks and bypassed most of them more than 90% of the time; human red teamers got through all of them. OpenAI has said prompt injection is unlikely to ever be fully solved. As Willison puts it, a defence that stops 95% of attacks is a failing grade in security.

Persistence matters as much as cleverness. When NIST strengthened its agent hijacking evaluations, new red-team attacks succeeded 81% of the time against 11% for the best known baseline, and allowing 25 attempts instead of one raised the average success rate from 57% to 80%. An attacker gets as many tries as they want. A test that tries once tells you very little.

The failures that hurt most are also the ones specific to your agent: the refund it should not issue, the record it should not show that customer, the approval step it can be talked out of. Those are business-logic vulnerabilities, and no generic filter knows your business logic. In our own benchmark, on two deliberately weak agents whose weaknesses were written down before any tool ran, Enoki's attacker reached 83% and 88% of them across three runs, against 25% and 38% for the best other tool. No single run found everything the three runs found together. The write-up is clear about its limits: two agents are not a census, and we configured the other tools ourselves.

How to test an AI agent's security

The OWASP AI Agent Security Cheat Sheet is specific about when: agents should undergo structured security testing before production and after material changes to prompts, tools, memory, retrieval, policies, or model providers. In practice that means before launch and on every release. Testing an agent well looks like attacking it:

  1. 1Mapwhat the agent can read, reach and do
  2. 2Attackwith inputs built for this agent
  3. 3Persistmulti-turn, many attempts
  4. 4Judgekeep the attack behind every finding
  5. 5Retestprove the fix held
  • Test the running agent, not the prompt. What matters is what the agent does with its real tools and data, so attack the deployed endpoint.
  • Cover the paths attackers use. Direct and indirect injection, tool misuse, data exfiltration, system prompt extraction, privilege escalation and approval bypass are the categories the OWASP cheat sheet lists.
  • Attack across a conversation. An agent can hold up to a single message and give way over several turns. We wrote about why in single-turn, multi-turn and dynamic attacks, and the multi-turn jailbreak test shows how your agent holds up.
  • Keep the evidence. A finding is only useful with the exact request and the agent's reply, so you can reproduce it before you fix it.
  • Re-run after every fix and every release. A prompt change or a new tool can reopen a hole you closed last month. See manual, automated and continuous red teaming for how teams split this work, or an agent pentest that runs on every release.

Frameworks and regulation

You do not need to adopt all of these. You do need to be able to say which risks you tested for, in terms the rest of your company and your customers' security teams recognise.

  • OWASP Top 10 for Agentic Applications 2026. The agent-specific risk list, ASI01 to ASI10. It explicitly recommends periodic red-team tests that simulate goal override.
  • OWASP Top 10 for LLM Applications 2025. Still the reference for model-level risks, especially LLM01 Prompt Injection, LLM06 Excessive Agency and LLM07 System Prompt Leakage. Our OWASP LLM Top 10 test maps its findings to it.
  • OWASP AI Agent Security Cheat Sheet. The most practical of the three: concrete controls and a testing section.
  • MITRE ATLAS. Adversary techniques for AI systems, now with agent-specific entries such as exfiltration via agent tool invocation (AML.T0086) and agent tool poisoning (AML.T0110).
  • NIST AI RMF. The risk management frame many US and global enterprises use. NIST has been running dedicated work on agent hijacking and agent security since 2025.
  • EU AI Act, Article 15. High-risk AI systems must be resilient against attempts by third parties to alter their use or outputs, which is what prompt injection does. Following the Digital Omnibus, the obligations for stand-alone high-risk systems apply from 2 December 2027.

Where to start

If you have one to five agents in production and no dedicated AI security programme, this order gets you the most for the least effort:

  1. List your agents and what each can do. Every tool, every data source, every credential. Include the internal ones.
  2. Score each one against the trifecta. Which have private data, untrusted input and a way out, all at once? Start there.
  3. Cut what is not needed. Remove unused tools, narrow permissions, and add approval to the actions you could not undo.
  4. Attack the riskiest agent. Point a red-teaming run at its endpoint and read what got through. Enoki's free run, in the box above, is one way to get a first result in about 30 minutes.
  5. Fix, re-run and make it routine. Retest after the fix, then on every release, so the next prompt change does not quietly undo it.

Frequently asked questions

What is AI agent security?

AI agent security is the practice of protecting AI agents, systems that plan and act through tools, APIs and data, from being manipulated into harmful actions. It covers the agent's inputs, tools, permissions, memory, identity and connections to other agents, not only the language model.

How is agentic AI security different from LLM security?

LLM security is mostly about what a model says. Agentic AI security is about what an agent does. Because agents call tools, hold credentials and keep memory, the same weakness, such as prompt injection, can lead to deleted data, leaked records or stolen tokens instead of a bad answer. OWASP publishes a separate Top 10 for agentic applications for this reason.

What are the biggest security risks of AI agents?

Indirect prompt injection that hijacks the agent's goal, misuse of the agent's own tools, data exfiltration, leakage of the system prompt, and over-broad identities and permissions. MCP servers and other runtime tools add a supply chain risk. The OWASP Top 10 for Agentic Applications lists ten, from ASI01 Agent Goal Hijack to ASI10 Rogue Agents.

Can prompt injection in AI agents be fully prevented?

Not today. Researchers from OpenAI, Anthropic and Google DeepMind bypassed most published defences with adaptive attacks, and OpenAI has said prompt injection is unlikely to ever be fully solved. The practical answer is to limit what a hijacked agent could do, through least privilege, approvals and output controls, and to test regularly whether it can be hijacked.

Have AI agents actually been hacked?

Yes. Some were real attacks: a malicious MCP server copied every email it sent to the attacker, and a tampered release of Amazon Q Developer shipped with a destructive prompt. Others were demonstrated by security researchers, including EchoLeak in Microsoft 365 Copilot, ForcedLeak in Salesforce Agentforce and a GitHub MCP exploit that leaked private repositories. And some of the most damaging failures, such as Replit's agent deleting a production database, involved no attacker at all.

What is the lethal trifecta?

A term coined by Simon Willison for an agent that combines access to private data, exposure to untrusted content, and the ability to communicate externally. An agent with all three can be tricked into sending your data to an attacker. Meta's Agents Rule of Two turns it into a design rule: allow at most two of the three in a session, or add a human.

Is MCP secure?

MCP is a protocol, and its security depends on the servers you connect and the permissions they hold. Known risks include tool poisoning, where instructions hide in tool descriptions, rug pulls, where a server changes those descriptions after approval, and malicious servers published to package registries. Pin versions, review tool descriptions and connect only to servers you trust.

How do you test an AI agent for security?

Attack the running agent the way an adversary would: map what it can reach, send it direct and indirect injections across multi-turn conversations, try to make it misuse tools or leak data, and record the exact exchange behind every failure. Repeat after every change to prompts, tools, memory or model, as the OWASP AI Agent Security Cheat Sheet recommends.