AI agent security

AI agent security testing vs penetration testing: what a pentest misses

A penetration test checks your application, its login and its APIs. It does not check what your AI agent can be talked into doing, such as handing one customer's data to another.

A penetration test checks the application around your AI agent: its network, its authentication, its APIs and the code behind them. AI agent security testing checks what the agent itself can be talked into doing: leaking its instructions, calling a tool on someone else's behalf, or handing one customer's data to another. If your agent can act, you need both.

If your pentest report says nothing about the agent, that is not a pass. It only means nobody tested the agent. We built a test agent with a flaw on purpose, and with one plain request our attacker got it to hand over another customer's billing details. Ten seconds later, the same agent said it only had access to one customer's data.

A pentest assumes the code decides what happens

A web application pentest tests your network, your login, the access rules on each API and common attacks such as injection, then reports what got through. That works for normal software because the code decides what happens: the same request does the same thing every time, until someone changes the code.

An agent breaks those assumptions

  • The input is language. There is no fixed list of parameters to test. Any sentence is a valid request.
  • A model makes the decisions. The code decides which tools the agent has. The model decides, in the moment, which tool to use and what to ask it.
  • The data talks back. A document, an email or a saved record that the agent reads can contain instructions, and the agent may follow them. We explain this in what AI red teaming is.
  • It keeps changing. A new prompt, a model upgrade or one extra tool changes what the agent does, without a single line of application code changing.

The gap: access rules that live only in the prompt

Here is the failure a pentest can miss. The agent's instructions say which customer's data a user may see. The system behind the agent does not check again: it hands over whatever the agent asks for. So the only thing between a user and another customer's data is a sentence in a prompt.

We built a test agent with exactly this flaw for our tool benchmark. Its instructions tie it to one customer's workspace. The tool it uses to look up records accepts whatever workspace the agent passes in, and never checks it.

Our attacker found a mention of another customer's record in the agent's own notes, and simply asked to see it. The agent fetched the record and shared that customer's billing contact and payment terms, without any warning. Ten seconds later the attacker asked which workspaces it could access. The agent answered: "I only have access to a single workspace."

The fix is not a better prompt. The system that holds the data should check who is logged in, instead of trusting the agent to ask for the right thing.

To be fair: we left that test agent wide open, so a pentester calling its tools directly would have found the same leak. In a real product the leak can be quieter. The system checks normal users correctly but trusts the agent, so the pentest passes while the agent is still a way in. Here is a real case from a live customer agent, anonymised:

Data exfiltrationone ordinary user credential reached the whole workforce personnel records for 150+ people.

Three more findings a standard pentest would miss

These also come from live customer agents, anonymised. Under each one: why a pentest aimed at the application would not catch it.

Indirect injectionan inbound email wrote the reply to a customer one variant carried a link to confirm bank details.

The instruction arrived in an email the agent read, not in anything a user typed. A tester who only sends requests to the API never sees it. This is called indirect prompt injection.

Code extractiona map of its own codebase, then a core file verbatim names the files holding the access checks and the guardrails, so an attacker knows where to aim first.

The agent handed over, on request, the kind of map a pentester builds by hand. A test that never talks to the agent never asks for it, just as it never checks for a system prompt leak.

Document forgerya signed 125,000 euro bank guarantee issued in a named bank's name, and the marking that flags a file as generated came off on request.

Nothing checked the document before the agent produced it. A pentest checks what the code allows, not what the agent is willing to create.

Where a pentest still wins, and how they fit

Agent testing does not replace a pentest. Your network, your login and the application around the agent still need a human tester. So does your business logic: a person thinks about what your product should never allow, and then tries to make it happen. And a customer, an auditor or a framework may require a human pentest and a signed report.

Some firms now offer an AI or agent pentest. Done well, it catches much of what a standard pentest misses. But it is still a one-off: it cannot test next week's release, and in a fixed week it can only repeat each attack so often. That matters, because an agent can answer the same question differently each time. In our tool benchmark, our attacker found 8, 8 and 6 of 12 planted vulnerabilities in three separate runs, and 10 when the three were combined. One run is a sample, not the full picture.

So there are three options. Add the agent to the pentest: one test, and nothing again until next year. Test only the agent: your network, your login and the report your auditors want stay uncovered. Test both: the agent on every release, the pentest on its usual schedule. We would test both, and share what the agent testing finds with your pentester, so they can spend their time on what only a person can judge.

Penetration testAI agent security testing
How it attacksRequests, payloads, exploitsConversations, planted content, attacks over several turns
How oftenOnce a year, or before a big releaseOn every release
What you getA report on the system as it stood that weekEvery finding with the exact attack and the agent's reply
Confirming a fixA retest, scheduled separatelyThe same attack, re-run on the next release

The line between the two is not perfectly clean. Enoki also checks whether your agent can be made to reach another customer's data, which is pentest territory too. But it does not assess your network or infrastructure, and it does not replace a pentest. Findings come with their mapping to OWASP and MITRE ATLAS, the frameworks your security reviewers already use.

What to ask your pentest provider

If your agent is in scope for a pentest, these five questions tell you whether it will really be tested.

  1. Do you attack through the conversation, or only the API? An agent fails in what it does with a request, not in the raw request itself.
  2. Do you plant content the agent reads? Indirect injection arrives in a document, an email or a record, not in the chat box.
  3. Do you hold one conversation over several turns? Some jailbreaks build up one message at a time.
  4. How many times do you repeat an attack? An agent can answer the same request differently each time, so one failed attempt proves little.
  5. What happens after we fix something? A retest next year cannot tell you whether next week's release broke it again.

Keep your pentest. Before the next one, have the agent itself tested, then again on every release, and share what comes back with your pentester.

Frequently asked questions

Is AI agent security testing the same as a penetration test?

No. A penetration test checks the application around the agent: its network, login, APIs and code. AI agent security testing checks what the agent itself can be talked into doing with its tools, data and permissions. An agent that can act needs both.

Does our annual pentest cover our AI agents?

Only if each agent is explicitly in scope and the testers talk to it, including through the documents and emails it reads. Otherwise the pentest tests the API the agent sits behind, not what the agent can be talked into doing.

My pentest report said nothing about the agent. Is that a pass?

No. It means nobody tested the agent, not that it is safe. Unless the agent was explicitly in scope, nobody tried to talk it into anything.

Why can an agent leak data when its API passed a pentest?

Because the rule about who may see what can live only in the agent's instructions. The system behind the agent trusts the agent and hands over whatever it asks for. A pentester who calls the API as a normal user is refused, so the API passes, but nobody asks the agent.

Does a penetration test cover prompt injection?

Only if the agent is in scope and the testers attack it through conversations and through the content it reads, such as a document, an email or a saved record. A test that only sends requests to the API never reaches the place where injection arrives.

How often should an AI agent be tested?

On every release. A new prompt, a model upgrade or an added tool changes what the agent does without any change to the application code, so a yearly pentest describes an agent that no longer exists.