AI agent pentest

A pentest that runs on every release, not once a year

An annual pentest describes an agent that has since been re-prompted, re-tooled and redeployed. Enoki runs the same attacks on every change, and re-runs a finding's own attacks to confirm the fix held.

FreeNo installNo code access

Point-in-time is the wrong shape for an agent

4 categories, run identically on every release, so each run can be compared with the last.

System Prompt Override

high

The agent ignores its system instructions and follows the attacker's instructions instead.

  • LLM01:2025
  • ASI01:2026
  • AML.T0051.000

PII Leakage

critical

The agent leaks personal data that was seeded into its context.

  • LLM02:2025
  • ASI06:2026
  • AML.T0057

Excessive Agency

high

The agent takes actions beyond what was requested, without stopping to ask permission.

  • LLM06:2025
  • ASI02:2026
  • AML.T0053

System Prompt Leakage

critical

The agent reveals its system prompt or its hidden instructions.

  • LLM07:2025
  • ASI06:2026
  • AML.T0056
  • AML.T0069.002

Free run

  • One agent, one short security assessment
  • Every finding with the attack that proved it
  • It tests the agent, not the app around it

Enoki platform

  • Assessments that run far deeper
  • The full conversation behind every finding
  • Unlimited reruns
  • REST API for your pipeline
See plans

Read further: Manual, automated and continuous testing compared

What happens after the first run

  1. Run it on every release

    The same suite, so each run can be compared with the last.

  2. Confirm the fix

    Enoki re-runs a finding's own attacks, so a fix is verified rather than asserted.

Questions before you run it

Does this replace our annual pentest?

Not the whole of it. It replaces the part that goes stale fastest, which is the agent's own behaviour under attack. The network, the auth and the application around it still want a human.

How is a fix confirmed?

By re-running the attacks that produced the finding, rather than by re-running the whole suite and hoping. A fix that only moves the failure somewhere else does not pass.

What does it cost after the free run?

The free run costs nothing and needs no code access. Everything beyond it is on the pricing page.

Your first assessment

Enter your agent's endpoint

One field, no install, no code access. You get the report and the attack behind every finding.

FreeNo installNo code access