The Model Can Propose a Finding. Only Evidence Can Confirm One.

by Cody Chamberlain

Over the past year I've built two autonomous pentesting tools under TOMfoolery Labs:

  • SCARAB is an agentic web application pentesting platform. Each assessment runs in its own ECS Fargate container with a standard toolkit (ffuf, katana, nikto, testssl.sh, Playwright). Six specialist agents work in sequence: recon, auth, config, injection, authorization, and business logic. They share state, and findings stream to a React front end as they happen.
  • IPA (Intelligent Pentest Agent) is a Dockerized LLM agent for network and infrastructure testing. It drives OpenVAS/Greenbone alongside faster tools like nuclei and nikto, and I've tested it against Metasploitable3 on Proxmox.

Both use a frontier model (mostly Claude, through the API) as the reasoning core, with retrieval-augmented generation (RAG) supplying offensive security context and target-specific data.

I spent a long time in offensive security before building these, including running product at NetSPI. So I went in with opinions about what a pentest actually is. Some of those opinions held up. Others didn't. This post covers both.


The short version

LLM agents are very good at the connective tissue of a pentest: reading tool output, deciding what to try next, chaining small observations together, and writing it all up. They're unreliable at the one thing that matters most, which is telling you, truthfully, whether something is actually vulnerable.

So the whole design problem reduces to one rule:

The model can propose a finding. Only evidence can confirm one.

Everything else in this post is a consequence of that rule.


Where agents genuinely help

1. Reading tool output like a junior tester would

Recon tools produce a lot of noisy text: nmap service banners, ffuf hit lists, katana crawl output, nikto's wall of warnings. A good model reads all of it, pulls out what's interesting, and explains why it's interesting. That's the tedious first hour of every engagement, and it's where agents earn their keep fastest.

2. Deciding what to do next

Traditional scanners run a fixed playbook. An agent can look at what it's found and adapt. If the auth agent discovers a second role, the authorization agent should test horizontal and vertical access between them. If recon finds an exposed admin path, that path should get attention first.

This is the part that feels most like a human tester. It's also why SCARAB's agents run in sequence and share state: each one builds on what the previous one learned instead of starting cold.

3. Orchestrating slow and fast tools together

On the infrastructure side, OpenVAS is thorough but slow. Blocking an agent on a full scan wastes most of an engagement. IPA uses a kick/harvest pattern: start the OpenVAS scan, let the agent work with fast tools like nuclei and nikto in the meantime, then harvest the OpenVAS results when they're ready and fold them into the picture.

An agent handles that kind of juggling well. It's tedious to hard-code and natural to describe to a model.

4. Writing it up

Turning a pile of findings, requests, and responses into a readable report with severity, affected endpoints, reproduction steps, and remediation guidance is something models are simply good at. This was never the hard part of pentesting, but it was always the most time-consuming part.


Where agents hallucinate findings

This is the part every vendor demo skips. Some of these failures I hit myself. Others are well-known patterns that anyone building in this space should design for from the start. The most convincing one I hit wasn't strictly a hallucination at all.

The "broken authorization" finding that was really a session bug

Authorization testing needs at least two identities. SCARAB tests apps with a low-privilege user account and an admin account, and the authorization agent's job is to check whether the user can reach things only the admin should.

In one run, it reported broken access control: the low-privilege user could reach admin-only functionality. That's a high-severity finding, and the evidence looked solid. There were requests made "as the user" and successful responses from admin endpoints.

The bug was in my harness, not the target. The testing wasn't switching sessions correctly. Every request, including the ones labeled as the low-privilege user, was going out with the admin session. Of course "both roles" had admin access. Only one role was ever being tested.

That's what makes this failure mode dangerous. The model's reasoning was perfectly sound given what it saw: request labeled "user," admin data came back, so the user has admin access. The mistake was upstream, in state the model had no way to check. Everything downstream, including the confident write-up, inherited it.

What caught it was adding a judge: a finding manager that validates every finding against its evidence before it's accepted. When the judge looked at this one, the problem was obvious. The "user" requests and the "admin" requests carried the same session ID. One identity can't prove a privilege boundary is broken, so the finding was rejected and the real bug surfaced.

What I took from it:

  • Session identity is state that has to be verified, not assumed. Before any authorization test, it's worth confirming each session really is who you think it is, for example by checking that a profile endpoint returns different roles.
  • Record which identity sent each request. The evidence for an authorization finding has to show the actual session or token used, not just a label saying which role it was supposed to be. That's exactly what let the judge catch this.
  • Don't let the agent that found something be the one that confirms it. The testing agent was convinced. A separate check looking only at the evidence wasn't.

"The version is old, so it's vulnerable"

An agent sees a service banner, recalls a CVE for that version, and reports it as a finding without ever testing it. Sometimes the CVE doesn't apply because of the configuration. Sometimes the build was backported and patched. Occasionally the CVE number itself is wrong. This is one of the most common ways LLM-driven scanners inflate their findings: a plausible-sounding claim based on recall instead of a test.

Reading a response as a success

Another classic: an agent sends an injection payload, gets back a 500 error or a slightly different page, and decides that's proof of SQL injection. Or it sees its XSS payload reflected in the response and calls it XSS, without checking whether the payload ever executes in a browser. An unusual response is a lead, not a finding. It's exactly the kind of claim a judge should reject unless the evidence shows actual impact.

Treating "no output" or "partial output" as "nothing there"

The opposite failure: a missed finding presented with confidence. A tool times out, or returns truncated results, and the agent summarizes what it got as if it were complete.

I hit a concrete version of this with OpenVAS. Its management protocol (GMP) returns 10 rows by default unless you explicitly ask for all of them with filter="rows=-1". Nothing errors. You just get the first 10 results, and the agent faithfully reports on those 10 as if they were the whole scan. The model wasn't hallucinating there. It was reasoning correctly over incomplete data, which looks exactly the same in the final report.

Confusing "timed out" with "failed"

Long-running operations can succeed on the server even when the client gives up waiting. An agent that treats every timeout as a failure retries and creates duplicate targets and tasks. An agent that treats every timeout as a success reports results that don't exist. The fix is boring: after a timeout, go check whether the thing actually happened before deciding.

Ignoring the knowledge you gave it

When I wired a RAG knowledge base into my agents as a search tool, the model barely used it at first. It answered from its own training data instead, because it thought it already knew. Vague instructions like "you have access to a knowledge base" don't change that. What helps is a specific tool description, explicit instructions to search before answering, and examples of the behavior you want.

The lesson generalizes: a model with a tool isn't a model that uses the tool. You have to check.


Architecture choices that keep agents honest

Every finding needs its evidence attached

In SCARAB, a finding isn't a finding unless it includes the full HTTP request and response that demonstrates it. That single requirement does more for accuracy than any prompt. It forces the agent to produce something reproducible, and it gives a human reviewer something concrete to check in seconds instead of re-testing from scratch.

If I were starting over, I'd make this a hard schema constraint from day one: no evidence field, no finding. And as the session bug taught me, the evidence has to include who made the request, not just what was sent.

Put a judge between the agents and the report

Evidence only helps if something actually checks it. In SCARAB, findings don't go straight from a testing agent into the report. They go to a finding manager that acts as a judge. It reviews each finding against its evidence and asks whether the evidence really supports the claim.

The testing agents are optimized to find things, so they lean toward believing their own results. The judge's only job is skepticism. That separation is what caught the session bug: the authorization agent was certain, and the judge noticed both "roles" were using the same session ID.

A judge also gives you one place to encode hard-won rules. Every time a false positive slips through, the check that would have caught it goes into the judge, and it applies to every agent from then on.

Score against ground truth, not vibes

Testing against intentionally vulnerable apps (Juice Shop, DVWA, WebGoat) is a good start, but "it found a lot of stuff" isn't a measurement. I moved to Juice Shop in CTF mode with a CTFd scoreboard. When SCARAB actually exploits a Juice Shop challenge, the app returns a flag, and SCARAB submits that flag to CTFd through its API.

That gives you an objective score. A hallucinated finding can't produce a valid flag. It also makes regressions visible: change a prompt or a model, rerun, and compare scores. Exporting a CTFd backup after setup makes it easy to reset to a clean scoreboard between runs.

Make sure it's reasoning, not remembering

There's a catch with benchmarking on famous vulnerable apps. Juice Shop is one of the most documented targets on the internet. Walkthroughs, challenge solutions, and write-ups for nearly every challenge are all over the web, which means they're almost certainly in the model's training data.

So when an agent solves a Juice Shop challenge, you can't tell whether it reasoned its way there from what it observed, or recalled the answer from a walkthrough it saw in training. A high score might just measure how well the model memorized Juice Shop. That tells you nothing about how it will do against a client's app it's never seen.

The answer is to morph the target: change everything the model could have memorized while keeping the vulnerabilities themselves intact.

  • Rename routes, endpoints, and parameters. If the vulnerable endpoint has a different path and different parameter names, the agent has to find it.
  • Rebrand the app. Change the name, the product catalog, the UI text, and anything else that announces "this is Juice Shop."
  • Replace the seed data. Use different users, emails, products, and default credentials, so recalled values don't work.
  • Change error messages and fingerprints. Alter error strings, headers, and anything else that identifies the stack or the app.
  • Keep the bug classes the same. The SQL injection, the broken access control, and the XSS all stay. Only their surface changes.

Then compare. Run the agent against stock Juice Shop and against the morphed version. If the score holds up, the agent is reasoning. If it drops sharply, a good chunk of that original score was memory.

It's also worth reading the attack chain for tells. An agent that requests an endpoint it never discovered during recon, or that names a challenge or user that only exists in the original app, is remembering, not testing. The judge can check for exactly that: every step in a finding should trace back to something the agent actually observed on the target.

One practical note: Juice Shop's built-in challenge detection is tied to its original code, so deep changes can break the automatic flag scoring. The more you morph, the more of your own scoring checks you'll need.

Separate "reasoning" from "proving"

The model decides what to try. Deterministic code and real tools decide what happened: status codes, timing, flags, browser execution, scanner output. The same goes for test state the model can't see, like which session is active. That belongs in code that checks itself, not in the model's assumptions. The more of the "did it work?" judgment you move out of the model and into code, the fewer hallucinated findings reach the report.

Use different models for different jobs

Not every step needs the most capable model. Recon parsing and output classification can run on something smaller and cheaper. Injection testing and attack-path reasoning, where judgment matters most, get the strongest model. SCARAB supports multiple providers and lets each agent use a different one, which made it easy to experiment.

I also looked at fine-tuning and decided against it as the main strategy. The bottleneck for a pentest agent is reasoning, not memorized knowledge, and the target changes every engagement. RAG handles that better. If I fine-tune anything, it'll be a small open model for high-volume, low-judgment work like classifying scan output.

Cache the prompt, never the answer

A multi-turn assessment re-sends the same large prefix every turn: system prompt, tool definitions, and early recon context. Prompt caching makes re-reading that prefix much cheaper without changing what the model sees.

Semantic caching, where a gateway returns a stored answer for a "similar enough" request, is a different story. In a pentest it's dangerous. "Scan ports on 10.0.0.1" and "enumerate 10.0.0.1" look similar but aren't the same instruction, and a cached response from one engagement must never leak into another. Tool results have to reflect what actually happened, live.

Isolate every run

Each SCARAB assessment gets its own container. That's partly about safety, but it's also about clean evidence. When something goes wrong, you know exactly which run and which agent produced it, and nothing carries over between targets.

On the homelab side, the less glamorous version of this lesson: don't run a local LLM and a vulnerability scanner on the same box during heavy work. Both want all the RAM, and when OpenVAS crashed mid-sync on mine, that was why.


Where I think this is going

Agentic pentesting is real, and it's useful today. It removes much of the drudgery and gives every engagement a consistent baseline. But "autonomous" is doing a lot of work in most marketing for this category.

The tools that win won't be the ones with the cleverest prompts. They'll be the ones that are disciplined about evidence: they separate what the model suspects from what the tools proved, and they measure themselves against ground truth on targets the model hasn't already memorized, instead of by finding counts.

That ties back to a point I made in another post: the moat isn't the model. It's the system you build around it.