Search This Blog

Powered by Blogger.

Blog Archive

Labels

Footer About

Footer About

Labels

Showing posts with label AI Pentesting. Show all posts

How We Find Critical Vulnerabilities with GLM 5.3 and Red Clippy

Over the last few months our red team exercises for BFSI customers have been run with an AI coding agent sitting in the loop. The findings that came out of them were the usual serious ones: broken authentication, unauthenticated access to sensitive data, an OTP bypass, SSRF, stored XSS, a login form that let us straight in with the password field left empty, a customer search that handed back the entire database when given a wildcard, and on one engagement a payment gateway secret key shipped inside a JavaScript bundle that every visitor's browser downloads.

None of that is exotic. Testers have been finding these things for twenty years. What changed for us was how the work got done, and more importantly, how it got kept.

Give a coding agent a shell and it turns into a fast, tireless tester. It runs the same tools you do. It will read a four megabyte minified bundle line by line without complaining, which is a thing no human on the team volunteers for. It will enumerate an API surface while you are still reading the scope document.

The trouble starts about forty minutes in. The context window fills up. The session compacts, or it ends and you start a fresh one the next morning, and the engagement goes with it. The new session re-scans hosts it already cleared. It re-tests things it already ruled out. Ask it which parts of the scope have been covered and it cannot tell you, because it does not know. And somewhere in a transcript nobody kept there is a confirmed injection that never made it into the report.

That is the problem Red Clippy(https://github.com/CSPF-Founder/red-clippy) exists to solve.

An engagement overview. All the screenshots here come from the project's demo database, not from a customer engagement.

It is not an AI pentesting framework

Red Clippy has no scanning engine of its own, no autonomous attack logic, and no opinion about what should be tested next. It will not find a vulnerability for you.

What it does is keep the record of an engagement while an agent does the testing and you direct it. Targets, scope, what has already been tested, findings, evidence. That is the whole job.

It is built for testers who already know what they are doing and want to use Claude Code, Codex CLI, or any other MCP-compatible client alongside their normal workflow. You define the target and scope in the panel, or paste the customer's scope list into the chat and let the agent enter it. From there you guide the agent however you like, the same way you would guide a junior on the team, and it writes down what it did as it goes.

That turns out to be useful for four things: knowing what has already been tested, checking the same finding across multiple domains and assets, keeping engagement history for periodic retesting, and not having to rely on the model remembering everything or on a folder of text files pretending to be a database.

The setup

Three pieces, all on one machine. GLM 5.3 from z.ai does the reasoning. Claude Code is the client, providing the shell, the file access and the agent loop. Red Clippy holds the record and connects to Claude Code over MCP.

Because it is a client rather than a model, and z.ai serves an Anthropic-compatible endpoint, you can point one at the other and keep the agent harness you already know. The setup is documented on the project page, so we will not repeat it here.

MCP runs client-side, so Red Clippy does not know which model is behind the agent and the tools behave the same either way. That means the discipline of the engagement is not tied to a model you happen to be using this quarter. If we move off GLM next year, the record, the coverage and the findings all survive the move.

The rules arrive before the first tool call

This is the part most people skip when they wire an agent into a workflow, and it is the one that changed our output the most.

An agent that has to ask for the rules of engagement generally will not bother. So Red Clippy hands over a Red Team Instructions document during the MCP handshake, before the agent makes its first tool call. It is one document, not a system prompt maintained in five places, and the most specific one wins: a per-engagement override if there is one, otherwise the organization default, otherwise the built-in.

The Red Team Instructions document, served to every agent on connect and overridable per organization and per engagement.

Most of it is unglamorous. The line that matters most on BFSI work is the one about taking the minimum access needed to show impact. An agent that proves an unauthenticated data exposure by retrieving three records and stopping has given you a finding. An agent that helpfully retrieves the whole table has given you a very different conversation with the customer.

The rest is tradecraft, and that is where several of our critical findings actually came from: read the main bundle rather than grepping it, trigger errors deliberately and read the whole response, strip the auth header and retry, then change the identifiers and see whose data comes back.

None of that is new methodology. It is what a competent tester does anyway. The difference is that it is in the agent's context on every connect, without anyone remembering to paste it in.

Setting up an engagement

You create the pentest, paste in the scope from the engagement letter, mark the in-scope domains and ranges, and mark the exclusions. You can type them yourself or let the agent enter them from the customer's list. Either way you read them before anything gets touched.

Scope units are assets, each with its own checklist, reachability marking and in or out of scope flag.

After that you drive, and the instructions are duller than people expect. "Do the initial recon first, subdomain enumeration across the in-scope domains." "Now go through the asset list, pick up whatever is still untested, and mark the checks off as you clear them." The agent runs its own tools from its own shell, the way it would anyway, and posts the results back as it works. Raw scanner output goes in with a single call and gets parsed automatically, whether it came from nmap, Burp, Nessus, OpenVAS, masscan, naabu or subfinder. The things you are actually testing become assets. Everything else stays an observation attached to an asset. Findings go in with severity, a CVSS vector, a proof of concept and the evidence that backs it.

There are 82 MCP tools, which is nearly everything the panel itself can do. That matters more than it sounds, because a tool set that only covers half the application forces you back into the browser mid-session to finish what the agent started. Anything the agent writes you can write yourself, and anything you write it can read. You can run the engagement entirely by hand, entirely through the agent, or switch between the two in the middle of a session.

What a record does that a transcript cannot

The interesting part is not that the agent is fast, although it is. It is that the work survives the session it was done in.

Recon noise stays out of the scope list, which is why it survives

Every content discovery run produces hundreds of paths. Every bundle you read produces endpoints, internal hostnames, technology fingerprints and, now and then, a secret. Throw all of that into an asset list and the asset list is useless by lunchtime.

Red Clippy separates the two. The things you are testing are assets. Everything else is an observation hanging off the asset it came from, with a kind and the tool that found it.


A few observations are findings in their own right, like a key that should never have been public. Most are leads, and the leads are what pay off later. On our engagements, unauthenticated access to sensitive data came from an API path pulled out of a bundle, called with no auth header, that returned data.

In a transcript that path scrolls away. As an observation it is still there tomorrow, attached to the right host, with the tool that found it recorded alongside.

Correlating things that happened days apart

The findings that matter are rarely one observation. They are usually two, made hours or days apart, that mean something together.

Reading the front-end bundle early in an engagement turns up internal hostnames. They go into the record and testing moves on. Days later, a parameter that fetches a remote image turns out to make outbound requests.

An SSRF is only worth what you can reach with it, and the hostnames from the first day are what you point it at. Making that connection requires the first day's record to still be there, and searchable, when you need it days later. That is exactly what a context window does not give you.

Red Clippy makes those joins explicit. Every finding carries tags for the asset it affects and the check it came from, so it files itself under both. The attack graph lets you link any two things, an observation, an asset or a finding, with a label of your own, then follow the links out from one or trace the route between two. The chain from "hostname found in the bundle" to "reachable through SSRF" to "admin interface behind it" is saved, rather than something to piece back together when the report is written.

Correlating across engagements

A customer is rarely one engagement. There is this quarter's, last quarter's, and the retest after that.

Because those engagements share a record, any host or IP can be asked about across all of them at once. Have we tested this before? What did we find? Was it reachable last time?

That pays off twice over. A critical bug in one API is a question about every other API the customer has: we confirmed one in a session, and a later session testing a different domain found the same bug there, because the first finding was something to check the new asset against rather than a paragraph in a transcript nobody reopened. And a host that was blocked last quarter but answers this one has almost never changed. It is a source address or a VPN, and knowing that saves an hour of chasing a WAF that is not there.

The same goes for paths. Whatever was recorded for a host in an earlier engagement can be pulled into the current one, so you get last round's content discovery for free and can see at once whether what you reported then is still live.

Coverage you can query instead of remember

There are 135 built-in checks mapped to OWASP WSTG, plus recon, network, cloud and OSINT checks, tracked per asset.


This is the best defence we have found against the way agent-driven testing actually goes wrong. It is not hallucination. It is skimming. An agent that stumbles onto an interesting SQL injection in the first twenty minutes will happily spend the rest of the session on it and then report a thoroughly successful engagement.

Asking what is left on an asset gives you the current state of every check, so "what have I not looked at on this host" becomes a question with an answer. Marking a check as not applicable counts as resolved, and that matters: "we looked, there is no file upload here" is a genuine testing outcome and belongs in the record rather than sitting in the untested pile forever.

The auth, authz and session categories are where OTP bypass and broken authentication live, and they are exactly the checks an excited agent skips on its way to something noisier. The blank password came out of exactly that part of the list: a login check nobody would call interesting, and a form that issued a valid session when the password field was submitted empty. Nothing would have gone back to that check if the record had not been sitting there saying it was untested. A count of resolved checks against the total is an honest statement about where an engagement stands. "I tested the application thoroughly" is not.

Findings that hold up to review

A finding has to stand on its own, because whoever reviews it will not have the tester sitting next to them explaining what they meant.



A finding carries severity and status, a CVSS vector, CWE and CVE, and separate fields for details, impact, proof of concept and remediation, because those are what a report needs and what a reviewer checks.


The proof of concept field is the one that does the work. It either contains steps that reproduce or it does not, and a reviewer can tell which without asking anyone. Evidence attaches to the finding itself rather than living in a folder someone has to match up later.


The rules for writing a finding sit inside the tool the agent calls to file one, so it reads them as it writes rather than somewhere far back in the session. They tell it to keep each field to a paragraph, keep hostnames out of the title, and say what needs to change in the remediation instead of pasting config and version numbers that may be wrong for the customer's stack.

What the human still does

The agent is fast and it is sometimes wrong, and the workflow assumes both. Everything it writes is an ordinary row in the browser that you can edit, reclassify or delete, and it picks up your corrections the next time it reads.

Three habits do the work. We read the findings themselves rather than the agent's account of them, because the finding rows are what the customer actually gets. We check the coverage before believing any of it, because an agent can come back with six findings having cleared nine checks out of 135. Findings are not coverage. And we set the severity ourselves, because whether something is a finding at all, and how bad it is, is a call a human should make.

The thing that does not change is responsibility. Scope marking and the Red Team Instructions are guardrails, not authorisation. Red team exercises run under a signed engagement letter, and the agent acts entirely on your authority. Everything it does is yours.

What the model does and what the record does

GLM 5.3 does the reasoning. Reading a minified bundle and noticing that a string is a live key. Stripping the auth header off a request and noticing the data still comes back. Putting a single wildcard into a customer search field, then reading the response closely enough to work out that it had returned every customer in the database rather than an error. Going back at an OTP flow after the obvious attempt failed. That is a model capability question, and a better model gives you better testing.

Red Clippy is what makes that add up to an engagement. It contributes memory, correlation, coverage and evidence discipline, and it contributes them identically regardless of what is driving the agent. The two together are why a session ending no longer takes the engagement with it.

Red Clippy was originally our own internal tool, built for our engagements. We have now put it out publicly, because we think other testers will get the same use out of it. It is open source, from the Cyber Security and Privacy Foundation. Source is at github.com/CSPF-Founder/red-clippy and the documentation, including how to set all of this up, is at cspf-founder.github.io/red-clippy. Bug reports and any other contributions are welcome.