The model never gets the last word.
I build the audit operator I hunt with: a thin controller over deterministic gates, with language models as advisers inside it. This page is architecture only. No target, program or finding is ever named here.
The model is advisory
A model can propose a hypothesis. It cannot pass a gate, write a finding, set severity or mark anything ready. One deterministic state store owns all of that.
Unknown is not pass
A gate that could not run reports unknown. Silence about a skipped step is treated as a defect, not as success.
Nothing submits itself
The pipeline ends in one of three states for a human to read. Submission is a person, every time.
Hold beats submit
A false submission is permanent; a held finding can be re-proved later. The system is built to prefer zero findings over one fabricated one.
Five stages, three ways out.
Scroll to run it. Each stage either hands evidence forward or stops the audit, and the stop is recorded with its reason.
The pre-hunt stage exists because 73.2% of my 97 logged rejections would have died there. The data ↗
Recon
Pin the commit, transcribe the program page, map the in-scope file tree. The audit gets an identity and a manifest before anyone reads code.
The output is a manifest and a ranked list of attack surfaces. Nothing is activated automatically; a human approves every target.
Pre-hunt gates
Fork novelty, severity floor, every exclusion clause, scope tree, prior audits, test intent. A failed gate stops the hunt before money or time is spent.
This is where 73% of my logged rejections would have died. The gates are deterministic checks over the program text and the repository, not model opinions.
Fuzz
A hand-written harness per project, plus generated invariant suites. Failing invariants become leads, never conclusions.
A green fuzz run is reported as coverage ("N mechanisms unexplored"), never as "nothing found". A killed campaign is a signal to investigate, not a harness bug.
Hypotheses
Assumption inversion first: for every in-scope function, what does it assume about caller, ordering and external calls, and does the code enforce it?
Model output is advisory. It can open a hypothesis only after deterministic validation; it can never close a gate, write a finding or change severity.
Realism gates
Attacker model, PoC privilege audit, impact reachability, realistic configuration, deployed state, program rules, and an adversarial reviewer that writes the likely rejection.
A real mechanism with unreachable impact is held, not submitted. The state exists because it happens: a mechanism proven in code that no real attacker could profit from.
Three outcomes. None of them is "submit".
The honest limits.
An AI security system that only lists its strengths is a marketing page. These are the ones I know about.
- It cannot see a duplicate that lives only in a program's private queue. The duplicate class is the largest in my log, and the hardest.
- Its calibration is in-sample. The rules were derived from the same rejections they are scored on.
- Machinery is not a result. Building the pipeline repeatedly ate the time it was meant to save, so it now runs under fixed budgets and I hunt by hand alongside it.
- Model judgement on severity and reachability is not trusted, which makes the realism gates strict and sometimes slow.
Building agents that have to be right?
I write about the engineering as it happens. Questions are welcome by DM or email.