Independent security researcher

I study why securityfindings don't get paid, and build the AIsystems that fix it.

Blockchain security research · agentic AI engineering

The log

97 submissions.
Every one rejected.

Four months of bug-bounty work, logged row by row. Each square is one rejected finding.

The question

Which check would have caught it first?

I went back through every row and assigned the earliest gate that would have stopped it, before the PoC, before the write-up.

The answer
73.2%

71 of 97 were catchable before I read a line of code.

The problem was mostly targeting, not analysis. So I built tools that check the targeting first. See the dataset ↗

  • 34 · Duplicate check: A known-issues list or public record already had it.
  • 29 · Program rules: Exclusion clause, severity floor or scope tree.
  • 8 · Prior audit & tests: An earlier audit or the test suite said it was known or intended.
  • 26 · Only after the code: Design intent, no attacker win, realism, or the payout floor.
Tools

Three open repos. Rules, not data.

None of them ships a program name, a chain or a reward table. You point them at your own private ledger.

viability-gate, live

Score the finding before you write the PoC.

Offline, deterministic, single binary, no API key. Every fire shows the matched snippet and the sentence it came from, so you can argue with it. A gate it could not run reports ?, never a pass.

How the tools fit together
~/findings · vg check
vg check finding.mdviability-gate 0.2.0  —  finding.md  OUTCOME   KILL            A hard blocker fired. Do not submit as written.  SCORE     60 / 100   (uncalibrated heuristic; not a probability; threshold 50)  HARD  P2   Privileged-role-gated against a program exclusion             ↳ matched: admin can             ↳ in sentence: An admin can set the pause flag to zero, which causes every vault operation to             → Check whether an unprivileged attacker can trigger the bug without role             → compromise. If not: hold. The one escape is UNKNOWING admin harm - a             → reasonable admin action with a non-obvious destructive second-order             → effect.             (docs/rules/P2.md; evidence: in-sample; origin rows: 3)  PASS  P1, P6, P12, P13, P3, P5, P11, P20, P7, P16, P18, P21, P22, P15, P4, P9, P8, P10, P14, P17, P19  SKIP  G1 (no ledger supplied)   G2 (no ledger supplied)   G3 (no scope tree supplied)   G4 (no prior-audit surface supplied)   G5 (no test suite supplied)   G6 (fork status unknown)   G7 (no own-history log supplied)   D1 (D1 is a corpus match)  Gates:  G1 ?   G2 ?   G3 ?   G4 ?   G5 ?   G6 ?   G7 ?   D1 ?  ────────────────────────────────────────────────────────────  Advisory, never authoritative. This tool produces an outcome with reasons;  it does not decide, and it cannot read your target's live rules.  Over-filtering is its own failure mode - the catalog is a prioritiser, not  a kill-switch. A triage rejection is not a final verdict. CLEAR means only  that no known blocker was found by this tool.

Real output from the viability-gate README. It is generated by running the binary, and CI fails if the two drift.

Measured, including the failure

I tested the tool against my own rejections.

The honest summary is the gap between the two bars. Reporting either number alone would misrepresent it.

Strict: named the gate that actually decided it23.7%
Lenient: named any gate-mapped reason to stop76.3%

59 rejections matched to their original finding text, scored on whether the tool flagged the gate the program closed it on. In-sample, so closer to a ceiling than an estimate. It is good at telling you to stop and bad at telling you why. Method and caveats ↗

Agentic AI engineering

An audit pipeline where the model never gets the last word.

Deterministic gates own state, severity and readiness. Model output is advisory. Nothing is ever submitted automatically.

What it catches, and what it can't
STAGE 01

Recon

Pin the commit, transcribe the program page, map the in-scope file tree. The audit gets an identity and a manifest before anyone reads code.

STAGE 02

Pre-hunt gates

Fork novelty, severity floor, every exclusion clause, scope tree, prior audits, test intent. A failed gate stops the hunt before money or time is spent.

STAGE 03

Fuzz

A hand-written harness per project, plus generated invariant suites. Failing invariants become leads, never conclusions.

STAGE 04

Hypotheses

Assumption inversion first: for every in-scope function, what does it assume about caller, ordering and external calls, and does the code enforce it?

STAGE 05

Realism gates

Attacker model, PoC privilege audit, impact reachability, realistic configuration, deployed state, program rules, and an adversarial reviewer that writes the likely rejection.

TERMINAL STATES

Three outcomes. None of them is "submit".

READY_FOR_HUMAN_REVIEWEvery gate passed. A person decides.
HELDMechanism real, impact not. Kept with its evidence.
ZERO_FINDINGSAfter a devil's-advocate pass on every critical function.
Follow the work

Fewer findings. All of them real.

New articles, datasets and tools, as they ship. For anything else, a DM on X or an email reaches me directly.