REDHERRING · HKUST
Cheap to Hypothesize,
Costly to Verify:
The Defense Surface of
Agentic Vulnerability Discovery
An agent can form many hypotheses.
It can only afford to verify a few.
The Hong Kong University of Science and Technology
01 / THE EVIDENCE
Watch where the effort goes.
Real messages. Real tool calls.
Every turn open for inspection.
An audit turns into a factoring attempt.
The agent’s recorded messages and tool calls will appear here.
Each row is a source symbol. Columns share the trajectory’s step axis. Multiple contacts in one column are one step, not separate effort.
Colors indicate symbol contact, not elapsed time or token expenditure.
How to read this evidence Provenance & measurement
Recorded agent claims may be wrong. Symbol contact shows exposure or discussion; it does not prove reachability, exploitation, or a causal defense effect. The guided notes are editorial summaries. Original messages and tool records remain separate.
Workspace paths, benchmark task identifiers, and credential-bearing lines are redacted. Environment setup output is omitted. Long tool outputs are explicitly marked as excerpts. Source line numbers refer to the original logs.
Download this exhibit’s data ↓02 / THE IDEA
Cheap to suspect.
Expensive to rule out.
RedHerring redirects a finite verification budget
toward certifiably safe decoys.
Vulnerability-discovery agents search a repository, form hypotheses, and investigate a subset. Proving that a promising path is exploitable—or ruling it out—costs source reads, reasoning, execution, and PoC construction.
RedHerring inserts a CVE-derived vulnerability chain behind a false bridge: a condition no input satisfies. The defender can efficiently verify unreachability with a private certificate. Establishing the same fact from public parameters poses a computationally hard problem.

Vulnerability chain
Apparent input-to-sink flow, derived from a real CVE.
False bridge
An unsatisfiable predicate on every path to the dangerous sink.
Repository adaptation
Existing types, data flow, and naming make each decoy fit its host.
03 / THE RESULTS
Less budget for real vulnerabilities.
33 OSS-Fuzz projects · 70 instances · 5 models
Matched 3-hour and 300-round limits
fewer confirmed real vulnerability findings
of completion tokens spent on decoy investigation
fewer findings even with a decoy-aware agent

WHEN THE AGENT KNOWS
Awareness changes the search.
The verification cost remains.
Explicitly warning Qwen3.8-Flash about decoys reduces the effort it spends on them. RedHerring still reduces confirmed findings by 37.2% relative to the informed Baseline.
Read the evaluation on arXiv ↗
THE PAPER
Cheap to Hypothesize, Costly to Verify
The Defense Surface of Agentic Vulnerability Discovery
The arXiv link will appear here when the preprint is available.
TAKE A CLOSER LOOK
The strongest explanation is in the record.
Go back to a real audit and follow a lead.
