REDHERRING · HKUST

Cheap to Hypothesize,
Costly to Verify:

The Defense Surface of
Agentic Vulnerability Discovery

An agent can form many hypotheses.
It can only afford to verify a few.

Kaikai Zhang* · Zihan Zhang* · Yuchong Xie · Zesen Liu
Shuangjie Yao · Zhixiang Zhang · Dongdong She

The Hong Kong University of Science and Technology

01 / THE EVIDENCE

Watch where the effort goes.

Real messages. Real tool calls.
Every turn open for inspection.

Loading recorded trajectory…

An audit turns into a factoring attempt.

The agent’s recorded messages and tool calls will appear here.

Other stepTool contactReasoning mentionBoth
START

Colors indicate symbol contact, not elapsed time or token expenditure.

ASSISTANT MESSAGE059 / 300
How to read this evidence Provenance & measurement

Recorded agent claims may be wrong. Symbol contact shows exposure or discussion; it does not prove reachability, exploitation, or a causal defense effect. The guided notes are editorial summaries. Original messages and tool records remain separate.

Workspace paths, benchmark task identifiers, and credential-bearing lines are redacted. Environment setup output is omitted. Long tool outputs are explicitly marked as excerpts. Source line numbers refer to the original logs.

Download this exhibit’s data ↓

02 / THE IDEA

Cheap to suspect.
Expensive to rule out.

RedHerring redirects a finite verification budget
toward certifiably safe decoys.

Vulnerability-discovery agents search a repository, form hypotheses, and investigate a subset. Proving that a promising path is exploitable—or ruling it out—costs source reads, reasoning, execution, and PoC construction.

RedHerring inserts a CVE-derived vulnerability chain behind a false bridge: a condition no input satisfies. The defender can efficiently verify unreachability with a private certificate. Establishing the same fact from public parameters poses a computationally hard problem.

RedHerring method overview: repository search reaches a decoy adapted to ordinary code; a false bridge separates its vulnerability chain from the dangerous sink; private construction information certifies unreachability.
FROM THE PAPER A vulnerability chain attracts verification. A false bridge keeps its dangerous sink unreachable. Repository adaptation makes the decoy read as ordinary program logic. Full figure ↗
01

Vulnerability chain

Apparent input-to-sink flow, derived from a real CVE.

02

False bridge

An unsatisfiable predicate on every path to the dangerous sink.

03

Repository adaptation

Existing types, data flow, and naming make each decoy fit its host.

03 / THE RESULTS

Less budget for real vulnerabilities.

33 OSS-Fuzz projects · 70 instances · 5 models
Matched 3-hour and 300-round limits

38.7–60.4%

fewer confirmed real vulnerability findings

30.6–51.5%

of completion tokens spent on decoy investigation

37.2%

fewer findings even with a decoy-aware agent

Paper results comparing Baseline and RedHerring across five models, with decoy effort shares and the relationship between diversion and reduction in confirmed findings.
MAIN EVALUATION Confirmed crash signatures serve as the operational proxy for vulnerability findings. Effort shares are averaged across instances; time shares are estimated. These are the paper’s aggregate results, separate from the illustrative trajectories above. Full figure ↗

WHEN THE AGENT KNOWS

Awareness changes the search.
The verification cost remains.

Explicitly warning Qwen3.8-Flash about decoys reduces the effort it spends on them. RedHerring still reduces confirmed findings by 37.2% relative to the informed Baseline.

Read the evaluation on arXiv ↗
Paper comparison of Qwen3.8-Flash with and without a decoy-awareness prompt.

THE PAPER

Cheap to Hypothesize, Costly to Verify

The Defense Surface of Agentic Vulnerability Discovery

Read on arXiv ↗

TAKE A CLOSER LOOK

The strongest explanation is in the record.
Go back to a real audit and follow a lead.

Explore the trajectories ↑

REDHERRING

Cite this work

arXiv:2609.35909 · BibTeX