Paired source-grounded evaluation

Benchmarks

This suite asks 20 frozen package questions. Each answer turns on a stale default, a lightly indexed implementation detail, or a disagreement between the reference manual and the source.

Suite
20 frozen cards
Coverage
155 probes · 22 packages
Model
GPT-5.6 Luna · medium
Comparison
Local corpus vs live web

Summary

Strong on evidence and token use. Mixed on time.

The pinned local corpus wins on quality, and the margin is largest when an answer must name the exact source. It also uses fewer tokens per answer. It does not answer faster.

Composite answer quality

+13.1 points

A blinded panel scores objective probes and holistic answer quality out of 100.

Biocontext97.6
WebSearch84.5

20 paired cards

Exact-source grounded correctness

+92.3 points

Correct claims whose citation matches the pinned source file.

Biocontext94.5%
WebSearch2.2%

107 source probes · 331 weight units

Total tokens per answer

67% fewer

Mean total tokens per answer across two complete paired runs.

Biocontext149k
WebSearch456k

40 answers per condition

Wall time per answer

21.8s slower

Mean end-to-end time per answer. Fewer tokens did not make local retrieval faster.

Biocontext107s
WebSearch86s

Biocontext answered faster on 17 of 40 pairs

Quality comes from one complete 20-card run. Efficiency pools two complete runs, for 40 answers per condition. Both arms used GPT-5.6 Luna at medium reasoning. The full figures show error bars and denominators.

Results

Answer quality

The quality run scores three things separately: general correctness, support in the supplied evidence, and an exact source match. The composite score adds a 70-point objective component to a 30-point holistic one.

Composite answer quality
Composite answer quality The registered 70-point objective component and 30-point holistic component, shown together.
Open SVG

Results

Retrieval efficiency

Both token figures pool the two complete paired runs. The bars keep cached input because it still counts as model usage. The scatter isolates newly processed context.

Pooled token efficiency
Pooled token efficiency Mean cached input, newly processed input, and output tokens per answer.
Open SVG

Results

Timing

The timing result is mixed. WebSearch finished sooner on average. Biocontext finished sooner on 17 of the 40 paired answers.

Pooled time efficiency
Pooled time efficiency Mean and median wall time across 40 answers per condition.
Open SVG