Composite answer quality
+13.1 pointsA blinded panel scores objective probes and holistic answer quality out of 100.
20 paired cards
Paired source-grounded evaluation
This suite asks 20 frozen package questions. Each answer turns on a stale default, a lightly indexed implementation detail, or a disagreement between the reference manual and the source.
Summary
The pinned local corpus wins on quality, and the margin is largest when an answer must name the exact source. It also uses fewer tokens per answer. It does not answer faster.
A blinded panel scores objective probes and holistic answer quality out of 100.
20 paired cards
Correct claims whose citation matches the pinned source file.
107 source probes · 331 weight units
Mean total tokens per answer across two complete paired runs.
40 answers per condition
Mean end-to-end time per answer. Fewer tokens did not make local retrieval faster.
Biocontext answered faster on 17 of 40 pairs
Quality comes from one complete 20-card run. Efficiency pools two complete runs, for 40 answers per condition. Both arms used GPT-5.6 Luna at medium reasoning. The full figures show error bars and denominators.
Results
The quality run scores three things separately: general correctness, support in the supplied evidence, and an exact source match. The composite score adds a 70-point objective component to a 30-point holistic one.
Results
Both token figures pool the two complete paired runs. The bars keep cached input because it still counts as model usage. The scatter isolates newly processed context.
Results
The timing result is mixed. WebSearch finished sooner on average. Biocontext finished sooner on 17 of the 40 paired answers.