# NRS relevance benchmark v1

Frozen before live execution on 2026-09-20. This is a synthetic, source-derived diagnostic benchmark, not an independently validated legal research benchmark.

## Comparison

Thirty ranking queries: 12 development and 18 held-out, with disjoint topic families. Four outside-corpus controls and six exact-citation controls are reported separately. Development results may guide later work; any resulting prompt or retrieval change needs a fresh held-out suite. Do not tune on this test set.

Primary comparison: original keyword candidate order versus Jev reranking of exactly the same 30 candidates. Both use the production retrieval, passage extraction, question definitions and pinned model. Report the app's separate keyword API results as a secondary baseline because its retrieval rules differ. Reranking cannot recover a provision absent from its candidate pool.

Primary endpoint: held-out known-target Hit@5. Also report Hit@1, Hit@10, reciprocal rank and nominated-target candidate coverage. Targets were nominated from preserved official text by the assistant, not a lawyer. They are incomplete provisional positives; these metrics are not precision or exhaustive recall. Report paired bootstrap percentile 95% intervals over query families (10,000 replicates, seed 20260920). Such intervals describe sampling within this small synthetic suite, not generalization to all research questions. Retain fallback searches in the comparison and disclose every failure.

## Labels and confidence

Grades: 0 unrelated; 1 topical but not useful; 2 useful support, definition, exception or procedure; 3 directly addresses the issue. Unknown grades remain unknown. The initial 60 source-derived section labels are provisional and are not passage labels. The blinded review page hides rank, model scores and proposed grades. Reviewers separately grade the complete section and the actual passage plus caption/heading seen by the model, identify themselves, and explicitly mark completed reviews. Hashes bind labels to queries, sources and model passages.

Human-label metrics: P@5 requires all five results judged. Pooled nDCG@5 requires every member of the declared pool judged; its ideal is limited to that pool, not the entire NRS. No missing judgment is treated as irrelevant. Model probabilities are not labels.

P(useful) is normalized P(grade 2)+P(grade 3). SDK confidence measures distribution concentration and is not P(correct). Once independent passage labels exist, report Brier score, five equal-width reliability bins, ECE and confidence versus exact-grade agreement, separately by split. Selected or incomplete review pools cannot establish population calibration. The frozen 0.8 useful-probability threshold is only a diagnostic, not a production cutoff. Outside-corpus controls show high-scored results for inspection; they are not automatically all irrelevant.

## Reproducibility and scope

Immutable input hashes, official-source snapshots, model/question version, raw validated provider responses, deterministic order, usage, failures and resumable request receipts accompany each run. Exact citation controls require no provider calls. Network submissions contain only frozen synthetic queries and public statutory passages. No researcher notes, client facts or credentials enter artifacts. The existing user-authorized key is read server-side. Budget: at most 272 attempted requests and 2 million reported input tokens, with conservative request-byte preflight. No automatic retries. This tests retrieval relevance, not current legal validity, completeness, legal outcomes or advice quality.
