A governance-oriented benchmark for turning an exam descriptor and concise findings into a faithful radiology report — scored where it matters clinically, not on prose. A missed or fabricated critical finding is a hard veto, never a soft deduction.
A missed or fabricated critical finding caps the score and forces FAIL — regardless of how polished the rest of the report reads.
No standalone style, fluency or "communication quality" dimension. The only discourse signals are minor, non-gating, fallback-only.
The combined dimension score is MIN(deterministic, judge): an optional LLM judge can lower a score but never inflate it past the gate.
Every run is re-scored through the gated combiner; a relabeled critical-miss is rejected before it reaches the leaderboard. scoringHash + suiteHash provenance.
Dimensions: CRIT 30% · QUAL 25% · TERM 20% · GUIDE 15% · RAG 10%, with hard failure gates. LAIBench is a technical benchmark framework — not a medical device, not regulatory approval, not clinical validation.
Two governance tiers, never mixed: controlled-eval (gated, first-party, aggregate-only) and public-smoke (synthetic, contaminable, diagnostic baselines — not a ranking). See the benchmark cards for case source, leakage risk, and adjudication status.
A beautiful report that misses a pneumothorax should fail.
A plain report that detects it should win.