Benchmark · AI-assisted radiology reporting

LAIBench

A governance-oriented benchmark for turning an exam descriptor and concise findings into a faithful radiology report — scored where it matters clinically, not on prose. A missed or fabricated critical finding is a hard veto, never a soft deduction.

5
weighted clinical dimensions
120
controlled pt-BR cases
90.0%
Laudos.AI clinical score
0
prose / aesthetic axes

Designed so form never rescues substance

01 — HOW IT SCORES
01

Hard critical-finding veto

A missed or fabricated critical finding caps the score and forces FAIL — regardless of how polished the rest of the report reads.

02

No prose / aesthetic axis

No standalone style, fluency or "communication quality" dimension. The only discourse signals are minor, non-gating, fallback-only.

03

Conservative combination

The combined dimension score is MIN(deterministic, judge): an optional LLM judge can lower a score but never inflate it past the gate.

04

Tamper-resistant numbers

Every run is re-scored through the gated combiner; a relabeled critical-miss is rejected before it reaches the leaderboard. scoringHash + suiteHash provenance.

Dimensions: CRIT 30% · QUAL 25% · TERM 20% · GUIDE 15% · RAG 10%, with hard failure gates. LAIBench is a technical benchmark framework — not a medical device, not regulatory approval, not clinical validation.

Leaderboard

02 — RESULTS

Two governance tiers, never mixed: controlled-eval (gated, first-party, aggregate-only) and public-smoke (synthetic, contaminable, diagnostic baselines — not a ranking). See the benchmark cards for case source, leakage risk, and adjudication status.

Research

03 — PAPERS
Our principle

A beautiful report that misses a pneumothorax should fail.
A plain report that detects it should win.