Legal AI Quality Intelligence · EU / UK / UAE
The evaluation lab for legal AI.
Bench measures how accurately AI systems perform legal work — by jurisdiction, against expert-labeled gold standards. You get a defensible accuracy score, a hallucination rate, and the exact corrections. Not opinions. Measurements.
The deliverable
A measurement, not a marketplace.
Your team submits model outputs. Bench returns a structured measurement report scored by a versioned rubric and verified by practising lawyers in the target jurisdiction — who remain production infrastructure, invisible and asynchronous. The product is the number and the correction, not access to people.
Every metric in a Bench report is computed by a deterministic, documented formula, so the same evaluations always reproduce the same score. That is what makes it citable in a board deck, a diligence room, or a vendor negotiation.
| Metric | What it tells you |
|---|---|
| Accuracy score | Overall correctness vs the jurisdiction-specific gold standard |
| Hallucination rate | Share of outputs with fabricated law, cases, or citations |
| Critical-error rate | Outputs with materially dangerous misstatements |
| Citation reliability | Cited authorities that are real, relevant, correctly used |
| Missing-law rate | Outputs omitting legally required elements |
| Error taxonomy | Failures classified: hallucination, misstatement, omission, wrong jurisdiction, bad reasoning |
| Gold-standard corrections | Expert-written correct answers for every failed item |
| Remediation memo | Prioritized fixes: data gaps, prompting, guardrails |
Featured benchmarks
Jurisdiction-specific. Expert-labeled. Versioned.
Generic benchmarks measure English-language legal trivia. Bench maintains 75 gold-standard items across three regulated jurisdictions, with held-out sets that never touch a public page — so vendors can't train on the test.
Process
Asynchronous by design. Three business days, typically.
| Step | What happens | Who does it |
|---|---|---|
| 01 | You submit model outputs — via API run, output upload, or public interface | Your team |
| 02 | Items are scored against the rubric; every score is verified or produced by a practising lawyer in the target jurisdiction, with QA sampling and calibration | The expert bench |
| 03 | The rubric engine computes the report deterministically — accuracy, hallucination rate, taxonomy, breakdowns | Instrumentation |
| 04 | You receive the measurement report with gold-standard corrections and a prioritized remediation memo | Delivered async |
No calls required, ever. Expedited turnaround is available on managed programs.
Who buys measurement
Built for teams whose AI touches the law.
The bench
Practising lawyers. Calibrated. Invisible.
Every gold standard is written or verified by a vetted practitioner in the target jurisdiction — admitted, in practice, paid per task. Experts pass a paid assessment and a calibration round before their scores count, and ongoing QA sampling keeps them honest.
Reports disclose expert credentials by class — jurisdiction, practice area, years of post-qualification experience — never by name. No profiles, no marketplace, no directory. The measurement is the product.
Credential classes — example report
- EU regulatory · 8y PQE · MiCA / funds
- UK data protection · 6y PQE · ICO practice
- UAE / DIFC commercial · 10y PQE
Inter-rater agreement from calibration rounds is published on the methodology page once the first cycle completes.