Legal AI Quality Intelligence · EU / UK / UAE

The evaluation lab for legal AI.

Bench measures how accurately AI systems perform legal work — by jurisdiction, against expert-labeled gold standards. You get a defensible accuracy score, a hallucination rate, and the exact corrections. Not opinions. Measurements.

The deliverable

A measurement, not a marketplace.

Your team submits model outputs. Bench returns a structured measurement report scored by a versioned rubric and verified by practising lawyers in the target jurisdiction — who remain production infrastructure, invisible and asynchronous. The product is the number and the correction, not access to people.

Every metric in a Bench report is computed by a deterministic, documented formula, so the same evaluations always reproduce the same score. That is what makes it citable in a board deck, a diligence room, or a vendor negotiation.

MetricWhat it tells you
Accuracy scoreOverall correctness vs the jurisdiction-specific gold standard
Hallucination rateShare of outputs with fabricated law, cases, or citations
Critical-error rateOutputs with materially dangerous misstatements
Citation reliabilityCited authorities that are real, relevant, correctly used
Missing-law rateOutputs omitting legally required elements
Error taxonomyFailures classified: hallucination, misstatement, omission, wrong jurisdiction, bad reasoning
Gold-standard correctionsExpert-written correct answers for every failed item
Remediation memoPrioritized fixes: data gaps, prompting, guardrails

Featured benchmarks

Jurisdiction-specific. Expert-labeled. Versioned.

Generic benchmarks measure English-language legal trivia. Bench maintains 75 gold-standard items across three regulated jurisdictions, with held-out sets that never touch a public page — so vendors can't train on the test.

Process

Asynchronous by design. Three business days, typically.

StepWhat happensWho does it
01You submit model outputs — via API run, output upload, or public interfaceYour team
02Items are scored against the rubric; every score is verified or produced by a practising lawyer in the target jurisdiction, with QA sampling and calibrationThe expert bench
03The rubric engine computes the report deterministically — accuracy, hallucination rate, taxonomy, breakdownsInstrumentation
04You receive the measurement report with gold-standard corrections and a prioritized remediation memoDelivered async

No calls required, ever. Expedited turnaround is available on managed programs.

Who buys measurement

Built for teams whose AI touches the law.

Legal-AI vendorsProve accuracy to buyers and boards with an independent, jurisdiction-specific score — before a prospect finds the failure mode themselves.diagnostic → managed
Compliance softwareYour product answers regulatory questions. Bench tells you where it hallucinates MiCA, GDPR, or VARA — and what the correct answer was.managed program
Law firms & in-houseEvaluate the AI tools you are about to trust with client work. Vendor-independent, expert-verified, private.diagnostic audit
AI labs & diligenceJurisdiction-deep evaluation data and audits for legal-domain capability claims, structured for repeat measurement.dedicated program

The bench

Practising lawyers. Calibrated. Invisible.

Every gold standard is written or verified by a vetted practitioner in the target jurisdiction — admitted, in practice, paid per task. Experts pass a paid assessment and a calibration round before their scores count, and ongoing QA sampling keeps them honest.

Reports disclose expert credentials by class — jurisdiction, practice area, years of post-qualification experience — never by name. No profiles, no marketplace, no directory. The measurement is the product.

Credential classes — example report

  • EU regulatory · 8y PQE · MiCA / funds
  • UK data protection · 6y PQE · ICO practice
  • UAE / DIFC commercial · 10y PQE

Inter-rater agreement from calibration rounds is published on the methodology page once the first cycle completes.

Start

Know your number before your buyers do.

Request a diagnostic auditJoin the expert bench