Use Cases Platform Security Price About

How eulaw.ai compares on real legal questions

We measured it rather than asserting it. The same questions were put to eulaw.ai and to two frontier models, each answer checked against the official source.

Measured 2026-08-30 · 345 questions per system

Answer quality across every question

A refusal, a non-answer and an error all count as a bad response rather than a neutral abstention. In legal work a confident silence is still a failed answer.

  • Correct
  • Partial
  • Bad
share of all 345eulaw.aifull retrievaleulaw.ai — Correct: 79.1%79%eulaw.ai — Partial: 5.5%eulaw.ai — Bad: 15.4%15%15.4% badGPT-5.6bare, no retrievalGPT-5.6 — Correct: 49.6%50%GPT-5.6 — Partial: 9.9%10%GPT-5.6 — Bad: 40.6%41%40.6% badClaude Sonnet 5bare, no retrievalClaude Sonnet 5 — Correct: 39.7%40%Claude Sonnet 5 — Partial: 4.3%Claude Sonnet 5 — Bad: 55.9%56%55.9% bad

Pooled over the questions every system answered, matched question by question, so a part-finished run cannot move the result by changing which questions are counted.

Where the difference comes from

The same run, split by what each question asks for. Higher is better throughout.

  • eulaw.ai
  • GPT-5.6
  • Claude Sonnet 5
% correct0255075100Name the rightactn=80eulaw.ai — Name the right act: 95%95GPT-5.6 — Name the right act: 53.8%54Claude Sonnet 5 — Name the right act: 41.3%41Legal substancen=91eulaw.ai — Legal substance: 69.2%69GPT-5.6 — Legal substance: 65.9%66Claude Sonnet 5 — Legal substance: 59.3%59Policy, votes,officialsn=93eulaw.ai — Policy, votes, officials: 78.5%79GPT-5.6 — Policy, votes, officials: 33.3%33Claude Sonnet 5 — Policy, votes, officials: 14%14Amendment chainsn=30eulaw.ai — Amendment chains: 40%40GPT-5.6 — Amendment chains: 3.3%3Claude Sonnet 5 — Amendment chains: 0%0Multi-turnconversationsn=51eulaw.ai — Multi-turn conversations: 96.1%96GPT-5.6 — Multi-turn conversations: 70.6%71Claude Sonnet 5 — Multi-turn conversations: 72.5%73100 turns deepn=100eulaw.ai — 100 turns deep: 78%78GPT-5.6 — 100 turns deep: 43%43Claude Sonnet 5 — 100 turns deep: 14%14

How recent the law is

A language model knows only what it was trained on, and legislation does not stop for a training cutoff. This is the sharpest split in the run, and it is also the least surprising one.

  • eulaw.ai
  • GPT-5.6
  • Claude Sonnet 5
% correct0255075100before 2010n=42010 to 2019n=62020 to 2024n=172025 onwardsn=13eulaw.ai — before 2010: 75%eulaw.ai — 2010 to 2019: 100%eulaw.ai — 2020 to 2024: 100%eulaw.ai — 2025 onwards: 76.9%GPT-5.6 — before 2010: 50%GPT-5.6 — 2010 to 2019: 66.7%GPT-5.6 — 2020 to 2024: 35.3%GPT-5.6 — 2025 onwards: 0%Claude Sonnet 5 — before 2010: 25%Claude Sonnet 5 — 2010 to 2019: 33.3%Claude Sonnet 5 — 2020 to 2024: 23.5%Claude Sonnet 5 — 2025 onwards: 0%

Does it invent acts that do not exist

Asked which acts amended a given law, the answer is a set with a known membership. Precision is the share of named acts that are real, which for a legal product is the number that matters most.

SystemActs named that are realReal acts foundInvented outright
eulaw.ai82.1% 85.7% 12
GPT-5.615.4% 8.6% 59
Claude Sonnet 5Declined every question in this track, so it invented nothing and answered nothing

Retrieval reduces fabrication by roughly a factor of five here. It does not remove it, and our own remaining count is tracked as a defect rather than presented as a floor.

Method and limits

Every system received identical question text, and every raw answer was normalised by one extractor before scoring, so no system is penalised for answering in prose rather than in a fixed format. The extractor also flags refusals.

The two models were tested through their APIs with no legal database, web search or retrieval. That isolates the value of the corpus rather than the model, and it is not how a lawyer uses either product. A second comparison against the consumer apps with search enabled will follow.

Known limits: one sample per question, so small gaps are noise. Identifier questions are graded against our own corpus, which measures whether the product surfaces what it holds rather than testing memory. The fact-coverage judge is itself a language model.