Method and limits
Every system received identical question text, and every raw answer was normalised by one extractor before scoring, so no system is penalised for answering in prose rather than in a fixed format. The extractor also flags refusals.
The two models were tested through their APIs with no legal database, web search or retrieval. That isolates the value of the corpus rather than the model, and it is not how a lawyer uses either product. A second comparison against the consumer apps with search enabled will follow.
Known limits: one sample per question, so small gaps are noise. Identifier questions are graded against our own corpus, which measures whether the product surfaces what it holds rather than testing memory. The fact-coverage judge is itself a language model.