Evaluations

Questions to ask when an AI tool is presented as accurate, reliable or safe.

What is it?

LLM evaluations are a test, they can measure how well an AI tool does a task. For example: "Does this tool find the correct case law?"

Why it matters

Vendors use evaluation results to support claims that a tool is "accurate" or "safe" - it is key for lawyers advising on decisions that rely on those claims to understand the ways in which such evaluations can be rendered unreliable or invalid. A test may not measure what is being claimed. Results might be limited to tailored test-scenarios, not real-world tasks.

First consider

  1. What is the specific thing being tested, or the claim being made (e.g. "Our new tool solves case citation hallucinations," are they saying 0% hallucinated case citations? or 10% or what?)
  2. Be wary of terms like “accurate”, “reliable”, “safe”, “unbiased”, “human-level”, “validated” or “production-ready,” ask: what do you mean by this? What specific, measurable criteria justify this claim?

Questions to ask

Use the sections below to examine the claim, the test, the evidence and the evaluator.

QUESTIONS 111

The test and its setup

  1. Show us random sample of concrete questions & answers that were used in your evaluation.
  2. Have you tested this? Can you give us the test results? What evidence do you have for your claim?
  3. What did you measure? How did you measure it?
  4. When did you do your testing? What model did you use? In what configuration? What tools were available to the model? What system instructions?Is the configuration they used when testing, the same as your configuration in deployment?
  5. Was it independently verified or did you do your own testing?
  6. What was your testing methodology, how many times did you test, how did you judge (i.e. score the outputs)
  7. Were the test outputs improved through retries, expert prompting, manual selection, editing or other assistance? Were the outputs what an ordinary user would expect?
  8. What user settings would impact the results?
  9. Can you show how much each major component contributes to performance, for example, the model, search function, proprietary data and human review?
  10. What changes have occurred since the evaluation? Which findings have been revalidated against the current product?
  11. Can we identify and preserve the configuration used for a particular matter or decision, even after you or an upstream supplier updates the service?

QUESTIONS 1218

What the scores mean

  1. Why is your chosen measure a valid indicator of the capability, risk or outcome we care about? What evidence supports that connection?
  2. Could a system score highly while still failing the real task? Show examples and explain how the evaluation detects that mismatch.
  3. How were test tasks mapped to the actual activities, difficulty levels and consequences in our intended workflow?
  4. If you combine several measures into one score, provide the components, weights and rationale. Can strong performance in one area conceal a serious failure elsewhere?
  5. How were pass thresholds chosen, by whom, and before or after viewing results? What practical consequence distinguishes a pass from a fail?
  6. What evidence shows that differences in your scores correspond to meaningful differences in performance or risk? Would a small score increase change any real decision?
  7. do most systems achieve similarly high or low scores?

QUESTIONS 1925

The test examples

  1. Where did the test examples come from?
  2. How were examples selected? Provide inclusion criteria, exclusions and counts at each selection stage.
  3. How closely does the dataset resemble our work in subject matter, complexity, language, document quality, length, recency and frequency of difficult cases?
  4. Were examples or answers generated by AI? If so, how were they checked for correctness, realism, diversity and systematic bias?
  5. How did you investigate whether test questions or answers appeared in model training, fine-tuning, retrieval materials, demonstrations or earlier development work? What uncertainty remains?
  6. Who established the correct or acceptable answers, what relevant expertise did they have, and what authoritative evidence supports those answers?
  7. How were ambiguous questions, reasonable professional disagreement and later changes in the correct answer handled? Can we inspect disputed examples and their resolution?

QUESTIONS 2633

Scoring and difficult cases

  1. Provide the complete scoring rubric, with examples of a full pass, partial pass, minor error and serious failure.
  2. How many assessors scored each example independently, how often did they disagree, and how was disagreement resolved?
  3. If an AI model scored outputs, identify its version, instructions and settings. Was it the same model, model family or supplier as the product being evaluated?
  4. How was the AI scorer validated against qualified human assessors on representative and difficult cases? Provide its own error rates, including acceptance of materially wrong answers.
  5. Could text within the submitted answer manipulate the scorer’s instructions or score? What tests addressed that possibility?
  6. Can we inspect anonymised item-level inputs, outputs, scores and assessor explanations, including failures and borderline cases?
  7. Provide information about refusals, timeouts, empty answers, partial completions and technical failures during testing.
  8. How does performance change with long files, large document collections, poor scans, tables, footnotes, tracked changes or information buried in the middle of a document?

QUESTIONS 3437

Coverage and supervision

  1. Which languages, user groups or other relevant subgroups perform materially worse? Provide sample sizes and uncertainty, including combinations of characteristics where relevant.
  2. Which jurisdictions, courts, practice areas and source types are covered? Identify material gaps and the actual update delay for new or amended law.
  3. Does the product distinguish law applicable at a historical date from current law, and how does it handle choice-of-law or jurisdictional ambiguity?
  4. What exactly must the supervising lawyer verify, with what expertise and access to sources?

QUESTIONS 3843

Independence and a real-world pilot

  1. Who commissioned, paid for and performed the evaluation? Disclose ownership interests, referral arrangements and other commercial relationships with the product provider.
  2. Did the evaluator also develop the product, advise on passing the test or sell remediation services? How were those conflicts managed?
  3. What access did the evaluator receive, and what was withheld? Did limited access prevent testing of important claims or controls?
  4. Has another qualified party reproduced the findings? If so, did it independently select the data and methods, or repeat the same supplied test?
  5. Is the assurance conclusion from any independent evaluator based on actual direct testing, document inspection, interviews? Which material claims were not independently verified?
  6. Will you do a pilot on a representative sample selected by us?