Three harnesses, one method.

Three open eval harnesses, side by side: legal RAG citations, what AI search says about a person, and LLM résumé screeners. Each row is one harness — the question it answers, the number it stands on, how it was run.

  • CrossSource

    When the RAG cites a court opinion, does that opinion support the claim?

    [judged]

    0.994

    citation precision, strict prompt (baseline 0.981) · n=25 · judge 15/15 vs human

    • 22 opinions · 25 questions
    • LLM judge, blind-validated 15/15
    • BM25 top-5 · baseline vs strict
  • mirror-eval

    What do AI search engines say about me, and is it true and sourced to something I control?

    [counted]

    13 → 63

    Perplexity citations, wave 1 → wave 2, of 63 probes · Claude 0 → 0

    • 4 engines · 63 search probes each
    • judges failed 3/40, 12/40 → counted instead
    • 2 waves · Aug → Sep 2026
  • screener-eval

    Do employer names and evidence links change what an LLM résumé screener scores?

    [null]

    < 1 pt

    score shift when every employer is swapped · 885 calls · 21 postings

    • 21 postings · 2 screeners · 5 reps
    • no judge in the loop
    • 95% CIs straddle zero

One judge passed and caught a bug. One judge failed and the study survived on counts. One needed no judge at all. Same method three times — that is the point.

Table — the three harnesses compared
Harness QuestionJudgeCounted metricSampleVerdict
CrossSource When the RAG cites a court opinion, does that opinion support the claim?LLM judge, blind-validated against a human: 100% (15/15)Claim-level citation precision, 0.981 → 0.994; recall 0.760 in both configurations22 opinions · 25 questionsJudge passed and caught a harness bug; prompting buys precision, retrieval owns recall
mirror-eval What do AI search engines say about me, and is it true and sourced to something I control?Two LLM judges from different model families — failed blind validation, 3/40 and 12/40Probes citing an owned surface, of 63 search-mode probes per engine — counted, judge-free4 engines · 63 search probes each · 2 wavesPerplexity 13 → 63, Claude 0 → 0; index access explains the split
screener-eval Do employer names and evidence links change what an LLM résumé screener scores?None — nothing is judged by a modelPaired delta in fit score (B − A), 10,000-resample bootstrap 95% CI1 résumé · 21 postings · 2 screeners · 885 callsEmployer swap and link deletion each moved the score under a point; which screener read it moved it 22

Build the judge. Validate it blind against a human. Count what you can. Report the rest as bands.

If you're shipping an LLM product, let's talk about what your evals miss.