LLM products fail quietly. I build the evals that catch them.
I'm Zoeb Nomi. At Instead — an AI-native tax research and planning platform, the first new entrant to clear IRS e-filing approval alongside incumbents — I own end-to-end output quality for a production tax-research LLM: citation accuracy, model benchmarking, and the evaluation loops that catch regressions before release. In tax research, a wrong citation isn't a UX bug. It's a compliance risk.
0.994
citation precision under strict citation discipline (CrossSource v0.1; baseline 0.981)
100% (15/15)
blind human–judge agreement validating the LLM judge
~95%
citation accuracy held on a production golden set
Three harnesses, one method.
Build the judge. Validate it blind against a human. Count what you can. Report the rest as bands.
-
CrossSource
[judged]
0.994
citation precision, strict prompt (baseline 0.981) · n=25 · judge 15/15 vs human
-
mirror-eval
[counted]
13 → 63
Perplexity citations, wave 1 → wave 2, of 63 probes · Claude 0 → 0
-
screener-eval
[null]
< 1 pt
score shift when every employer is swapped · 885 calls · 21 postings
One judge passed and caught a bug. One judge failed and the study survived on counts. One needed no judge at all. Same method three times — that is the point.
All three harnesses are public repositories. This is the last year of commits, live from GitHub.
PM who ships code
-
Instrument
Flag → classify by failure mode → weekly review, scored on a four-dimension rubric: answerability, accuracy, citation quality, actionability.
-
Taxonomize
A failure taxonomy instead of a single score — so you know whether the fix belongs to prompting or to retrieval.
-
Validate the judge
An unvalidated eval reports wrong numbers with full confidence — precisely the failure mode the eval exists to catch.
If you're shipping an LLM product, let's talk about what your evals miss.