LLM products fail quietly. I build the evals that catch them.

I'm Zoeb Nomi. At Instead — an AI-native tax research and planning platform, the first new entrant to clear IRS e-filing approval alongside incumbents — I own end-to-end output quality for a production tax-research LLM: citation accuracy, model benchmarking, and the evaluation loops that catch regressions before release. In tax research, a wrong citation isn't a UX bug. It's a compliance risk.

0.994

citation precision under strict citation discipline (CrossSource v0.1; baseline 0.981)

100% (15/15)

blind human–judge agreement validating the LLM judge

~95%

citation accuracy held on a production golden set

Three harnesses, one method.

Build the judge. Validate it blind against a human. Count what you can. Report the rest as bands.

  • CrossSource

    [judged]

    0.994

    citation precision, strict prompt (baseline 0.981) · n=25 · judge 15/15 vs human

  • mirror-eval

    [counted]

    13 → 63

    Perplexity citations, wave 1 → wave 2, of 63 probes · Claude 0 → 0

  • screener-eval

    [null]

    < 1 pt

    score shift when every employer is swapped · 885 calls · 21 postings

One judge passed and caught a bug. One judge failed and the study survived on counts. One needed no judge at all. Same method three times — that is the point.

All three harnesses are public repositories. This is the last year of commits, live from GitHub.

PM who ships code

  • Instrument

    Flag → classify by failure mode → weekly review, scored on a four-dimension rubric: answerability, accuracy, citation quality, actionability.

  • Taxonomize

    A failure taxonomy instead of a single score — so you know whether the fix belongs to prompting or to retrieval.

  • Validate the judge

    An unvalidated eval reports wrong numbers with full confidence — precisely the failure mode the eval exists to catch.

If you're shipping an LLM product, let's talk about what your evals miss.