Writing¶
Occasional posts on measuring LLM systems — written from the pipelines I run myself, with the runs behind them public wherever they can be. The decisions they arrive at end up in the decision records; these are the working out.
2026¶
9 September 2026 — My LLM eval cried wolf. Here's what I measured.
A case went from 5/5 to 2/5 with nothing changed. Measuring the noise floor of an LLM-judged eval, what it caught the week after, and where it still cannot see.