<?xml version='1.0' encoding='utf-8'?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title>Writing — digline</title>
    <link>https://digline.dev/blog/</link>
    <description>Occasional posts on measuring LLM systems: what the numbers do when nothing has changed, and what it costs to find out.</description>
    <language>en</language>
    <atom:link href="https://digline.dev/feed.xml" rel="self" type="application/rss+xml" />
    <lastBuildDate>Fri, 18 Sep 2026 00:00:00 +0000</lastBuildDate>
    <item>
      <title>The case that errors is usually the case that was failing</title>
      <link>https://digline.dev/blog/denominator-trap/</link>
      <description>A case that errors drops out of the denominator, so precision and recall get measured over a smaller suite than the run you are comparing against. The score goes up.</description>
      <pubDate>Fri, 18 Sep 2026 00:00:00 +0000</pubDate>
      <guid isPermaLink="true">https://digline.dev/blog/denominator-trap/</guid>
    </item>
    <item>
      <title>Bad evals, my own: five exercises from two LLM judges</title>
      <link>https://digline.dev/blog/bad-evals-my-own/</link>
      <description>I applied the reading Dan Luu applies to other people's benchmarks to my own two LLM judges. Numbers first, explanations after.</description>
      <pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate>
      <guid isPermaLink="true">https://digline.dev/blog/bad-evals-my-own/</guid>
    </item>
    <item>
      <title>My LLM eval cried wolf. Here's what I measured.</title>
      <link>https://digline.dev/blog/my-llm-eval-cried-wolf/</link>
      <description>A case went from 5/5 to 2/5 with nothing changed. How I measured the noise floor of an LLM-judged eval, what it caught the week after, and where it still can't see.</description>
      <pubDate>Wed, 09 Sep 2026 00:00:00 +0000</pubDate>
      <guid isPermaLink="true">https://digline.dev/blog/my-llm-eval-cried-wolf/</guid>
    </item>
  </channel>
</rss>
