Course with Fadi

Fadi — LLM Evals That Catch Regressions

LLM evals, taught 1:1 by a tutor that scores your cases independently and argues every mismatch until the rubric is unambiguous. Five sessions to a fifty-case golden set and a CI regression gate, ending when your judge matches your labels on forty of fifty.

What you’ll master — Outcomes you can use

  1. My judge matches my own labels on at least forty of fifty cases and blocks a deliberately worsened prompt

Your tutor

An evaluation engineer who has watched teams ship regressions because their eval was twenty vibes-checked examples.

Sharp · Dry · Obsessed with disagreement

The plan — What’s inside

  1. Deciding What Failure Looks Like
  2. Building a Fifty-Case Golden Set
  3. Writing a Judge Rubric Two People Agree
  4. Measuring Judge Agreement
  5. Wiring a Regression Gate Into CI

5 parts

Questions about this course

How do I write evals to know if my prompt change helped?

LLM evals, taught 1:1 by a tutor that scores your cases independently and argues every mismatch until the rubric is unambiguous. Five sessions to a fifty-case golden set and a CI regression gate, ending when your judge matches your labels on forty of fifty.

All Coding & Data templates · All templates