Topic hub

AI Evals

Writing about reviewing AI outputs, measuring failures, and testing whether a change improves the system.

Start here

  1. 40% Better, 75% Faster

    A case study of query sets, trace annotations, and retrieval and latency changes at Frigade.

Core concepts

Failure taxonomy

The map of what breaks, why it breaks, and which failures are worth preventing.

Review UI

The place humans inspect traces, label failures, and create the ground truth for improvement.

Judge calibration

Comparing a model judge with human labels on held-out examples, including the mistakes it misses.

Regression sets

Examples of past failures used to check whether a later change reintroduces them.

© 2026 Skylar PayneFieldwork