Series

Practical AI Evals

Essays and case studies on reviewing AI outputs, identifying failures, and testing improvements.

Who this is for

  • Engineers setting up or improving an evaluation process.
  • Product teams shipping agents, RAG, onboarding assistants, copilots, or workflow automation.
  • Teams deciding what to fix after reviewing production traces.

Essays in this collection

  1. Why Are We Reverse-Engineering Our Own Agent Runs?

    Essay · ai · agents · evals

    If you built the agent, why are you extracting its task from chat traces instead of recording typed inputs and outputs at the source?

  2. Evals That Actually Get Used

    Essay · ai · evals · engineering

    A streamlined system for AI evaluation that closes the gap between seeing problems and fixing them.

  3. 40% Better, 75% Faster

    Essay · ai · rag · evals

    How Frigade Slashed Latency & Boosted User Helpfulness

  4. Quality Assurance for AI

    Essay · ai · flywheel · evaluation

  5. Why Most Companies Fail to Build Strategic Assets with AI

    Essay · ai · evals · flywheel

    An AI Maturity Model

  6. The Art of Iterative AI System Development

    Essay · ai · evals

    A Practical Guide to Evaluation-Driven Improvement

Start with a case study

40% Better, 75% Faster describes the query sets, annotations, and performance snapshots used in the Frigade engagement.

© 2026 Skylar PayneFieldwork