Failure taxonomy
The map of what breaks, why it breaks, and which failures are worth preventing.
Topic hub
Writing about reviewing AI outputs, measuring failures, and testing whether a change improves the system.
A case study of query sets, trace annotations, and retrieval and latency changes at Frigade.
The map of what breaks, why it breaks, and which failures are worth preventing.
The place humans inspect traces, label failures, and create the ground truth for improvement.
Comparing a model judge with human labels on held-out examples, including the mistakes it misses.
Examples of past failures used to check whether a later change reintroduces them.
A streamlined system for AI evaluation that closes the gap between seeing problems and fixing them.
How Frigade Slashed Latency & Boosted User Helpfulness
An AI Maturity Model
A Practical Guide to Evaluation-Driven Improvement