Start with a case study
40% Better, 75% Faster describes the query sets, annotations, and performance snapshots used in the Frigade engagement.
Series
Essays and case studies on reviewing AI outputs, identifying failures, and testing improvements.
If you built the agent, why are you extracting its task from chat traces instead of recording typed inputs and outputs at the source?
A streamlined system for AI evaluation that closes the gap between seeing problems and fixing them.
How Frigade Slashed Latency & Boosted User Helpfulness
An AI Maturity Model
A Practical Guide to Evaluation-Driven Improvement
40% Better, 75% Faster describes the query sets, annotations, and performance snapshots used in the Frigade engagement.
Use the AI Evals topic hub for more writing on evaluation.