Assuring AI quality takes more than testing
From QA to EvalOps: how we document that an AI system solves the right task at the required quality, including after it goes into production.

Summary
An AI solution is not good because it works in a demo. It is good when we can document that it solves the right task at the required quality, including after it goes into production.
A classic software test has a single correct answer: given a specific input, we expect a specific output. Generative AI has no single correct answer. Two different responses can both be good, and a response can sound convincing and still be wrong. A green test suite therefore does not tell you whether the AI system solves the task.
EvalOps makes quality measurable and repeatable. Success is defined from the actual workflow and tested on a dataset that resembles reality. Evaluation runs automatically on every change, blocks a deployment on regression and continues after the system is in production.
Contents
- 1Why classic QA falls short
- 2From gut feeling to measurable quality
- 3Start by defining what success means
- 4Build an evaluation dataset that resembles reality
- 5Measure both the end result and the components
- 6Automate evaluation, and calibrate it against humans
- 7Make evals part of the development process
- 8Evaluation does not stop at go-live
- 9The most important eval may lie outside the AI system
- 10EvalOps and the AI Act
- 11EvalOps does not replace security and classic QA
- 12From AI project to production discipline
Written for
For those who build or are responsible for an AI system that is in production or about to go into production: AI teams working with RAG solutions and agents, and the people who own the workflow the system supports. It is particularly relevant for organisations in regulated casework, where quality must be documented to supervisory authorities and appeals bodies.