How to evaluate an AI system before you trust it.
Trust is a number you measured, not a feeling you had in the demo.
Why retrieval quality, not the model, decides whether your system tells the truth.
| Method | Retrieval and generation scored separately, offline then online |
|---|---|
| Dataset | Real support tickets and search logs, labelled with the client team |
| Metric | Passage recall, answer correctness and correct refusals |
| Cadence | Weekly, on a fixed question set |
Retrieval quality, not the model, decides whether a RAG system tells the truth. We score retrieval and generation apart, on real questions, on a schedule, and treat a correct refusal as a result.
Everyone tests the model. Almost nobody tests the retrieval step with the same rigor, and retrieval is where most of the wrong answers come from. So we separate two questions that get lumped into one: did we retrieve the right passages, and did we answer correctly given what we retrieved.
We build the evaluation set from real questions, not invented ones: support tickets, search logs, the questions your team already gets asked. We label the passages that should come back for each question, so retrieval has its own score, separate from generation.
A RAG system can pass every spot check a person runs by hand, because a person tends to ask questions the system is good at, on documents the system already indexed well. The gaps show up on the questions nobody thought to ask.
Once retrieval and generation are measured separately, the fixes get cheaper and faster to find. A chunking problem stops looking like a model problem. A stale index stops looking like a hallucination you cannot explain.
We also test what happens when the answer is not in the documents. A system that guesses instead of saying “I don’t have that” is worse than no system, because it hides its own failure behind good grammar. Refusing correctly is a result we measure, not an afterthought.
None of this removes the need for judgment. Someone still has to decide what “correct” means for your documents, your customers and your risk tolerance. We build the evaluation with your team, not for it, because that judgment does not transfer well from one company to another.
We run the evaluation on a schedule, not once at launch. New documents arrive, old ones get retired, and a phrasing that used to retrieve the right passage can quietly stop working. Treating evaluation as a one-time gate is how a system that passed every test in March starts giving wrong answers by autumn, with nobody able to say when it began.

Tell us what you are building. We reply with a plain answer on whether AI is the right tool.
Talk to usor write to hello@maistik.studio