Skip to content
Experiment

A small evaluation harness

Notes from building a harness that scores retrieval, reasoning and refusals separately, on fifty real cases instead of five thousand invented ones.

What it is

A script and a spreadsheet: labelled cases go in, three scores come out. No dashboard, no service, nothing that needs to run in production. It exists to answer one question before a system ships: does it actually work on the cases that matter, and which one broke first.

How it works

Each case in the spreadsheet carries three things: the passages the answer should come from, the answer itself, and whether a refusal is the correct outcome. The harness runs the system under test against every case, compares the result to the expected one, and writes a row per case rather than a single average. A drop in the score can be traced back to the exact case that caused it, not just to “something got worse this week”.

Limits

It does not judge free-text answers on its own; a person still labels the disagreements between what the model said and what the case expected, and that label becomes part of the record. It has only run against documents and support tickets so far, not against time series or sensor readings, so treat the retrieval half as unproven outside text.

Network of connections spreading from Valencia across the globe

Want this on your *data*?

Everything in the lab is public. If one of these should run inside your environment, on your data, talk to us.

Talk to usor write to hello@maistik.studio