kullback is the open-source harness behind Leibler.

Your traces in.
A verified environment out.

Traces are the logs your agent already writes. Kullback reads them, rebuilds the tools, data and rules the agent used, replays the logs to prove the copy is faithful, then scores any model inside that copy on what it changed.

new traces, fixes, disputes tracesyour runs buildtools, state, checks runany model verdictcode, end state reportyou decide

The loop. Code decides pass or fail from the final data, not from the transcript. A judge model can take a pass away, never give one.

What a week of traces turns into.

Every value points back to the line in your trace it came from. Nothing is invented.

Checked by replay before any model is scored.

Measured on Sierra's public retail traces (tau2), where the real tools and database exist to compare against. Seen: runs used for the build. Held out: runs the build never saw.

checkseenheld out
tool signatures0/15
starting rows0/252
writes0/370/20
reads0/1250/65
errors0/60/4

Runs where a model had to stand in for a missing tool are shown in the report, never counted as passes.

All of it is public.

The rebuild is measured.
Scoring models in it is next.

Generated tool code, verifiers on real tasks, and runs of candidate models. The numbers get published either way.