Evals — the measurement scoreboard
Frozen clips from different viewpoints, labelled once, scored on every pipeline change.
A new approach ships when the numbers beat the gates — not when the demo looks good.
Release gates
Eval clips
- Process the clip through the normal pipeline once, then freeze its output: copy
data/outputs/<job>/result.json to the manifest's result path.
- Label it per
docs/benchmarks/PHASE0_BENCH.md (shuttle every ~5th frame + around hits, player box centres, optional pose / 3D shots) into the labels path.
- Score everything:
python scripts/bench/run_bench.py --manifest bench/manifest.json --record.
Run history
GPU spend