What it does
Runs the three jobs at the end of a video ad pipeline: checking whether a produced clip matches what was specified, structuring the test so its result can actually be read, and reading the result honestly afterwards. It never scores a dimension it has not measured. It writes the exact commands that produce the technical facts and the extracted frames, then judges from what comes back — and where something cannot be judged from frames at all, such as lip-sync or the audio mix, it returns not-reviewed with the reason rather than a plausible number. An invented score passes a clip nobody checked. Its test setup moves ONE variable per cell, because a winner across two differences is a winner nobody can learn from, and it refuses a readout on a test that has not accumulated enough data to be read. A result below readability is not a weak signal; it is noise, and noise acted on is worse than no test. It closes the loop. The readout produces the feedback that goes back to the angle, hook and script stages, so the next cycle starts from a result instead of from guesswork.
Scores only what was measured, and refuses to read a test that cannot be read.
Example prompts
- Write me the QA measurement commands for these clips.
- Here is the ffprobe output and the frame descriptions — score this clip.
- Set up the test so each cell moves one variable.