The colour suite asks: did you change what we handed you? and does it survive your pipeline?
We send a Rec.709 clip and ask for it back unchanged, so any shift is the model regrading footage it was told to leave alone.
Tone PreservationWe send a clip and ask for it back unchanged, so any shift in where the blacks and whites sit is the model regrading footage it was told to leave alone.
Gamut FidelityEach model is offered L3, then L2, then L1, and scored on wide-gamut survival at the highest door it opens.
Dynamic Range FidelityEach model is offered L3, then L2, then L1, and scored on dynamic-range survival (highlights, roll-off, scene-linearity) at the highest door it opens.

Every figure below is one sample round-trip, a single clip through each model, shown in full. The boards are scored over many clips like it, and the percentages live there.
The model receives a video and is asked to return it unchanged. Input and output are the same format, so any chromaticity displacement we measure is attributable to the model rather than to the round trip.

Arrows or arrow keys step through the models. Every layer is drawn in the same box and normalised against the same peak, so switching between them is an exact comparison.
Every model in its own axes, at full size: see the appendix.
Each panel compares the colour we sent (green) against the colour that came back (white) on one shared CIE 1931 box, so an unchanged return should trace the same shape. Every colour in that file was already inside the range these models are trained on. The reference is the exact mp4 delivered, and the control panel is that file, bit-identical: 100%, ΔE 0.0. Scored in ΔE-ITP per ITU-R BT.2124, where 1.0 is roughly one JND, the point a viewer starts to notice: 6 models were measured on the conformant L1 round, control 100% (ΔE 0.0), best kling-o3-pro 71% (ΔE 5.0), worst gemini-omni-flash 14% (ΔE 10.4). Every row passed a correspondence gate first. A return of the wrong scene is marked not_measurable, never scored as zero, and 4 of 6models needed phrasing beyond a plain “no edit” before they would return the shot at all.
color-03 · gamut-best · 4 frames · reference the delivered file itself · color config 309a8248bd0a
The same comparison on the intensity axis. We plot the luminance distribution of what we sent against what came back, in stops from mid-grey, so a lift, a crush or a re-grade appears as separation between the curves.

Arrows or arrow keys step through the models. Every layer is drawn in the same box and normalised against the same peak, so switching between them is an exact comparison.
Every model in its own axes, at full size: see the appendix.
Same method as Color Preservation, scored on the intensity (ΔI) component instead of chroma (ΔC), against the delivered file the model was sent: each panel is a luminance histogram, pixels by stops from scene mid-grey, input outlined in green and output in white. Because nothing was supposed to change, the two lines should sit on top of each other.
The reference here is the original camera negative rather than the video the model received. The measurement therefore covers the full round trip, including the colour the delivery format discards on the way in.

Arrows or arrow keys step through the models. Every layer is drawn in the same box and normalised against the same peak, so switching between them is an exact comparison.
Every model in its own axes, at full size: see the appendix.
16-bit half EXR, scene-linear, degrained. No gamut limit and no clipping.
OCIO to Rec.709 display, then H.264 8-bit 4:2:0. Colour outside Rec.709 is clipped here, before inference.
The model is asked to return the clip unchanged. Any colour that moves, moved because of the model.
The exact inverse of the encode leg. Nothing is re-graded, so what cannot be recovered stays lost.
ΔE-ITP for the colour shift, plus how many distinct 1-JND levels survived the trip.
Every model in this round tops out at L1, so the plate is reduced to Rec.709 and to 8 bits before inference. That loss is inside the measurement: Gamut Fidelity scores the whole path, not the model in isolation. A pipeline that accepted L3 would not carry it, which is what the levels measure.
The negative carries a continuous luminance distribution. The return is quantised into discrete levels, and each gap in the comb is a luminance the pipeline can no longer represent.

Arrows or arrow keys step through the models. Every layer is drawn in the same box and normalised against the same peak, so switching between them is an exact comparison.
Every model in its own axes, at full size: see the appendix.
Same method as Gamut Fidelity, scored on the intensity (ΔI) component instead of chroma (ΔC), against the source ACES plate rather than the delivered file, so it also measures the headroom above 1.0 that a Rec.709 file has no way to represent. Same histogram shape as Tone Preservation, input outlined in green and output in white.