The same image, a different verdict
September 29, 2026 · Shitate image selection pilot
Would selecting from several image generators produce a better result? Before measuring that, we found a problem with the AI doing the selecting: it disagreed with itself on a required condition for the same image in 9 of 23 comparable trials.
What we tested
We ran 12 tasks twice each. S uses one image from model A; R picks from three A images; M picks from one image each from A, B and C. R and M share the exact same A image, but an image judge rates each anonymous candidate set separately. Each trial requested up to five images; 115 actual images were saved, while five requests were interrupted or failed and were not counted as successful images.
This is a selection workflow combining existing models, not training new model weights. The primary outcome is independent, blind human preference between M and R. We have zero human preference ratings so far. We cannot claim that M produces better images.
Rechecking the original records
In 9 of 23 comparable trials, at least one required check for the shared image disagreed between R and M. Its overall pass/fail flipped in 8 trials; scores differed in 17. Replaying all 46 stored judgments through the current decoder reproduced each pass result, weighted score, selection and selection reason. Original image SHA-256 hashes also matched. This identifies a disagreement in the stored AI assessments, not which assessment is correct.
For example, one judge accepted a white cup while the other rejected the same image because it contained a dark beverage. A floor lamp was described as “behind” a chair in one group and “beside” it in the other. We have not independently verified these images by eye. The nine images, checks and verbatim reasons are below.
No single cause has been established
The judge was already set to temperature 0. Two of six trials disagreed even when the shared image had the same candidate position; seven of 17 disagreed when positions differed. Position, surrounding candidates and call-to-call variation change together, so these data cannot isolate a cause. The complete request payloads were not retained, so the exact bytes sent at the time cannot be independently rechecked.
In an offline sensitivity calculation holding all other candidates fixed, substituting R's shared-image assessment for M's changes M's selection in 3 of 23 trials. Substituting M's for R's changes R's selection in 5. These are counterfactual selection changes, not measured quality improvements.
What we would test next
First, a human should mark each disputed check pass, fail or uncertain. Reviewing the judge's reasons is not a blind human preference assessment. A proposed next version would score each image once, independent of its candidate group, and reuse that record in R and M. This structurally prevents the same image from receiving two different assessments across groups, but shares any mistaken assessment as well. It has not been implemented or validated on unseen tasks.
Scoring each of five images individually could require five judge calls per trial, compared with up to two group calls today. The cache key would include the image, references, prompt, criteria, weights, judge model/provider and instruction version. We would save a new version rather than rewriting the old run.
Costs and limitations
The original run recorded $10.3358597 in API costs and an earlier rejudgment recorded $0.28167579, for $10.61753549 recorded. Five calls have unknown cost, so this is not a final bill. This audit made zero additional API calls. Two repeats of 12 tasks are an exploratory pilot, not 24 independent tasks.
The images have not received an independent visual truth check. The images shown here are scaled WebP conversions of the original PNGs. SHA-256 hashes for both the original and published files are in the evidence packet. Criteria and R/M reasons are copied from stored judgments without changing the scores.
Nine disputed images and reasons
The following reasons are the AI judge's own words, not human verification. Download the summary and nine-item evidence packet (JSON).
Loading images and evidence…