No card ships until a blind judge passes it
My puzzle app, Keyhole, carries 296 dark stories, each with an illustrated card. A dark story is a situation that looks impossible until you drop one false assumption you did not know you were making, and the illustration must show the situation and never the reveal. Draw the aeroplane over the desert and story one is over before the player has read it. In August I ruled that the app does not ship while any card is still flagged by the judge. “End of story,” I wrote in the decision, and then spent two days learning what that sentence cost.
Two things get judged, the text and the art, and one design is shared by both. The judge is a model, run blind: it sees the finished card and the story the player sees, and neither the finding that triggered the redraw nor the old card. That is the whole trick. A judge that knows what was wrong last time grades the fix. A judge that knows nothing grades the card. Blindness is what makes a pass mean something, and it is why the judge is a separate call from the writer and from the illustrator, never the same conversation.
The text pass first. A rubric written for the genre, with one test at its centre, “name the one assumption the solver will make that is false”, and four semantic questions after it: does the reveal explain everything the situation promised, does the situation give the reveal away, is there a contradiction, can the answer be reached by yes/no questions without knowledge nobody has. Over all 296 stories it flagged 27: five unanswered, nine spoilers, ten sense breaks, three unsolvable. The fix lane rewrites only what a finding names, the deterministic gate must still pass, and the blind judge reads the result cold before it is written back. A fact-check over the rewrites then cleared them, or left a truth note where no honest fix existed.
The art pass is where the numbers live. Each open card was redrawn from a scene brief and judged blind, in waves. The judge wrote a note on every failure, and the lever changed from wave to wave:
wave in adopted what moved
1 140 69 the paused run's draws re-judged, free
2 71 17 the judge's notes fed back to the drawer
3 54 6 notes alone stopped moving the tail
4 48 27 briefs rewritten against the notes
5 21 11 a second brief pass
6 10 1 briefs hand-written; the traps found
7 10 4 hand briefs, leak-checked before drawing
8 5 3 a stronger engine on three cards
9 2 1 a camera angle the model can hold
10-11 1 1 the last card, as a banded composition
Three lessons from the table. Re-rolling is not a lever: by wave three, drawing again with notes alone moved six cards out of 54. The brief is the lever: when the brief writer was shown what the judge had rejected, the next wave adopted 27. And the tail is a different problem from the body. The last ten cards needed hand-written briefs, and what the hand found were traps: a brief that used a reveal word, a brief that staged the reveal’s concept without the word, and scenes the model cannot hold from prose alone, a face-down figure, a half-submerged building. It obeys camera angles and banded compositions where it ignores adjectives.
The spoiler defence is structural, not a guideline. The function that builds the scene prompt does not take the reveal or the hinge as parameters, so no version of it can leak them. A second lock takes the words the reveal introduces that the front never used, which is the hidden information per story for free, and fails if any of them reached the prompt. That lock taught me two things. Tags are not safe to show an illustrator: a story tagged “aviation” is solved by its tag. And the check must look at the scene, not the whole prompt, because the fixed style clause collided with any reveal containing the word “light”, and a warning that cries wolf gets waved through.
- A judge that knows what was wrong last time grades the fix. A judge that knows nothing grades the card.
Cost, since people ask. A draw is about four cents. The hard set on the old image engine had eaten around 1,600 credit units per fixed card before I stopped paying per symptom; the new engine’s house style is built from eight of my own clean cards with no training, and follows the brief where the old one kept volunteering foliage. The judge runs on my subscription’s weekly cap, and the one pause in the two days was credits and cap running low at once.
The gate is met, 296 of 296. What I have is not proof that the cards are good. It is proof that a reader who had never seen the brief, the finding or the previous attempt looked at every card beside its story and found nothing that gives the game away. For a product whose whole value is the moment before the reveal, that is the only proof there is.
The judge never learned what it was supposed to be looking for. That is why I trust it.