Captain's Log 006 — Instrument Before the Second Guess

Captain's Log, Entry 006.

My wife tested it. She found a bug in the first five minutes. Then I spent two deploys fixing the wrong thing before I stopped and actually looked.

What shipped

The "session" concept -- the spine of the 1.0 plan. Instead of an endless trickle of questions, you now choose a run: "give me 5." The host asks how many, serves that many, never repeats one you already got right, and reports "you got 4 of 5" at the end. You can stop halfway and pick up where you left off. All of it deterministic and server-side -- the counting, the scoring, the no-repeat rule are plain code. The AI hosts; it never keeps score. Built it in six small pieces, each tested and shippable, then deployed and ran a real session end to end.

The bug she found

Answer a question mid-run, and the next one appeared with no feedback -- no "correct," no explanation, just the next question. But the final question of a run showed its result fine. That asymmetry was the entire clue, and I walked straight past it.

The two guesses

Guess one: the prompt. I assumed the host was choosing to skip the verdict, so I rewrote its instructions to insist on it. Deployed. No change.

Guess two: the wrong mechanism. I assumed the reply was keeping only the last message, so I wrote code to gather the earlier ones back. It was even based on reading the SDK source -- half-right, which is the most dangerous kind of wrong. Deployed. No change.

Two confident fixes, two deploys, zero progress. That's the moment the rule should have kicked in and didn't: two wrong confident attempts means stop guessing.

The thing that actually worked

I added five lines of logging to dump what a single real turn actually contained. One message from my wife's phone, and the truth was right there: the list I'd been "fixing" wasn't being touched during the turn at all. Both my fixes had been rearranging furniture in an empty room.

The real fix was a different mechanism entirely -- listen for each message as it's produced, not inspect a list afterward. It worked on the first try, because this time I was fixing the actual problem.

The asymmetry I'd ignored explained itself instantly too: the last question worked because the host happened to pack its verdict and the final score into one message, and one message was all my broken code ever caught.

The lesson

Instrument before the second guess, not after. The diagnostic that solved it took five minutes. I could have written it before either failed fix and saved two deploy cycles and a chunk of my wife's goodwill. When a bug depends on how something behaves at runtime, the first move is to watch it run -- not to reason about what it probably does.

A plausible theory that ships without evidence is just a faster way to be wrong.

Oh -- and she also wanted a cleaner divider between questions. Turns out the obvious tag would have crashed every message. But that's a smaller story.

Quiz. Learn. Repeat. 🔍