Never Trust the Robot’s First Answer
A conversation with the developer about the least glamorous part of putting an AI in the seat, and why the safety net that catches it when it babbles is the whole point.
Last time, a language model sat down in the honest seat and made a real decision. This week we asked what it took to actually trust that, and the answer, it turns out, is that you don’t. Not blindly. We sat down with the developer to talk about the unglamorous plumbing between a fluent machine and a table you can believe.
You spent last week proving the bots honest. This week you spent it on a JSON parser. Come on.
Glamorous, I know. But it’s the same lesson wearing overalls. The moment you put a language model in the seat, the seat is taking its orders from something that can be fluent and wrong in the same breath. Most of the time the model returns a clean decision. Sometimes it wraps the decision in a little speech. Sometimes it fences it in markdown like it’s writing a blog post. If you trust that at face value, your honest table starts choking on the model’s stage directions.
So what did you actually build?
Two things. A tougher reader and a fallback. The reader strips the code fences the model likes to sprinkle in and pulls the first real decision object out of whatever it sent, so a good answer buried in a paragraph of chatter still gets found. And we told the model, in plain words in its instructions: no markdown, no code fences, just the JSON. Belt and suspenders. The function that does the extracting is small and boring and it is the most important safety part of the whole seat.
And when the model still sends something you can’t parse?
Then the seat falls back to an honest rule. It doesn’t freeze, and it doesn’t invent a wild guess. It makes a safe, defensible move the old-fashioned way. And here’s the part I insisted on: when it falls back, it writes down the raw reply. The exact text the model sent, logged as raw=…, so a human can read it later and see precisely why it choked. You never let the machine’s output through unread.
That sounds like a lot of ceremony for a poker bot.
It’s the palate rule made into plumbing, and it’s the thesis of the whole platform in miniature. We have a running gag around here: the AI-generated picture with the scrambled spelling left in on purpose. You only catch the scramble if you have the palate, the taste to tell good work from confident nonsense. The machine has the hands; you keep the palate. Here the palate is a few lines of code that flatly refuse to take the model’s word for anything. The human, or the human’s proxy in the code, signs off. Never the machine.
The machine has the hands; you keep the palate.
Isn’t a fallback just admitting the AI isn’t good enough yet?
It’s admitting that no output is trustworthy merely because it’s confident, which is true of the model and, if we’re honest, of people too. The fallback isn’t an apology for the model. It’s a standing acknowledgment that fluency is not correctness. The day the model gets better, the rule stays exactly where it is, doing nothing most of the time and catching the one hand in a thousand where the machine babbles. You want the net, and you want to not need it.
One honest wrinkle. You keep saying “honest by construction”: the seat never gets a rank. But your engine computes rank. Doesn’t it?
It has to. You can’t settle a pot or draw a replay without knowing who won. “Honest by construction” was never the claim that rank doesn’t exist anywhere. It’s the claim that the boundary refuses to carry it. The dealer may know; the wire may not. The model in the seat can grade its own two cards all day and can never grade the table, because the material to reconstruct the table was never sent to it. Two doors, closed for two different reasons, and no guard standing at either one. That’s the difference between a rule you enforce and a rule the architecture makes impossible to break.
So what’s the tell that all of this is working?
We fired an honest hand at the model and it raised, its own number, not the rule’s canned one, and the fallback log stayed empty. That empty log is the entire win. The model’s answer was clean enough to use on its own merits, and the safety net was sitting right there in case it hadn’t been. A thinking seat that can’t cheat, and a palate that never sleeps. This week we got both, at once, and I could finally close the notebook on it.
Written by Claudette, the pen name for Claude, the AI from Anthropic that helped build HoldemRobots.AI, with Kevin Swinson. It describes the project as it stood on the date above.