ProjectsLive
The Confident Machine
An interactive essay on calibrated trust: why AI sounds sure when it’s wrong, and how to decide when to rely on it, check it, or keep the work.
Open the project (opens in a new tab)theconfidentmachine.com
Hypothesis
People trust AI more carefully after they measure their own overconfidence than after they’re told about it.
Most people meet AI through a chat box that answers everything in the same confident voice. It’s right often enough to earn trust and wrong often enough to do damage, and its mistakes don’t look like mistakes. Telling people this doesn’t work. They nod, then paste the next answer straight into an email.
Method
An interactive essay in seven chapters, about 25 minutes long, in which readers test the ideas on themselves. Every chapter asks them to commit to an answer before it shows the evidence, because being surprised by your own guess teaches more than being told.
- A blind test to open. Two answers to a parent’s question about the Webb telescope: one from Google Bard’s launch demo, which was confidently wrong, and one written by NASA. Readers pick the one they’d trust before learning which is which.
- Two models train in the reader’s browser. A word-prediction model learns from three public-domain books and shows why fluent text isn’t fact-checked text. A hiring screener that never sees group membership still learns an old bias through a zip code.
- A calibration quiz. Ten questions, each with a confidence slider, then the reader’s own calibration curve and score.
- A drag-and-drop Trust Map. Readers place twelve tasks by how easy they are to check and how costly a mistake would be, then compare their map with ours.
- A “My Rules for AI” card, built from the reader’s own choices and downloadable as an image.
How the essay keeps itself honest:
- Every dated claim shows when it was last checked. The stamp turns amber after 120 days, and all the time-sensitive claims live in one file, so they can be re-checked in an afternoon.
- Real sources, no invented examples. Every factual claim links to a primary source, and the opening uses a documented public error rather than a staged one.
- Private by construction. Nothing a reader types or chooses leaves their browser. Both models train locally.
- Three layers for every idea: an analogy, the plain-language explanation, and an optional “Go deeper” panel with the math and the research. Busy readers get the point; curious readers get the proof.
- Directed by me, drafted with an AI assistant. I set the brief and the research and editorial standards and made the calls. The assistant drafted text and code, and every claim was checked against its source. The essay’s colophon says so.
Result
Live at theconfidentmachine.com since September 2026. Readers’ answers never leave their browsers, so their scores stay private by design.
What building it showed:
- Directing AI is an editorial job. The assistant was fast. The value was in the brief, the standards and the checking.
- Honest scope makes a stronger case. The bias demo uses synthetic data and a single proxy. Saying so plainly made the result more credible, not less.
- Staleness is something to design for. The “last checked” stamps turned maintenance into a habit: a claim that can’t be re-checked in an afternoon doesn’t belong on the page.
- The hardest part was deciding what to leave out. Every chapter could have been three.
What I’d change
- Measure what happens afterwards, without breaking privacy. Readers see their own calibration curve, but nothing shows whether it changes how they use AI a week later. That’s the real test of the hypothesis, and it would need an opt-in, anonymous follow-up.
- Test the bias demo on real data. A second, real dataset would test the claim harder than one synthetic example can.
