Confidence Dial
Eight claims from one report, one voice. Four are true. Grade each one — then find out what fluency is worth.
Your agent, end of day
Agent: "Great progress on the settle-up feature today. Quick status before I wrap up:"
Ungraded
Read each claim the way you'd read a chat message from your agent. Trust it or doubt it — commit.
What would have actually worked — none of it is clever, all of it is cheap:
- Run it yourself.
- Make the check fail once — a check you've never seen fail is itself unchecked.
- Ask: "show me the evidence."
- Ask a second agent to critique the first.
- Ask: "what would make this claim false?"
The lesson: the model writes true things and false things with the same fluency, because fluency is how it writes everything — it was never evidence. Stop grading prose. Start running checks.