Connect with us

Hi, what are you looking for?

Tech

The Hidden AI Problem: What Happens When We Stop Checking What AI Tells Us

The Hidden AI Problem: What Happens When We Stop Checking What AI Tells Us
The Hidden AI Problem: What Happens When We Stop Checking What AI Tells Us

In 2023, a pair of lawyers filed a court brief with citations in it. Nothing unusual there, except that six of the cited cases did not exist. They’d been generated by a chatbot, complete with names, court names, and legal reasoning, all of it invented. The lawyers said they hadn’t realized the tool could do that. A judge sanctioned them. The story made headlines everywhere, mostly because it was so easy to laugh at.

A year later, a search giant’s new AI-powered answer feature told millions of people, with complete composure, that you can add glue to pizza to keep the cheese from sliding off, and that eating at least one small rock a day is good for you. The examples were absurd enough to go viral. The company quietly fixed the worst ones.

Both of these got a lot of attention. And both of them, I think, told us the wrong lesson.

The obvious lesson is “AI makes mistakes.” Fine. We know that now. The hidden lesson — the one that didn’t make headlines, and the one this article is actually about — is that in both cases, a human was supposed to be in the loop, and in both cases, nobody checked. The lawyers checked nothing. The people reading the rock advice mostly laughed and scrolled. The failure that matters isn’t the hallucination. It’s the fact that we keep building systems that confidently produce wrong things into a world where checking has quietly become optional.

The part that gets reported

Let’s give the machines their due before we talk about us. AI systems do make things up. It has a technical name — hallucination — and it’s a structural feature of how these models work, not a bug that’s being worked out. A system trained to predict the next most plausible word has no built-in mechanism for distinguishing plausible from true. It produces both with the same effortless confidence. Sometimes the words are right. Sometimes they’re smooth fiction.

There are well-documented cases where this has caused real harm, not just embarrassment. An Australian mayor publicly called out a chatbot that had confidently claimed he was a convicted criminal. A radio host sued over a fabricated endorsement. The list grows every few months. None of this is a secret. The industry acknowledges the problem, researchers study it, and the public has largely absorbed the idea that AI “can be wrong.”

And yet, somehow, the knowledge doesn’t translate into behavior. That’s the gap that deserves our attention.

The part that doesn’t get reported

Here’s the uncomfortable question: when did you last check something an AI told you?

Not in the sense of noticing it was wrong. In the sense of actively verifying — opening a second source, running a search, doing the work of confirmation before you accepted the answer and moved on. For most people, the honest answer is “rarely.” And that’s not a personal failing. It’s a well-documented human flaw that we’ve been studying for decades, long before anyone heard of ChatGPT.

It’s called automation bias. Researchers in aviation and medicine identified it years ago: when people have a reliable automated system in front of them, they stop monitoring it. They assume it’s right. They go along with it even when their own senses suggest otherwise. Pilots have flown planes into terrain while trusting an automated system that was quietly turned off or misconfigured. Doctors have overridden their own clinical judgment because a machine suggested a different answer. The pattern is stubborn, and it doesn’t respond to being told “automation makes mistakes.” Everybody already knows that. They do it anyway.

The reason is simple: checking costs more than not checking. It takes time, attention, and effort. An AI answer arrives with zero friction, perfectly formatted, in the tone of a confident expert. Double-checking it means leaving the pleasant flow of work and doing the boring labor of verification. The human brain, which is an inveterate cost-cutter, declines the expense. Most of the time it’s right to decline. The answer was correct. It’s just that one time in twenty it isn’t, and by then the habit is already set.

Why smoothness feels like truth

There’s another layer underneath, and it’s almost unfair. Psychologists have shown that people judge information partly by how fluent it feels — how easily it moves through the mind. Smooth, well-structured, confidently worded content feels truer than hesitant, messy content, regardless of the actual facts. It’s called the fluency effect, and AI is basically a fluency machine.

Think about the difference between an AI answer and a genuinely researched one. The AI answer is clean. It arrives in tidy paragraphs with no hedging, no half-finished thoughts, no visible uncertainty. A human expert, asked the same question, will often pause, qualify, offer caveats, say “it depends.” That hesitation is honest. It’s also, to the fluency-judging brain, evidence of weakness. The machine’s unbroken confidence is, to the same brain, evidence of strength.

So the technology isn’t just making mistakes. It’s producing mistakes that are optimally packaged to be believed. And that’s a genuinely new situation. Human error was usually accompanied by human tells — a shaky memory, a sloppy sentence, a “I think this is right but I’m not sure.” Machine error arrives wearing the exact same confident face as machine correctness. There’s no tell. The system itself can’t distinguish the two, and neither can the interface. Which means the only filter left is the human, and the human, for all the reasons above, has stopped filtering.

The vanishing reference point

There’s one more reason this is getting worse rather than better, and it’s the one I find most unsettling. Checking requires something to check against.

When most content was written by humans, there was a background of “normal text” you could compare against. Something generated by a machine stood out, at least a little, by its artificial smoothness. That’s fading fast. AI-generated text is now a large share of what people read — the news aggregations, the summaries, the SEO content, the answer snippets at the top of your searches. The background is becoming machine text too. There’s a real sense in which AI summaries flatten the messy, contradictory, lived detail of human conversation into clean paragraphs — and the clean paragraph is exactly what feels most believable.

When everything reads like it was written by the same smooth, confident machine, your internal “this looks fake” radar stops having a baseline to work from. The signal gets lost in the noise of its own copies. And the misinformation worry from the first article in this series becomes concrete here: it’s not that one big fake fools you. It’s that the ordinary stuff you read every day stops being a trustworthy reference point for judging the rest.

When the answer becomes an action

Now take all of the above and multiply it by the direction AI is moving.

The chatbots people got used to answer. The systems being built now, and increasingly marketed as the main event in 2026, act. They don’t just tell you what to do — they do it. The industry has shifted from tools that generate text to systems that run whole workflows on their own, and even chatbots have moved well beyond words into doing things. An agent that books the flight, files the form, answers the email, or moves the money isn’t presenting you with a claim you can check. It’s presenting you with a done.

This changes the geometry of the problem completely. Checking an answer is optional in a way that feels safe, because nothing has happened yet. Checking an action means checking it after the fact, when the flight is booked, the form is sent, the money has moved. The cost of verification doesn’t just go up. The window to do it at all — while there’s still something to prevent — shrinks to nothing. Over-reliance stops being a matter of believing the wrong fact. It becomes a matter of letting the wrong thing get done, to you or on your behalf, before anyone notices it was wrong.

That’s the escalation hiding inside the original, harmless-looking problem. You don’t start out trusting the machine. You start out not checking the easy answers, because the easy answers don’t matter. Then the easy answers quietly become decisions, and the decisions become actions, and the whole staircase has been walked by a person who never once said “let me verify this” out loud.

What checking actually requires

I’m not going to tell you to verify everything. Nobody has time for that, and it would defeat the point of the tools. But there’s a way to keep yourself in the loop without turning every prompt into a research project, and it comes down to a small set of habits.

Check the things that would hurt if wrong. This is the whole system in one line. An email rewrite doesn’t need verification. A medical question, a financial decision, a legal statement, a fact you’re about to put in front of other people — those do. Sort every request by the cost of a wrong answer, and only spend verification budget on the expensive ones.

Ask the model to show its work. Not in the sense of trusting its reasoning, but in the sense of asking for sources, dates, and concrete anchors you can actually chase down. A model that can’t point to anything specific is telling you something.

Look for the tell. It’s usually confidence. The most dangerous AI output is the one that hedges nothing. Genuine knowledge, in humans and increasingly in machines, tends to come with qualifications. Suspiciously clean answers deserve a second look, especially on subjects where real experts are usually more careful.

After the fact counts as checking too. If you let an agent take an action on your behalf, review the outcome. This sounds trivial. It’s not — it’s the only remaining point in the loop where the human still has leverage, and it’s the one we skip most often.

None of this is about distrusting AI. It’s about noticing that the real risk was never the machine being wrong. The machine will be wrong sometimes; that’s a solved problem in the sense that we’ve always known it. The risk is the machine being wrong and nobody noticing, because the whole system — the confidence, the fluency, the volume, the shift to autonomous action — is optimized to make checking feel unnecessary. The story of the lawyers and the fake cases isn’t a story about AI. It’s a story about what happens to human judgment when a persuasive machine makes it easy to stop using it.

The series continues from here. The opening article mapped out why people worry about AI at all, and the trust paradox explained why they keep using it anyway. This piece is the part in between — what happens to the human habit of verification once the using starts. Next up: the jobs question, the privacy question, the misinformation question, and the control question, each examined on its own terms.

You May Also Like

Blog

What follows is an illustrative reconstruction — two fictional employees at one fictional company, told in alternating scenes over eighteen months. The characters are...

Tech

Here’s a scene I keep running into, in various forms, since around 2023. Someone usually someone perfectly sensible — is staring at a chatbot...

Tech

At dinner last week, the subject came up the way it always comes up now. Someone’s bank froze their card because a fraud model...

Blog

Course overview Enrollment open. No prerequisites required. 2026 is the first year “AI automation” stopped being an engineering specialty. The things that previously needed...