skip to main content

AI Will Always Make Mistakes

AI gives imprecise, ballpark answers by design

Most people are familiar with Google’s AI rollout that led to AI Overviews suggesting that users eat rocks and glue the toppings onto their pizza. It’s funny to laugh at outdated AI models, but it’s a lot less funny when your kid flunks their term paper, or you get fired for citing fake research at your job. With billions and billions of dollars being poured into AI companies, will they ever be able to eliminate these catastrophic errors?

The answer is no. The mathematical constraints of large language models (LLMs) mean that they will always produce plausible falsehoods, but that doesn’t mean that they’re completely useless. We’re going to dive into the way that AI chatbots generate their responses and how you can use them safely and responsibly.

Imprecise by design

Generative AI can be a powerful tool, but at its core it is probabilistic. An LLM doesn’t actually know anything except how likely it is for certain words to appear in a certain order. It’s basically a fancy version of the predictive text on your phone.

Instead of just guessing the next word based on the previous one, it tries to look at several paragraphs worth of text (or more) all at once and then generate several paragraphs of new text to go along with it. This often happens behind the scenes, as in the case of chatbots. While it might look like a back-and-forth conversation, under the hood it’s just copying the conversation into one big document and running the same autocomplete function on it.

In a 2025 study by OpenAI researchers Adam Tauman Kalai, Edwin Zhang, and Ofir Nachum as well as Santosh S. Vempala from Georgia Tech, the authors demonstrate that there is a lower bound for the error rate that even with perfectly accurate training data never reaches zero. In other words, no matter how good the data you feed into an AI is, it will still give you the wrong answer to a question some percentage of the time. And if you have less than perfect training data (like Reddit posts, for example), the error rate could be much higher.

Of course, we should be used to technology making mistakes all the time. Returning to the example of texting on your phone, you can occasionally make a generic text response by just clicking on the autocomplete suggestions, but they rarely work for specific questions. While just clicking on autocomplete suggestions to see where it takes you is good for a laugh, we wouldn’t trust it to send a detailed response. Likewise, generative AI does a pretty good job of piecing together generic answers for well-known questions, but starts to fall apart when your questions get more specific or obscure.

Good at generalities, bad at specifics

Since an LLM’s probabilities for how to string words together is based on its training data, you’re much more likely to get a reasonable response about a topic that has been widely written about online because the words from those sites have probably been scraped and used for training data. If you try to ask specific questions about a more narrow topic, however, you will start to see more errors pop up in the response.

Let’s demonstrate this by asking Google’s Gemini about a topic I would safely consider myself an expert on: the website HighSpeedInternet.com.

Okay, so far so good. It’s the sort of generic information that you could have just found by going to our About Us page, but at least it’s accurate. Let’s try to get more specific.

Again, it sounds a bit like Gemini is just paraphrasing our own About page, but nothing too egregious. And if we scroll down …

… there’s the bit about me. But here Gemini has made a mistake while trying to paraphrase my bio. I did get a shoutout from Wired back in the day, but it was for my writing in a videogame, not my commentary on infrastructure. Let’s poke at this a little more.

“Highly regarded in the tech journalism space.” Oh, Gemini, you flatter me (which is a whole issue unto itself), but let’s scroll down a bit.

Again, Gemini is pretending not to plagiarize my bio by playing Mad Libs with tech journalism buzzwords, but it’s still making the same mistake. Let’s get straight to the point.

Screenshot of a Gemini conversation about writer Peter Christiansen.

Gemini didn’t answer my question, but it did double down and say that my tech research was “directly praised and referenced by Wired.” That’s a much more bold and more easily falsifiable statement. Also, what does it even mean by “consumer telecom boundaries?” Let’s prod it a bit harder.

Screenshot of a Gemini conversation about writer Peter Christiansen.

There we go. Gemini has finally acknowledged that one of its previous statements was incorrect. Of course, it then invented a new claim that is equally dubious, but you get the idea. Also, blaming my bio for saying something it didn’t say? Classy, Gemini…

Fortunately, no one has to write a school report on my life, but you can see how seamlessly false claims can pop up in AI responses. If I weren’t the one reading the response, it probably would have gone by unnoticed. In an educational setting, these sorts of errors could ruin your grades. In a professional setting, they could ruin your career.

AI doesn’t hallucinate

The plausible falsehoods regularly generated by LLMs are commonly referred to as “hallucinations,” which is an incredible bit of marketing for AI companies, but a terrible term for understanding how AIs work. A hallucination implies a problem in the normal functioning of the senses that creates a false perception that doesn’t reflect reality. In contrast, AI hallucinations are part of the normal functioning of an LLM. They are not an example of the model experiencing a glitch; they are examples of the model doing exactly what it’s supposed to do.

LLMs take a bunch of text, convert it to a string of numbers (known as tokens), and then determine what numbers should come next in the sequence, based on the model’s own internal weights. These weights are set during the training and evaluation period, where models will be rewarded for giving good responses and penalized for giving poor responses. The idea is that a false statement would be penalized, but as the OpenAI researchers note in their paper, as with many tests, confident guesses are often rewarded, while admitting you don’t know is penalized.

As such, an LLM doesn’t care if the answer it gives you is completely false. It will even admit a statement is false, and then make the same claim again, as we saw in the previous example. The only thing it cares about is generating a response that fits the pattern it was trained to follow.

The technical term for this sort of disregard for the truth, as theorized by the moral philosopher Harry Frankfurt, is “bullshit.” Bullshit is different from lies or deceit, because a liar knows the truth and is trying to hide it. To someone spouting bullshit, it doesn’t really matter if their statements are true or false. The purpose of bullshit is just to convince or impress its audience. It’s like the slick used car salesman who will say anything to get a sale, or the charming rogue who can sweet-talk himself out of any situation. Thus, in the philosophical sense, AI is bullshit.

Importantly, AI isn’t just bullshit when it gets things wrong. It’s bullshit all the time. The only difference is that sometimes the resulting output is useful to its users and sometimes it is not. It’s all the same to the AI.

Looking for a way to stay safe while online?

Using a VPN masks your IP address and safeguards your online privacy. Find out how to get started with a VPN from our tech experts.

How to use an LLM safely and effectively

On Sept. 9, 2026, the New Mexico Supreme Court held a defense attorney in contempt for submitting a brief that contained false testimony and fabricated witnesses. The brief in question was generated by feeding the facts of the case into ChatGPT. The lawyer claimed that he had no idea AI tools could generate false information. That argument didn’t seem to convince the judge.

If LLMs and chatbots are mathematically guaranteed to give false information, should we be using them at all? Well, we certainly shouldn’t be using them as if they were actual experts on a topic, but that doesn’t mean that they don’t have their uses.

As their name states, large language models are, first and foremost, language models. They are amazing tools for interacting with software through natural language processing. In fact, the suggestions in the OpenAI researchers’ paper include creating specific tooling for AI applications (like chatbots) to handle simple math problems and provide accurate answers to a list of predefined questions. This effectively bypasses the mathematical bounds for accuracy inherent in base models and makes it possible for an AI application to give only correct answers—at least to certain questions.

It’s important to note that the actual answers in this case are coming from the tooling, not the base model. While we generally treat these tools as mere add-ons to the LLM, the language model is primarily serving as an interface between the user and the software that’s answering their questions. This is where this kind of AI really excels.

Of course, if you’re just using a general-purpose AI application like Gemini or ChatGPT, they can still be useful. You just can’t take anything they say at face value.

Anyone who’s internet savvy has experience doing this with plain old Google Search (before it got turned into another chatbot). Your average Google Search will turn up hundreds of thousands of results, but only a few will be relevant to what you actually want. In fact, while Google’s algorithm does a pretty good job of bringing quality results to the top, that doesn’t mean you always want to click on the first link.

An important part of internet literacy is understanding how to interpret information like search results. A page of search results might list a reputable online business and a virus-laden scam site right next to each other, but internet users have developed the skills to recognize a trustworthy site.

Similarly, AI responses might contain both true and false information, but it makes disentangling them much more difficult. Users need to be able to think critically about these responses and determine which information is trustworthy. Don’t think of a chatbot as a learned professor explaining established facts. Think of it as a used car salesman reading you the results of a Google search on a topic he knows nothing about.

LLMs and AI chatbots have their uses, but when you have a hammer, it’s tempting to treat everything as if it were a nail. If you use AI, use it where it actually improves your productivity. Understand how often responses include false information. If the AI makes an interesting point, click through to the original source to verify. And whatever you do, don’t ever turn in AI-generated work without double-checking it. Especially if you’re a lawyer.

Author -

Peter Christiansen writes about telecom policy, communications infrastructure, satellite internet, and rural connectivity for HighSpeedInternet.com. Peter holds a PhD in communication from the University of Utah and has been working in tech for over 15 years as a computer programmer, game developer, filmmaker, and writer. His writing has been praised by outlets like Wired, Digital Humanities Now, and the New Statesman.

Editor - Jessica Brooksby

Jessica loves bringing her passion for the written word and her love of tech into one space at HighSpeedInternet.com. She works with the team’s writers to revise strong, user-focused content so every reader can find the tech that works for them. Jessica has a bachelor’s degree in English from Utah Valley University and seven years of creative and editorial experience. Outside of work, she spends her time gaming, reading, painting, and buying an excessive amount of Legend of Zelda merchandise.