AI Is a Privacy Nightmare
Tech companies want your data one way or another
Oct 7, 2026 | Share
Technology
Some of the biggest legal battles over AI currently have to do with privacy in one form or another. Elon Musk’s Grok chatbot is being sued both in the U.K. and in the U.S. over its use in the creation of nonconsensual sexual images of women and children. An AI company in Illinois was sued for secretly scanning faces from people’s online photos. Meta has been sued for allegedly using its AI-powered smart glasses to collect data to train its AI models.
Protecting your online privacy is important, but AI is making it a much more difficult task. We’re going to talk you through the privacy risks of AI and some of the things you can do to help protect yourself online.
AI companies want your data
Large language models (LLMs), at their core, are statistical pattern predictors. They generate content like a very sophisticated version of the autocomplete function on your phone. The models themselves are primarily a huge database of numbers used to calculate the probabilities for this autocomplete function.
These numbers are set during the training phase of the model’s development, which is a bit like tuning a machine with a million little dials. When the machine is fed a document where two words show up next to each other, the corresponding knobs get a tweak. The more training data the model gets, the better its settings become for predicting text.

Unfortunately for AI companies, investing in model training produces diminishing returns. A 2023 study by Andrew J. Lohn notes that, according to data from OpenAI, the makers of ChatGPT, a model that received $10 million worth of training could correctly solve a programming problem on average 64% of the time (PDF). If the company spent $100 million on training a new model, that model on average would solve the problem 74% of the time. This means that if a company wants to gain a tiny edge over the competition, it needs vastly more compute capacity and training data.
Unfortunately for everyone else, a lot of this data comes from us. AI companies send out bots to crawl the internet to harvest our blog posts, our Instagram photos, our LinkedIn profiles, and any other data that’s publicly visible on the internet. Chances are you didn’t put that information on the internet for the purposes of training AI, and you might not like what the AI does with your information either.
In 2023, Workday was sued for allegedly using its AI-powered job screening tools to unlawfully filter out applicants based on race, age, and disability. These AI tools were trained on the data of countless people just like the plaintiffs, and perhaps even the plaintiffs’ own data, though it’s probably impossible to know due to AI companies’ lack of transparency.
This lack of transparency itself can infringe on people’s rights. When organizations like credit bureaus collect data on you, they’re required by federal law to let you know about the scores they give you and allow you to dispute them. AI tools can fulfill a similar function.
In January 2026, hiring platform Eightfold AI was sued for collecting personal data on users from across the internet and ranking them on a scale from 0–5 without their knowledge. If a company used Eightfold’s tools, any candidate with a low ranking would be automatically rejected before a human ever saw their resume.
And since Eightfold claims its data sets include “the profiles of more than 1 billion people working in every job,” chances are they’ve got your data too (PDF).
AI makes normal privacy concerns worse
While AI companies have created new privacy concerns, they also make existing ones worse. Since the rise of social networks like Facebook and Twitter, users’ privacy has slowly eroded as their data became more valuable. Third-party data brokers have made the situation worse, spreading our data to sites and organizations we never even interacted with.
With all the money in the tech industry going to AI companies, it’s unsurprising that these same data brokers are now selling all our information to these companies to be fed into LLMs. Or in many cases, the social networking platform that already has your data is developing its own AI, and most likely the terms and conditions of using that site give it the right to all your data.
In June 2026, Reddit accused Anthropic of training its Claude AI on its users’ personal information without their consent. Reddit had already made deals with companies like Google and OpenAI to license that data (so this isn’t exactly a win for Redditors’ privacy), but Anthropic allegedly decided that it could just take the data without paying.
AI users aren’t safe either
AI companies don’t always have to harvest or buy valuable data. Sometimes they just wait for their own users to give it to them.
In September 2026, OpenAI announced that it had solved the Navier-Stokes equations, a longstanding math problem that had stumped mathematicians for years. It apparently used a method being used by a mathematician at NYU … who just happened to be using OpenAI coding tools in his work. While he has not openly accused OpenAI of stealing his work through its coding tools, according to a Forbes article, OpenAI “cannot rule out that the de-identified data from its usage helped improve the models.”
Businesses that make widespread use of AI often use more expensive enterprise tools specifically to prevent proprietary data from leaving the company. This incident highlights the privacy risks of using any AI not hosted on your own servers.
Individual users should likewise be wary of entering any sensitive or personal information into an AI prompt.
AI is a great tool for bad actors
While plenty of big companies are engaging in morally questionable but technically legal AI practices, many scammers have latched onto AI. AI can help scammers reach a much larger audience while also making their tools more powerful. Instead of sending out generic phishing attacks, the scammers can tailor their attacks to specific people.
Scammers aren’t the only bad actors to make extensive use of AI. Elon Musk’s xAI is being sued both in the U.K. and in the U.S. over the use of its Grok chatbot in the creation of nonconsensual sexual images of women and children. The spread of these images has caught the attention of regulators from countries around the world and caused public backlash against the company.
Looking for a way to stay safe while online?
Using a VPN masks your IP address and safeguards your online privacy. Find out how to get started with a VPN from our tech experts.
How to protect yourself from AI
As with online privacy more broadly, there is only so much you can control as an individual. And with issues like Meta’s smart glasses allegedly being used to harvest visual data, even remaining offline isn’t a surefire way to keep your data private.
Still, there are precautions you can take to keep your personal information safe, as well as public policies that could stop some of the more egregious violations of people’s privacy.
Be careful with your data
You should always be careful about what data you post on which platforms, but even more so with the widespread use of generative AI. Your photos and personal information could be used by scammers to impersonate you or fed into image generation tools to bully or harass you or used in any number of nefarious ways.
In general, try to be aware of the privacy policies of the platforms you use and how likely they are to share your information with third parties or use it for their own AI training. Make use of the privacy controls available to you and post information with those levels of protection in mind.
Be very careful with any information you give to a chatbot. AI companies have a bad record of harvesting their users’ data without their consent. Even businesses have to be careful about sensitive company information being put into AI prompts, so you should too. As a general rule, don’t give an AI any personal information you wouldn’t want the whole world to know.
Finally, if you are using AI, there are privacy-focused options, like duck.ai, that don’t hold onto your personal information and generally aren’t as intrusive. The trade-off is that by respecting user privacy, they often give slightly more generic answers (though the bar is pretty low).
Support privacy-focused legislation
Many of the severe privacy issues related to AI are structural issues that can only be addressed through regulation and legislation. It’s important to realize, however, that not all proposed AI laws are actually privacy-focused. For example, many of the proposals to protect children from the dangers of social media and AI revolve around age verification requirements. In general, these laws would force you to give personal, identifying information to the same tech companies that won’t stop trying to steal your personal, identifying information.
Similarly, approaches that look to protect artists and musicians from having their work stolen by AI image or music companies often focus on simply strengthening copyright laws, which would often take power away from individual artists in favor of large music publishers.
We also don’t want regulations that negatively impact the normal working of the internet. Web crawlers, for example, are essential tools for the function of search engines, for academic research, and for digital preservation through organizations like the Internet Archive. We don’t want to ban the use of crawlers on the internet; we just want to stop AI companies from harvesting our data and often crashing our sites in the process.
In a White Paper for the Stanford University Institute for Human-Centered Artificial Intelligence, Jennifer King and Caroline Meinhardt make three suggestions:
- De-normalize data collection by default.
- Make the entire AI data supply chain transparent and accountable.
- Support the development of systems that empower individual data rights.
AI companies can act with blatant disregard for privacy because our current systems have no built-in mechanisms to hold them accountable. We don’t cut other technologies this much slack. You can’t walk into someone’s house and take a photograph of them. When someone sues you for libel, you can’t say “the printing press did it.”
When new technologies disrupt society, we’ve had to make new laws to protect individual rights and the public good. In that sense, there’s nothing that new or different about AI.
Additional resources
Author - Peter Christiansen
Peter Christiansen writes about telecom policy, communications infrastructure, satellite internet, and rural connectivity for HighSpeedInternet.com. Peter holds a PhD in communication from the University of Utah and has been working in tech for over 15 years as a computer programmer, game developer, filmmaker, and writer. His writing has been praised by outlets like Wired, Digital Humanities Now, and the New Statesman.
Editor - Jessica Brooksby
Jessica loves bringing her passion for the written word and her love of tech into one space at HighSpeedInternet.com. She works with the team’s writers to revise strong, user-focused content so every reader can find the tech that works for them. Jessica has a bachelor’s degree in English from Utah Valley University and seven years of creative and editorial experience. Outside of work, she spends her time gaming, reading, painting, and buying an excessive amount of Legend of Zelda merchandise.



