You can check. Not for every dataset (the biggest ones are proprietary), but for the open datasets that power most of the models you have heard of, there are real tools that give you real answers. Here is what you can search, how to do it, and what grain of detail you get back.
| Dataset | What It Contains | Tool | What You Get |
|---|---|---|---|
| LAION-5B | 5.85 billion image-text pairs scraped from the public web. Used to train Stable Diffusion, Google Imagen, and others. | haveibeentrained.com | Individual image match. Upload your image and the tool returns whether it appears in LAION-5B, with source URLs. You get a yes/no at the file level. This is the highest-resolution search available for any medium. |
| LAION-5B (advanced) | Same dataset, programmatic access. | LAION clip-retrieval | Search by text query or image. Returns matching image URLs from the dataset. Useful for checking whether your work appears under variations of your name, brand, or style descriptors. |
Spawning AI runs the haveibeentrained.com tool and also maintains a Do Not Train registry. If you find your work in LAION, you can register an opt-out that Stability AI has committed to honoring in future model training. Spawning also builds Kudurru to block data scrapers in real time.
| Dataset | What It Contains | Tool | What You Get |
|---|---|---|---|
| Common Crawl | Billions of web pages crawled monthly since 2008. Common Crawl is the foundation for C4, Dolma, The Pile, RefinedWeb, and most open LLM training sets. If your website or blog is in Common Crawl, it is almost certainly in one or more training datasets. | Common Crawl CDX Index | URL-level detail. Search for your domain or specific page and see every crawl date, HTTP status, and WARC file location. Example query: enter mysite.com to see every page from your domain that was crawled. |
| Dolma (Allen AI) | 3 trillion tokens from Common Crawl, Reddit, academic papers, code, books, and Wikipedia. Used to train OLMo. Fully open and downloadable from HuggingFace. | Dolma on HuggingFace | The dataset is downloadable for inspection. Its provenance is documented: Common Crawl snapshots, Pushshift Reddit, Semantic Scholar, StackExchange, Project Gutenberg. If your content is in any of those sources, it is likely in Dolma. |
| The Pile (EleutherAI) | 825 GB of text from 22 sources: books, web pages, academic papers, code repositories, Reddit, Wikipedia, and more. | The Pile paper & sources | The Pile lists all its constituent datasets. Cross-reference: if your content appears in any of the 22 sources (Books3, OpenWebText2, Pile-CC, PubMed, GitHub, StackExchange, etc.), it is in The Pile. |
Common Crawl is the upstream source for nearly every open text dataset used in LLM training. C4, Dolma, The Pile, RefinedWeb, RedPajama, and many others all start from Common Crawl snapshots and apply their own filtering. A hit in the Common Crawl index means your content almost certainly flowed downstream into multiple training datasets. The CDX index is free, public, and responds in seconds. There is no excuse not to check.
| Dataset | What It Contains | Tool | What You Get |
|---|---|---|---|
| Pushshift Reddit Dataset | Every public Reddit comment and submission from Reddit's founding through early 2023. Ingested into both The Pile and Dolma. Each record includes the username, subreddit, timestamp, and full text. | Archives hosted on Academic Torrents; subset on Kaggle | Username-level detail. The author field maps every comment and post to a Reddit username. If your username appears in these archives, your writing is in the training data for many LLMs. |
Reddit is uniquely searchable at the username level because Pushshift preserved the author field. No other social media platform offers this resolution. Twitter data was largely excluded from training datasets due to API restrictions. Facebook and Instagram content is in Meta's proprietary datasets with no public search tool. Pinterest images may be in LAION but there is no username-level check.
| Dataset | What It Contains | Tool | What You Get |
|---|---|---|---|
| Books3 | 191,000+ books used to train models from Meta, Bloomberg, and others. A subset of The Pile. | The Atlantic Books3 Search | Author-level. Search by author name to see which of your books appear in the dataset. |
| Project Gutenberg, PubMed, Semantic Scholar | Public domain books, biomedical abstracts, and academic papers. Ingested by both The Pile and Dolma. | Source sites directly | If your work is on these platforms, it is very likely in training datasets. |
| Dataset | What It Contains | Tool | What You Get |
|---|---|---|---|
| 4 music training datasets | ~12 million tracks from YouTube and Spotify, downloaded via automated tools. The largest dataset alone holds ~12M tracks. Google and Stability AI have acknowledged using material from these collections. | The Atlantic AI Watchdog | Artist-level. Search by artist name to see if your music appears. Daft Punk, Aphex Twin, Radiohead, Wu-Tang Clan, and many independent musicians have already been found in these datasets. |
| Dataset | What It Contains | Tool | What You Get |
|---|---|---|---|
| The Stack v2 (BigCode) | 6.4 TB of permissively licensed source code in 358+ programming languages from public GitHub repositories. Used to train StarCoder, StarCoder2, and other code-generation models. | Am I in the Stack | Repository-level. Enter your GitHub repo URL or username. The tool tells you whether your repos are in the dataset. You can also request removal through the BigCode opt-out repo. Opt-outs are honored in future dataset releases. |
| GitHub subset of The Pile | Public GitHub repositories scraped and filtered for The Pile. Included in training data for GPT-NeoX, GPT-J, Cerebras-GPT, and others. | The Pile GitHub repo documents the exact scraping parameters | If your public repos existed before 2020-2021, they are likely in The Pile. Check The Pile paper for source descriptions. |
Code is actually one of the more searchable mediums. BigCode built the "Am I in the Stack" tool specifically for developer agency. If you write open source code, check it. Opt-out is real and honored.
The tools above only work because some datasets are open. Things scraped from the public web. The private walled gardens are entirely different. Here is the honest accounting.
These companies build the models. Their training datasets are proprietary:
These companies collect your data constantly. Whether they train AI on it is something they decide behind closed doors. For every platform below: no public search tool exists. You cannot verify what they have used.
| Platform | What They Collect | AI Training Status |
|---|---|---|
| YouTube | Every video you upload, every comment, every watch history entry. Transcripts are auto-generated and stored. YouTube is almost certainly the largest single source of training video and audio data in existence. | Google has never confirmed or denied using YouTube content to train Gemini. Public research datasets (YouTube-8M, HowTo100M) are indirect evidence. Some transcripts may appear in Common Crawl if the video page was crawled. But your actual videos, your voice, your face — no search tool exists. Google decides what gets used. You get no say. |
| Customer Service Calls | Recordings, transcripts, sentiment analysis. Millions of hours. Your voice, your accent, your complaints, your account numbers. Recorded "for quality assurance" and increasingly sold to data brokers and AI companies. | Companies like LivePerson, Five9, and Genesys sell call center data to AI trainers. Your bank's IVR recording, your insurance claim, your tech support rant. If a company recorded your call, that recording may be in a training dataset. You will never know. There is no opt-out. The only way to avoid this is to never call a 1-800 number. |
| Gmail | Every email you send and receive. Google stopped scanning email for ad targeting in 2017, but the data is stored and Google's privacy policy allows internal use for product improvement. | Google has made conflicting statements. In 2023 they updated their privacy policy to say they use "publicly available information" to train AI. That language was later expanded. Whether "publicly available" includes Gmail content that you wrote but did not publish is a question Google has chosen not to answer clearly. Assume your Gmail content has been used. You cannot check. |
| Google Docs & Drive | Every document, spreadsheet, and file you store. Google's Workspace terms allow internal analysis and improvement of Google products. | No public disclosure. Google has stated that Workspace data is not used to train consumer-facing AI, but the distinction between "consumer AI" and "internal AI" is not defined in any document you can read. You cannot check. |
| Pixel Phones & Google Assistant | Voice commands, location history, app usage, call logs, SMS content (if you use Google Messages), photos (if you use Google Photos), browsing history, and every tap and swipe in every app. Pixel phones are the most instrumented consumer devices ever made. | Voice data from Assistant interactions has been used for speech recognition training, with human reviewers listening to recordings. Google paused human review in 2019 after public outcry but resumed with additional consent language. Whether this data has been used to train LLMs or multimodal models is unknown. Google does not disclose which device data feeds into which training pipelines. You cannot check. |
| Fitbit | Heart rate, sleep patterns, step count, GPS routes, menstrual cycles, weight, blood oxygen. Your body's data, continuously recorded. | Google acquired Fitbit in 2021. They have stated Fitbit health data will not be used for advertising. They have never stated that it will not be used for AI training. Health data is the most sensitive category of personal information. It is also extraordinarily valuable for training health-related AI models. You cannot check. You cannot opt out without discarding the device. |
| Apple (iPhone, Siri, iCloud) | Your photos, messages, voice recordings, location, app usage, health data, and payment history. Apple markets privacy as a selling point. | Apple runs most processing on-device or through their Private Cloud Compute architecture, which they claim does not retain your data after processing. Apple Intelligence features use on-device models. They have stated they do not train on user data. Whether this remains true, and whether it applies to all future Apple AI products, is a promise from a corporation. It is not independently verifiable. You cannot check. |
| Meta (WhatsApp, Messenger, Facebook DMs) | Private messages, group chats, voice calls, and video calls. WhatsApp claims end-to-end encryption for message content, but metadata (who you talk to, when, how long, from what device) is collected and stored. | Meta has stated that private messages are not used for training Llama. But Meta has also been caught making claims about data practices that turned out to be incomplete. WhatsApp metadata is collected and could be used for AI without touching message content. There is no way to verify. You cannot check. |
| GitHub Copilot | Your code prompts, completions, code snippets, and interaction context. As of April 2026, GitHub uses this data from Copilot Free, Pro, and Pro+ users to train AI models. Private repository code at rest is not used. Business and Enterprise tiers are excluded. | You CAN check and opt out. Go to github.com/settings/copilot/features and disable "Allow GitHub to use my code snippets for product improvements." This is one of the few platforms where opt-out is built into the product settings. If you do nothing, your Copilot interactions are training data by default. |
If a company collects your data, and that company builds AI models, and they will not let you check whether your data was used — assume it was used. The burden of proof is on them. They have the data. They have the models. They will not open either to inspection. You do the math.
The platforms you cannot check are black boxes. But the open datasets (Common Crawl, LAION, The Pile, Dolma) are the upstream source for most text and image training. If your work was ever publicly accessible on the web, start by checking Common Crawl. If your image was ever posted online, check LAION. A hit in the open datasets tells you something real. A miss tells you nothing — your work could still be in a proprietary dataset. Absence of evidence is not evidence of absence.
Even if you cannot check every dataset, you can stop some platforms from using your future content for training:
This prevents new conversations from being used for training. It is not retroactive.
Microsoft automatically opts you in. Your documents are scraped for AI training unless you disable this.
Adobe's updated Terms of Service allow them to use your artwork for AI training. This setting opts you out. Not available for Adobe Stock contributors.
LinkedIn has already trained models with user content. This stops future use. Not retroactive.
Opt-out is available only to EU users under GDPR. For users in the US and elsewhere, Meta provides no opt-out. To file an objection if you are in the EU: Settings & Privacy → Privacy Center → "How Meta uses information for generative AI models and features" → Right to Object.
Some creators are going further than searching. Tools from the SAND Lab at the University of Chicago let you alter your work in ways that degrade AI training on it:
These tools are controversial and their long-term effectiveness is debated. They are an active area of research.
Whether you find your work in a dataset or not, document everything. Screenshot your search results. Save the URLs. Note the dates. See our Documenting Infringement guide for a complete checklist.
Found your work? The next step is documentation and, if appropriate, a takedown.
Documenting InfringementLast reviewed: July 2026. Tool URLs and datasets are current as of this date. Common Crawl and LAION are continuously updated; the specific snapshots used by each training dataset may differ from what is live today. This guide provides information. It does not constitute legal advice. Beaumont & Sheridan is not a law firm.