Ask ChatGPT a question and it answers in one confident paragraph. No sources, no ten blue links. It's fair to wonder where that answer comes from, especially if you sell software and your buyers now ask ChatGPT which tool to pick. And once you know, the more useful question follows: how do you get inside it?

Where does ChatGPT get its information?

ChatGPT gets its information from two places: a training dataset it learned once, covering the web up to a fixed knowledge cutoff, and live web results it pulls in real time when it decides to search. It isn't looking answers up in a database. It predicts the next words from patterns it absorbed during training, and when it browses, it grounds that prediction against current pages.

Those two modes behave very differently, and that difference is the whole story. The trained model is fast, fluent, and frozen in time. The browsing model is slower but current. Most of what feels like ChatGPT "knowing" something comes from the first. Most of what lets a brand that launched last quarter show up at all comes from the second.

What is ChatGPT trained on?

ChatGPT is trained on three buckets of data: publicly available text from the open web, data licensed from partners, and feedback from human trainers. That's OpenAI's own description of how its models are built, and it's more honest than most of the guessing around it.

The open web is the big one. Think Wikipedia, news archives, forums like Reddit, public code, digitized books, and a long tail of ordinary websites. Licensed data fills gaps the open web leaves. The human feedback layer, known as RLHF, doesn't add facts so much as shape behavior: which answers are helpful, which are unsafe, what tone to use.

One thing to hold onto. The model doesn't store these sources like files in a cabinet it can reopen. It learned statistical patterns from them, then threw the originals away. That's why ChatGPT can sound authoritative and still be wrong. It's reconstructing what an answer probably looks like, not retrieving a document it saved.

Does ChatGPT get its information from Google or the live internet?

Increasingly from Google. When ChatGPT searches the live web, its retrieval leans heavily on Google's index for current and product questions, with Bing and OpenAI's own crawler in the mix too. Without browsing turned on, it doesn't touch the internet at all, and answers from the trained model, frozen at its knowledge cutoff and blind to anything after that date.

For a brand, the cutoff is where the trouble starts. If your category moved, your product shipped, or your positioning changed after the model was trained, the frozen version of ChatGPT won't know. It'll answer from an older snapshot of the web, or from whatever it can pull in the moment it searches. Which is why being present in the sources it can reach, right now, matters more than being "well known" in some general sense.

Where ChatGPT actually pulls its answers from

When ChatGPT browses for an answer, it favors pages it can read cleanly and reasons it can trust: established sites, clear structure, and claims that other sources back up. Which of your pages it reaches for then depends on the exact question being asked.

Ask it about a specific product or brand and it often pulls from that company's own pages: its docs, pricing, and help center. Ask it to compare or recommend within a category and it leans on third-party sources: review sites, roundups, and community threads. In our own look at AI citations, that split ran sharper on ChatGPT than on the other tools, which lean harder on third-party sources across the board.

The practical read is simple. Your own site does more work on ChatGPT than most people assume, so it has to be crawlable and clear. And the closer a buyer gets to a decision, the more the answer depends on what other people say about you.

How to become a source ChatGPT cites

To become a source ChatGPT pulls from, you need two things: to be reachable, and to be worth reaching for. Reachable means crawlable and indexed. Worth reaching for means you publish things a model has a reason to quote. There's no "rank number one" here. You're either in the generated answer or you're not.

Four moves do most of the work. First, get into the index the retrieval layer uses, which means Bing as well as Google, with clean pages AI crawlers can read. Second, earn third-party mentions, because on most engines that's still where the citations come from, and unlinked brand mentions now count as a signal in their own right. Third, publish primary research. A number nobody else has is the single most quotable thing a company can own, because the model has nowhere else to get it. Fourth, keep your entity consistent: same name, same category, same facts across your site, your listings, and the third-party pages that describe you.

If your brand's missing from AI answers today, that's usually diagnosable rather than mysterious. Authority comes first, though: the mentions, links, and reviews that make the web treat you as credible are what actually get your pages pulled. Build the entity first around a consistent story, then the citable pages do their work. The disciplines are the same as they always were. The signals are different, and the entity is the asset.

You can't fix what you can't see. Before you spend a penny on AI search, know exactly what ChatGPT tells your buyers about you, and where it pulled that from. That's where every VisibleIQ engagement starts. See how we work.

FAQ

Does ChatGPT cite sources? When it browses, yes, it'll often show links to the pages it used. When it answers from training data alone, no. It's reconstructing from patterns, not quoting a document, so there's no source to show.

Does ChatGPT use Google? Yes, more than people assume. For up-to-date and product questions, its live search leans heavily on Google's index, so being visible in Google is a major part of being visible in ChatGPT. Bing helps too.

How often does ChatGPT's information update? The trained model updates only when OpenAI releases a new version with a later knowledge cutoff, which is occasional. Live browsing, when it's on, is current to the moment of the search.

Is it safe to tell ChatGPT confidential things? Treat anything you type as potentially used to improve the model unless you've turned that off or you're on a plan that excludes it. For sensitive company data, assume caution.

Rafael De Jesus
Founder
Organic growth and AI search optimization specialist. Rafael has added seven figures in ARR to B2B SaaS and AI-native companies, built a 200K+ following, and generated 3 billion+ impressions across search and social. He writes about how to shape what AI says about your brand.