In Pew Research Center's February 2026 survey of 5,119 U.S. adults, 49 percent said they use AI chatbots, up from 33 percent in the summer of 2024. (Pew asked the question more broadly in 2026, so the comparison is approximate.) Far fewer people could say what happens between the moment they press Enter and the moment an answer appears. That is worth knowing, because it explains both why these tools are so capable and why they can be wrong with total confidence.
It writes one piece at a time
OpenAI's own explanation is plain. Its model "uses that understanding to predict the next most likely word when generating a response, one word at a time." Strictly, the pieces are not words but tokens. "A token can represent a character, part of a word, a whole word, or punctuation."
OpenAI's rule of thumb for English is that one token is about four characters, so 100 tokens is about 75 words. Google gives a similar estimate for its Gemini models: about four characters per token, and 60 to 80 English words per 100 tokens. The figures differ because each company splits text its own way, and other languages and computer code usually take more tokens per word.
Answering is a loop. The model looks at everything in the conversation so far, scores the possible next tokens, picks one, adds it to the text, and repeats until the reply is finished. There is no separate step where it looks up a fact and checks it. The fluent paragraph you read is the output of that loop.
Where the patterns come from
Before a chatbot answers anyone, it is trained by "learning patterns from large amounts of information, including text, images, audio, and video." OpenAI names three kinds of training data: information publicly available on the internet, information it accesses through partners, and information that users, human trainers, and researchers provide or generate. It also says its models "do not store or retain copies of the data they are trained on." Training adjusts the model's internal settings, called parameters, rather than filing documents away.
The design behind today's chatbots is the Transformer, introduced in 2017 in the paper "Attention Is All You Need" by eight researchers at Google Brain, Google Research, and the University of Toronto. Their new architecture was "based solely on attention mechanisms, dispensing with recurrence and convolutions entirely." In plain terms, attention lets the model weigh every part of the text against every other part when deciding what comes next, which is why it can keep track of a long question.
Why it sounds so helpful
A freshly trained model only continues text. Turning it into an assistant takes another step. In OpenAI's 2022 InstructGPT paper, people wrote example answers and ranked the model's outputs, and those rankings were used to train it further, a method called reinforcement learning from human feedback. The result surprised many readers: evaluators preferred the answers of a 1.3 billion parameter model tuned this way over the 175 billion parameter GPT-3, "despite having 100x fewer parameters."
Confabulations are a natural result of the way generative models are designed.
Why it makes things up
The everyday word is hallucination. The National Institute of Standards and Technology uses confabulation, which it defines as "the production of confidently stated but erroneous or false content." NIST is direct about the cause: generative models "generate outputs that approximate the statistical distribution of their training data; for example, LLMs predict the next token or word in a sentence or phrase." A likely sounding answer and a true answer are not the same thing.
OpenAI's own researchers point to a second cause in how models are tested. "When models are graded only on accuracy, the percentage of questions they get exactly right, they are encouraged to guess rather than say 'I don't know.'"
What it does not know
Every model has a knowledge cutoff. In OpenAI's words, "The models are trained on data up to a certain point and responses do not incorporate information about events beyond that, unless tools are used." Web search is the most common of those tools: "ChatGPT can search the web to answer questions with current information and links to relevant sources." Whether a given chatbot searches depends on the product, the plan, and the settings. When it does, the summary is still written by the same next token loop, so open the links it cites.
What happens to what you type
This is the part most users skip. On OpenAI's services for individuals, "we may use your content to train our models." You can opt out by turning off "Improve the model for everyone" under Settings, then Data controls, or through OpenAI's Privacy Portal. OpenAI says it does not use ChatGPT Business, Enterprise, Edu, or API data to improve its models by default.
Anthropic changed its consumer terms on August 28, 2025. Users of Claude Free, Pro, and Max now choose whether their chats are used for training. If they allow it, the data is kept for five years; if not, the existing 30 day retention period applies. The choice can be changed at any time in Privacy Settings, and Incognito chats are not used for training. Business and API products at both companies follow separate terms.
How to use one well
- Treat a fluent answer as a draft, not a source. Fluency is what the system is built to produce.
- Ask for sources, then open them. A citation that does not lead anywhere is a warning sign.
- Ask whether it searched the web. If not, assume nothing after its knowledge cutoff is included.
- Check your data settings before you paste in client, medical, or financial details.
- Use it where you can check the work: summarizing text you supplied, drafting, outlining, and brainstorming.