AI Basics
One prompt, followed all the way through
Scroll
One prompt, followed all the way through
This page follows a single prompt from the moment you press enter to the last word of the reply.
Your text is cut into tokens, each token becomes a list of numbers, those numbers pass through dozens of layers, and the last layer scores every possible next token. One is picked and added to the input, and the whole thing runs again. A reply is that loop, repeated until the model chooses to stop.
The sections below take it a piece at a time, in the same order as the long explainer: the network, how it learned, attention, how it answers, and what that means for you. For the full interactive version, read How LLMs Think.
FoundationsInside the model is a neural network: billions of small units, each connected to units in the next layer. Each unit receives numbers, does a little arithmetic, and passes a result on.
Before it can answer anything, the model is trained on an enormous amount of text — books, articles, websites, code — to do one thing: predict the next word.
Modern LLMs depend on a mechanism called attention. As the model reaches each word, it looks back over the words before it and decides which ones matter for understanding this one.
When you chat with an LLM, it is not looking answers up in a database. It generates the reply one token at a time, each time asking: given everything so far, what comes next?
An LLM is a pattern-matching engine built from simple arithmetic, scaled up enormously. Knowing how it works tells you where to trust it and where to check it.