AI Basics

How does an LLM work?

One prompt, followed all the way through

Scroll

One prompt, followed all the way through

This page follows a single prompt from the moment you press enter to the last word of the reply.

Your text is cut into tokens, each token becomes a list of numbers, those numbers pass through dozens of layers, and the last layer scores every possible next token. One is picked and added to the input, and the whole thing runs again. A reply is that loop, repeated until the model chooses to stop.

The sections below take it a piece at a time, in the same order as the long explainer: the network, how it learned, attention, how it answers, and what that means for you. For the full interactive version, read How LLMs Think.

Foundations
ONE PROMPT, ROUND THE LOOP text as typed tokens pieces vectors numbers layers × dozens scores every token pick one append it the chosen token joins the input, and the loop runs again the reply, one pass at a time: The capital is Paris . Everything an LLM does is this loop. The weights inside the layers were set during training, and they do not change while it answers.
Six stages and a loop. The reply grows by one token each time round.

It starts with a web of weighted connections

Inside the model is a neural network: billions of small units, each connected to units in the next layer. Each unit receives numbers, does a little arithmetic, and passes a result on.

  • Each unit multiplies its inputs by their weights, adds them up, and applies a simple bend
  • The weights are connection strengths, and they are everything the model learned
  • Units sit in layers, and information flows from one layer to the next
Neural networks
ONE UNIT, UP CLOSE 0.8 × 0.9 0.2 × −0.4 0.5 × 0.3 Σ sum, then bend 0.69 to the next layer 0.8×0.9 + 0.2×(−0.4) + 0.5×0.3 = 0.79 bend(0.79) → 0.69 Each unit is a weighted sum and a bend. A model is billions of them, layer after layer.
One unit: three weighted inputs, a sum, a bend, and a result passed on.

Learning by predicting the next word

Before it can answer anything, the model is trained on an enormous amount of text — books, articles, websites, code — to do one thing: predict the next word.

  • It sees The cat sat on the ___ and learns that mat is far likelier than helicopter
  • When a guess is wrong, the error flows back through the network and nudges every weight
  • After trillions of words, grammar, facts and patterns of reasoning are captured in the weights — imperfectly
Training
BEFORE AND AFTER TRAINING The cat sat on the ___ before trainingmat0.26floor0.25helicopter0.24banana0.25 after trainingmat0.71floor0.24helicopter0.01banana0.01 Each wrong guess sends an error back through the network, nudging every weight. Repeated over trillions of words, the right continuations rise to the top. Nobody teaches it grammar or facts. They are side effects of predicting well.
The same four candidates, scored before training and after.

Attention: working out what each word means here

Modern LLMs depend on a mechanism called attention. As the model reaches each word, it looks back over the words before it and decides which ones matter for understanding this one.

  • In the boat drifted toward the river bank, bank leans on river and means the edge of a river
  • In she deposited the check at the bank, the same word leans on deposited and means a place that holds money
  • Each word only looks backward, so meaning is always settled from what came before
Attention
SAME WORD, DIFFERENT NEIGHBORS weights are illustrative Theboatdriftedtowardtheriverbankriver 0.61 · boat 0.14 · drifted 0.08bank → the edge of a river Shedepositedthecheckatthebankdeposited 0.44 · check 0.31 · She 0.06bank → a place that holds money Each word looks back at the words before it and borrows meaning from the relevant ones. That is how the same word comes to mean two different things.
The same word, two sentences. What came before decides which meaning it takes.

From prediction to conversation

When you chat with an LLM, it is not looking answers up in a database. It generates the reply one token at a time, each time asking: given everything so far, what comes next?

  • Your prompt passes through every layer of the network in a fraction of a second
  • Roughly, earlier layers handle surface patterns and later ones more abstract meaning
  • The last layer scores the next token, one is chosen, and the loop goes round again
Inference
ONE WORD OUT, ADDED, ROUND AGAIN capital of France? → model The pass 1 … France? → The model capital pass 2 … → The capital model is pass 3 … The capital is model Paris pass 4 Each pass runs the whole input through every layer and produces exactly one token. A long answer is a loop that ran many times — that is the wait you see.
Four passes, four tokens. Each pass sees everything produced so far.

Why it matters

An LLM is a pattern-matching engine built from simple arithmetic, scaled up enormously. Knowing how it works tells you where to trust it and where to check it.

  • It is strong at language: writing, summarizing, translating, brainstorming and code
  • It predicts what sounds right, not what is right, so a fluent wrong answer is always possible
  • Clear instructions and the right context make it markedly better — the next levels are about exactly that
Big picture
LIKELY IS NOT THE SAME AS TRUE The capital of Australia is ___ illustrative scores Canberra 0.71 Sydney 0.24 Melbourne 0.04 a fluent wrong answer is always on the menu strong at writing, summarizing, translating, code check it on facts, numbers, quotes, citations Sometimes the wrong answer gets picked. Ask clearly, and verify what matters.
The likeliest answer wins most of the time. The wrong one still has a real chance.