AI Basics
A next-token predictor, trained at scale
Scroll
A next-token predictor, trained at scale
A large language model is a generative model for text. It was trained on an enormous amount of writing to do one narrow thing: given some text, predict what comes next.
It does not produce a whole answer at once. It scores every token it knows as a candidate for the next position, picks one, appends it, and repeats with the longer text as its new input. Summaries, translations, code and conversation are all that single step run many times until the model produces a stop token.
“Large” refers to the number of adjustable values inside it, typically billions, and to the volume of text it learned from. What it learns is which text is likely, which is mostly but not always the same as which text is true. Much of the rest of this curriculum is about engineering around that gap.
Foundations