Post

What Is a Language Model? — LLMs Part 1

What Is a Language Model? — LLMs Part 1

3 min read — LLM Concepts, 3 Minutes Each, Part 1

The Definition

A language model is a neural network trained to do one thing: given some text, predict what comes next. We’ll define a neural network in Part 2. For now, think of it as a program that learns from examples instead of being hand-coded.

LLM is short for “large language model” — a language model built at a very large scale. Everything in this series is about that large-scale kind, since that’s what people mean by “LLM” today.

That’s the whole object. A next-word predictor, scaled up.

The Idea

An LLM predicts text one small piece at a time. Each piece is called a token. We’ll define a token properly in Part 3. For now, think of it as a word or part of a word.

To generate text, the model repeats one step: look at the text so far, predict the next token, add it to the text. Then repeat.

Writing, answering, and coding are all this same step, repeated many times.

How It Works (the short version)

  • You give it text: "The capital of France is"
  • It ranks every possible next token by probability: Paris (92%), located (3%), a (1%), …
  • It picks one token and adds it to the text
  • It repeats this, one token at a time, until the answer is done

The model learned these probabilities during training. It read enormous amounts of text. Each time, it guessed the next word. Each time it guessed wrong, it adjusted itself slightly. It did this billions of times.

This explains what the model does. It doesn’t yet explain how it does it — that’s the rest of this series.

Example

Prompt: "Roses are red, violets are"

The model’s next-token probabilities might look like this:

1
2
3
4
blue    → 71%
violet  → 12%
purple  → 8%
loud    → 0.01%

It picks blue. Now it predicts the token after “blue.” Same process, one step later.

Why It Matters

An LLM is a next-token predictor. Nothing more, nothing hidden.

This explains a lot. It explains why an LLM can finish your sentence. It also explains why an LLM can confidently state something false. Both are the same action: predicting a plausible next token.


Next: Neural Networks in One Page — what a neural network actually is. (Coming soon.)

This post uses several terms loosely: neural network, token, training, “large.” Each gets a precise definition in the part of the series listed above.

This post is licensed under CC BY 4.0 by the author.