Post

Tokens, Not Words — LLMs Part 3

Tokens, Not Words — LLMs Part 3

The Definition

A token is the actual unit a language model reads and predicts. It is not always a word. It can be a whole word, part of a word, or a single punctuation mark.

Before a model is trained, its builders fix a list of allowed tokens, called a vocabulary. Every model has its own vocabulary, typically tens of thousands of entries. A tokenizer is the program that splits any input text into tokens from that vocabulary — a fixed procedure, decided once before training, that never changes once the model is in use.

Why Subword, Not Whole Words

A vocabulary of whole words alone can’t handle a word it has never seen before — a typo, a made-up name, a rare technical term; there’s simply no entry for it. A vocabulary of individual characters could handle anything, but at a real cost: a word like “unhappiness” would need 11 separate tokens, one per letter, instead of a handful of larger pieces. Subword tokenization splits the difference: common words stay whole, and rare or unfamiliar ones break into smaller, reusable pieces — “unhappiness” becomes 3 tokens, not 11 or 1, shown below.

Vocabulary Size Varies by Domain

“Tens of thousands of entries,” from the top of this post, covers a real range. English-focused models tend toward the smaller end — Llama 2 uses about 32,000 tokens, Falcon about 65,000, and OpenAI’s Generative Pre-trained Transformer (GPT) line started at about 50,000 with GPT-2. Models built to handle many languages well need far more: GPT-4’s tokenizer uses about 100,000, Llama 3 about 128,000, and heavily multilingual models like mT5 and BLOOM (BigScience Large Open-science Open-access Multilingual Language Model) go up to around 250,000. Roughly, a multilingual vocabulary needs 2 to 3 times as many entries as an English-only one to avoid the same over-fragmentation problem non-English text already has in a smaller vocabulary.

Code doesn’t really need a bigger vocabulary — it needs the right entries allocated differently. Whitespace and indentation make up a large share of any code file, but are rare enough in ordinary English prose that a tokenizer built only for English gives them almost no dedicated tokens. Point that same tokenizer at Python and every level of indentation costs several tokens just to represent blank space. A tokenizer meant to handle code well, like GPT-4’s, deliberately adds tokens for common runs of whitespace, so indentation costs close to one token instead of many.

Highly specialized domains — biomedical research, law, physics and math — run into a similar problem from a different angle: a general-purpose vocabulary was never built from enough of that domain’s own text, so its distinctive terms fragment into meaningless pieces, the same way an unusually long or unfamiliar word does. A drug name, a legal doctrine, or a block of mathematical notation can all end up split into pieces that don’t line up with any meaningful unit in the field. Models built for these domains often use their own custom vocabulary instead, trained on that domain’s own text — a biomedical model called SciBERT, for instance, uses a 30,000-token vocabulary built specifically from a mix of scientific and biomedical papers, and specialized legal models have used custom vocabularies of a similar size built from legal text.

One more distinction worth being precise about: a search engine’s “tokenization” is not this at all. It’s a much simpler, older idea — splitting text into words for a search index, usually by whitespace and punctuation, sometimes reduced further so different forms of a word match each other. It shares the word “tokenization” with what this post describes, but not the mechanism, and not the goal.

How a Tokenizer Splits Text

Given text, a tokenizer’s job is to turn it into a sequence of tokens from its fixed vocabulary:

flowchart LR
    TXT["text"] -->|tokenizer| TOK["tokens"]

A tokenizer can only ever produce tokens that are already in its vocabulary — it has no way to invent a new one on the spot, even for text it’s never seen before. Common short words usually become one token each. Longer or rarer words often split into pieces — roughly like this, for illustration (the exact split varies by tokenizer):

1
2
3
4
"the"            -> [the]
"cats"           -> [cat, s]
"unhappiness"    -> [un, happi, ness]
"antidisestablishmentarianism" -> [anti, dis, establish, ment, arian, ism]

The exact split depends on the tokenizer’s vocabulary — different models split the same word differently. A future post covers the actual algorithm most tokenizers use to build that vocabulary.

A vocabulary is fixed before training and never grows afterward, so what happens when the input has something outside it — an emoji, an unfamiliar writing system, a garbled string of characters? Older tokenizers replaced anything they didn’t recognize with a single generic “unknown” token, throwing that piece of text away entirely. Modern tokenizers avoid this: their vocabulary also includes every individual raw byte, so any input can always be represented, worst case one byte at a time. Nothing is ever truly unrecognizable — text with no matching pieces in the vocabulary just costs far more tokens to represent, one byte at a time instead of one word or subword at a time.

The tokenizer isn’t part of the neural network itself. It runs once, as a fixed step before any of the math in the network’s layers happens: text goes in, a list of integers comes out, and only then does the network — through the vocabulary lookup a future post covers — actually begin its calculations.

From Tokens to Numbers, and Back

Once split, the tokenizer replaces each token with a number called its token ID — its position in the vocabulary list. The network never sees the text “cat”; it sees a token ID, like 4187. Every distinct token has exactly one such number, and the same token always gets the same number: the mapping is fixed once the vocabulary is built, not reassigned from run to run. That number is arbitrary, though — token 4187 isn’t “closer” to token 4188 than it is to token 90210, any more than “cat” being word #4187 in a dictionary would make it related to word #4188. Fixed and unique, but still meaningless: the integer only distinguishes one token from another; it carries no meaning by itself. That’s exactly why a further step is needed before the network can do anything useful with it:

flowchart LR
    TOK["tokens"] -->|"position in<br/>vocabulary"| NUM["integers"]

Generating text reverses this exactly: the model predicts a number, that number maps back to its token, and the token’s text gets appended to the output. This is the mechanism behind the step Part 1 described as picking a token and adding it to the text:

flowchart LR
    NUM["integers"] -->|"vocabulary<br/>lookup"| TOK["tokens"] -->|"append"| TXT["text"]

Numbers in Text vs. Token IDs

“Number” means two unrelated things here, worth telling apart. A digit that appears in the text — like the “87439” in a prompt — is just text: the tokenizer splits it the same way it splits any other text, with no special awareness that it’s a number at all. It might become ["874", "39"], purely because those digit groupings happened to be common in the tokenizer’s training data — not because 874 and 39 mean anything mathematically. The token ID is a second, separate number: the vocabulary assigns each of those pieces its own arbitrary position, same as any other token, completely unrelated to the digits it spells out:

flowchart LR
    N["text: 87439"] -->|tokenizer| P["tokens: 874, 39"]
    P -->|"vocabulary<br/>position"| I["ids: 23891, 118"]

Even a token whose text happens to look like a number follows the same rule: the token spelled "874" doesn’t get token ID 874, or anything close to it. Its ID is just whichever slot that string landed in when the vocabulary was built, with no connection at all to the digits it spells out — as arbitrary as any other token’s ID.

This causes a real, documented problem: because the split is based on training-data frequency, not place value, the same three-digit chunk can land in the hundreds place in one number and the units place in another — "87439" might split as ["874", "39"] in one case and ["87", "439"] in a different one. The model isn’t doing arithmetic on digits the way a calculator would; it’s predicting token sequences, and those tokens don’t line up with place value consistently. This is a real contributor to LLMs making arithmetic mistakes that would surprise a calculator.

Special Tokens

Not every entry in a vocabulary is a piece of ordinary text. A tokenizer’s vocabulary also includes a handful of special tokens that mark structure instead of content — most importantly, an end-of-sequence token that marks where the text stops. When the model predicts that token instead of an ordinary one, generation stops. That’s the actual mechanism behind Part 1’s “until the answer is done”: the model isn’t following some separate stopping rule, it’s just predicting one more token like any other, and that token happens to mean stop. Like any other token, how likely it is to come next is shaped entirely by training, not by a hard-coded rule — nothing stops it from being predicted mid-sentence except that training never gave it a reason to.

Putting It Together

Input tokens and output tokens aren’t two different kinds of thing — they’re positions in the exact same vocabulary, just moving in opposite directions. What the model reads and what it predicts share one fixed list; there’s no separate “output vocabulary.”

That makes generation one repeating loop:

flowchart LR
    IN["your text"] --> ENC["tokens -> integers"]
    ENC --> NET["the network"]
    NET --> DEC["integers -> tokens"]
    DEC --> OUT["new text<br/>↻ appended, fed back<br/>in as the next input"]

Everything between “integers in” and “integers out” is the network itself: each integer becomes a list of numbers — an embedding, still a future post — and that list is exactly the x1…xn a neural network layer already defined. The same weights, bias, and nonlinear step from that post are what actually process every token passing through, all the way until a new integer comes out the other side. The tokenizer turns that integer back into a token and text, and the loop repeats — until the model predicts the end-of-sequence token described above, and generation stops.

Why Tokens Matter

Everything measured about an LLM — how long a conversation it can hold and its input and output limits (both covered in a future post), and the price of a request (covered in a future post too) — is measured in tokens, not words or characters. A short, common word usually costs one token. A rare word or a typo often costs several, splitting into smaller pieces the same way the longer examples above did. Most non-English text costs more too, but for a different reason: the vocabulary was built mostly from one language’s text, so other languages’ ordinary, common words are under-represented in it — not because they’re unusual, just because there was less of them to learn from. The same sentence can end up costing noticeably more tokens in one language than another as a result.

Token counts can be surprising for another reason, too: a word’s leading space is often folded into its own token rather than costing a separate one — " cat" and "cat" are usually two different tokens in the vocabulary, not the same token plus a space. Where the spaces fall changes the count, not just which words appear.

Capitalization can do the same thing. Many tokenizers preserve case rather than ignoring it, so "The" and "the" are also usually two separate vocabulary entries, not one token that happens to look different. A sentence in Title Case or written in all-caps can end up costing more tokens than the identical words in lowercase, for the same underlying reason as the leading-space case: the vocabulary was built from exact character sequences, and a different case is a different sequence.

As a rough sense of scale, English text averages around 4 characters per token — a little under one token per word. That number moves a lot with the specifics above: more punctuation, more unusual capitalization, or more non-English text all push the token count up relative to the character count.


Next: Paper: Byte Pair Encoding — the algorithm behind most tokenizer vocabularies. (Coming soon.)

This post uses “embedding” loosely — covered in a future post.

This post is licensed under CC BY 4.0 by the author.