What a Neural Network Actually Is — LLMs Part 2
The Definition
A neural network is a function: it takes in numbers and produces numbers out. Each number going in is called an input — nothing more than a plain number. A network might take two inputs, or thousands.
Its name is a loose analogy to brain cells, but the math underneath is ordinary multiplication and addition, with one twist. That twist is what makes a network more than a calculator, and it gets its own section below.
The network itself is built from smaller steps called layers. Each layer turns one set of numbers into a different set, and one layer’s output becomes the next layer’s input:
flowchart LR
IN["numbers in"] --> L1["layer 1"] --> L2["layer 2"] --- DOTS["⋯"] --> LN["layer N"] --> OUT["numbers out"]
style DOTS fill:transparent,stroke:none
style IN fill:transparent,stroke:none
style OUT fill:transparent,stroke:none
What a Layer Is Made Of
A layer is made of neurons — its basic units. This post covers the most common kind, a fully connected layer: each neuron looks at every number the layer receives, and produces exactly one number in return, so a layer’s neuron count sets its output count regardless of how many numbers came in.
Scale makes this concrete: a layer with 1,000 inputs and 100 neurons does not split those 1,000 numbers across the neurons. Every neuron sees all 1,000 of them, so connecting them takes 1,000 × 100 = 100,000 weights — one weight per neuron per input.
Here’s a layer with n inputs and m neurons. All n inputs sit inside one box, and that box connects to every neuron:
%%{init: {'flowchart': {'nodeSpacing': 15, 'rankSpacing': 50}}}%%
flowchart LR
INPUTS["inputs<br/><br/>x1<br/>x2<br/>⋮<br/>xn"]
INPUTS -->|"all n inputs"| N1["neuron 1<br/>weights w1<br/>bias b1"]
INPUTS -->|"all n inputs"| N2["neuron 2<br/>weights w2<br/>bias b2"]
INPUTS ~~~ DOTS["⋮"]
INPUTS -->|"all n inputs"| Nm["neuron m<br/>weights wm<br/>bias bm"]
N1 --> y1["y1"]
N2 --> y2["y2"]
DOTS ~~~ yDots["⋮"]
Nm --> ym["ym"]
style DOTS fill:transparent,stroke:none
style yDots fill:transparent,stroke:none
Every arrow carries all n inputs — every neuron reads every one of them, identically. What differs is what each neuron does with them: its own weights, labeled w1, w2, and so on inside its box, alongside its own bias. Neuron 2’s weights are never shared with neuron 1’s, even though both read all n inputs. Across all m neurons, that’s n × m weights in total, plus one bias per neuron (m more).
For a network’s first layer specifically, x1 through xn is one token’s embedding (Part 5).
Two adjustable numbers make up each neuron’s share of that count. A weight is one adjustable number per input, multiplied against its matching input — so a neuron’s weight count always equals its input count, no more, no fewer, and that count is fixed once the network is built. A bias is one more adjustable number, unique to that neuron, added in after the weights are multiplied. Multiply, then add the bias — that’s as far as plain arithmetic gets you, which is exactly where the twist comes in.
What a Neuron Is Made Of
A neuron is made of its weights and bias, covered above, plus one more piece. So, what’s the twist? After multiplying and adding, a neuron applies a nonlinear step: a fixed rule applied to a single number. A common one, called ReLU (Rectified Linear Unit), keeps the number if it’s positive and otherwise uses zero. Real LLMs often use smoother variants, such as GELU (Gaussian Error Linear Unit), but the job is the same.
Weights and bias are the two per-neuron pieces from the previous section — every neuron has its own. The nonlinear step is different: it’s not a per-neuron adjustable number at all, just a fixed rule (like ReLU, defined above) applied the same way to every neuron in the layer. Training changes weights and biases; it doesn’t change which rule the nonlinear step uses.
This step earns the name “twist” because of what happens without it. Multiplying and adding alone can only ever produce a weighted sum — stack as many such layers as you like, and the whole network still collapses mathematically into one single weighted sum of the original inputs. The nonlinear step breaks that collapse, which is what lets stacked layers compute more than a single layer could.
Put the three pieces together, and a neuron does exactly three things in order: multiply each input by its weight, add up the results along with its bias, then apply the nonlinear step. That produces the neuron’s one output number. Collect every neuron’s output, and that’s the layer’s output. As a formula, neuron i’s output is:
1
yi = f( (x1 × wi1) + (x2 × wi2) + ... + (xn × win) + biasi )
wi1 is neuron i’s weight for input x1, biasi is neuron i’s bias, and f is the nonlinear step. i ranges from 1 to m — one such formula per neuron in the layer.
Together, all of a network’s weights and biases are called its parameters. An LLM with billions of parameters is a neural network with billions of these adjustable numbers.
Altogether, one neuron’s full path looks like this:
flowchart LR
X["inputs<br/>x1 … xn"] -->|"weights<br/>(unique to<br/>this neuron)"| SUM(["multiply & add,<br/>then + bias<br/>(unique to<br/>this neuron)"])
SUM --> F["nonlinear step<br/>(same fixed rule<br/>for every neuron<br/>in the layer)"]
F --> Y["y"]
One Fixed Size, Any Length of Text
A neuron’s fixed weight count raises a natural question: how does the same network handle a short prompt and a long document alike? Not by changing n. Text of any length is first broken into a sequence of same-size pieces, called tokens (Part 3), and each token is turned into a fixed-size list of numbers, called an embedding (Part 5) — the same length every time. A layer’s neurons always take in one such fixed-size piece; what varies is how many pieces there are, not the size of any single one, and every piece runs through the same weights independently and at the same time, not one after another. Combining information across however many pieces are present is a separate mechanism, attention (Part 7) — not something this post’s layers do on their own.
Not Every Layer Is Fully Connected
A fully connected layer isn’t the only design. A convolutional layer connects each neuron to only a small, local patch of its input, instead of the whole thing, and reuses that same small set of weights across every patch — well suited to images, where nearby pixels matter more than far-apart ones. A recurrent layer processes one piece of a sequence at a time, feeding its own previous output back in alongside the next piece — built for data where order matters, like reading text one word after another.
LLMs use neither of these to combine information across a sequence. They use attention (Part 7) instead, which lets every position look directly at every other position at once — without a fixed local patch, and without processing one piece at a time in order.
How Layers Chain Into a Network
Layer 1’s outputs (y1 … ym) become layer 2’s inputs, and layer 2 works exactly the same way as layer 1 — same neuron mechanics, its own separate weights and biases:
flowchart LR
X["network inputs<br/>x1 … xn"] -->|"n weights<br/>per neuron"| L1["layer 1<br/>m neurons"]
L1 -->|"y1 … ym<br/>m weights<br/>per neuron"| L2["layer 2<br/>k neurons"]
L2 --> Z["outputs<br/>z1 … zk"]
Because layer 2 receives m numbers, each of its k neurons needs m weights of its own — one per incoming number, exactly like layer 1. Nothing is shared between layers; each has its own parameters. Layer N repeats this same pattern until the last layer’s output is the network’s final answer.
This is where the twist pays off: stacking layers with the nonlinear step lets a network approximate very complex functions. In theory, even one very wide layer could do this, but in practice stacking gets there far more efficiently.
None of this is useful yet, though — the network’s weights start out as random numbers. Training is the process of adjusting those weights until the network’s output gets closer to what you want. We’ll cover training itself in Part 12, and work through its arithmetic by hand in Part 16.
Worked Example
Take one neuron with 2 inputs:
flowchart LR
X1["x1 = 0.8"] -->|"× w1 = 0.7"| SUM(["add up<br/>plus bias −0.9"])
X2["x2 = 0.6"] -->|"× w2 = 0.5"| SUM
SUM -->|"−0.04"| ACT["nonlinear step<br/>keep if positive, else 0"]
ACT --> OUT["output = 0"]
The same numbers, worked out:
1
2
3
4
5
6
7
8
9
10
11
x1 = 0.8, w1 = 0.7
x2 = 0.6, w2 = 0.5
bias = -0.9
sum = (x1 × w1) + (x2 × w2) + bias
= (0.8 × 0.7) + (0.6 × 0.5) + (-0.9)
= 0.56 + 0.30 + -0.9
= -0.04
nonlinear step (keep if positive, else 0):
= 0
Change the inputs, keep the weights and bias fixed:
1
2
3
4
5
6
x1 = 0.9, x2 = 0.9
sum = (0.9 × 0.7) + (0.9 × 0.5) + (-0.9) = 0.18
nonlinear step: 0.18 is positive, keep it
= 0.18
Same weights, same bias, different inputs, different output. The weights and bias are what the network “knows” — they don’t change while the network is running, only during training. A neuron in a real LLM typically has thousands of inputs, not two.
Why Neural Networks Matter
Every LLM is built from exactly this: layers of neurons, each multiplying its inputs by weights, adding a bias, then applying the nonlinear step. Part 1’s next-token predictor and Part 7’s attention mechanism are both computed the same way underneath — numbers multiplied by adjustable weights, nothing more exotic than what this post covered.
Next: Tokens, Not Words — how text becomes the numbers a network can use. (Coming soon.)
This post uses “training” loosely — covered in Part 12. It uses “token” loosely — covered in Part 3.