The Neuron
Many neural networks are built from a small mathematical unit called a neuron. Think of deciding "should I order pizza tonight?" You weigh a few facts: how hungry you are, how much money you have. Some facts matter more to you than others — that "how much it matters" number is called a weight. You also have a baseline craving no matter what — that's the bias. Add it all up, and a nonlinear activation shapes what is passed to the next layer.
decision = squash( hunger × importance₁ + budget × importance₂ + baseline )
That weighted-sum pattern is one building block. Large language models combine billions of learned parameters across attention and feed-forward layers rather than behaving as a pile of literal yes-or-no decisions.
Layers & the Forward Pass
One neuron makes one tiny judgment. Put a row of them side by side and you get a layer — a panel of judges all looking at the same facts, each caring about different things. Stack several panels and you get a company hierarchy: junior analysts spot simple patterns, managers combine those into bigger ideas, and a director makes the final call.
Real example — a network looking at a photo: the first layer notices edges and blobs of color. The next layer combines edges into shapes like ears and whiskers. A deeper layer combines those into "this looks like a cat." Nobody programmed "ear" or "whisker" — the layers figured out those stepping stones on their own.
One full trip of information from input to answer is called the forward pass. Press play:
Training: Loss & Gradient Descent
A brand-new network has all its dials set randomly, so its answers are garbage. Training is like learning darts blindfolded, with a friend shouting how far you missed by:
1) Throw (make a prediction). 2) Friend shouts your miss distance — that score is the loss. 3) Figure out which part of your throw caused the miss — arm too high? too much force? That blame-tracing is backpropagation. 4) Adjust each part slightly and throw again. Repeat millions of times until you're hitting bullseyes.
The adjust-and-retry part has a fancy name — gradient descent — but it's just walking downhill in fog: you can't see the bottom of the valley, but you can feel which way slopes down, so you take a step that way. How big a step? That's the learning rate. Try it:
Tokens & Embeddings
Networks only eat numbers — so how do they read? Two steps.
Step 1 — chop: text is snapped into tokens, LEGO-brick pieces of words. Common words are one brick ("the", "cat"); rare words get built from smaller bricks ("unbelievable" → un + believ + able). Each brick has an ID number in a big fixed dictionary. This is why AI services charge "per token" — it's per LEGO brick.
Step 2 — locate: each token ID looks up its embedding — a long list of numbers that works like GPS coordinates on a map of meaning. On this map, "king" sits near "queen", "pizza" near "burger", and "happy" far from "furious". The model learned this map itself, just from reading. Finally, since a bag of coordinates has no order, each token gets a seat number (positional encoding) so the model knows "dog bites man" ≠ "man bites dog".
Self-Attention
Read this sentence: "The robot picked up the ball because it was heavy." What does "it" mean — the robot or the ball? Your brain answered instantly by glancing back at earlier words and weighing which fits. Attention is exactly that glance-back, done with math. For every word, the model asks: "which other words in this sentence help me understand this one?" — and borrows meaning from them, in proportion to how relevant they are.
Under the hood it works like a room full of people wearing name tags. Each word shouts a question ("I'm a pronoun — who's a heavy object around here?"), reads everyone's name tag to see who's relevant, then listens to what the relevant ones actually say. In the papers these three are called Query, Key, and Value — but it's just shout, scan tags, listen.
The Transformer Block
Attention alone isn't enough — it gets wrapped with three helpers into a block, one factory station on an assembly line. The best mental model: attention is the group discussion (words share info with each other), and the feed-forward part is individual homework (each word goes off and thinks alone about what it just heard). Do discussion → homework, and you've got one block. Stack that station 30–80 times and you've got a frontier model.
Tap each part of the schematic to see its job in plain words:
The Full Model & How It Writes
Here's the secret that surprises everyone: ChatGPT, Claude, Llama — under the hood, each is a very fancy autocomplete running in a loop. The recipe: words → coordinates (Ch. 04) → through the stack of blocks (Ch. 06) → out comes a probability for every word in the dictionary as the next word. Pick one, glue it onto the sentence, run the whole thing again. Every AI answer you've ever read was written one token at a time this way.
How it picks is controlled by temperature — an adventurousness dial. Low: always take the safest word (reliable, a bit boring). High: give unlikely words a real chance (creative, occasionally unhinged). Try it:
Model Adaptation & Grounding
Fine-tuning changes model behavior
A freshly pretrained model is like a brilliant medical graduate — enormous general knowledge, zero bedside manner, no specialty. Fine-tuning is the residency: continue its training on a smaller, focused set of examples so it learns a specific job — answer like a support agent, write in your brand's voice, output your exact report format. Three ways to do it, differing in how much of the brain you retrain:
Full fine-tune
LoRA
QLoRA
| Method | Dials retrained | Hardware needed (8B model) | When to use |
|---|---|---|---|
| Full fine-tune | 100% | high-memory or multiple GPUs | When broad weight updates justify the cost |
| LoRA | a small subset | often one capable GPU | Efficient task or behavior adaptation |
| QLoRA | a small subset | less memory than standard LoRA | Adapter training under tighter memory limits |
| Prompting / RAG | 0% | none when using an API | Evaluate before training |
A simplified model-development pipeline
What "RAG" means — in one picture
Models only know what they read during training, and they can't see your files. RAG (retrieval-augmented generation) is the workaround, and it's beautifully unglamorous: when a question comes in, search your own documents, and paste the relevant pages into the prompt along with the question. The model answers using what's in front of it — like an open-book exam instead of a memory test.
Your Path
Build intuition (2–3 wks)
Watch, then rebuild the core ideas on this page in simple Python.
Understand transformers (2–3 wks)
Follow a small GPT implementation and compare the code with a visual walkthrough of the architecture.
Adapt something real (2–3 wks)
Start with prompting and retrieval, then use LoRA or QLoRA when evaluation shows that training is warranted.
Ship and evaluate (ongoing)
Build a small end-to-end product with structured outputs, tools, evaluations, observability, latency and cost budgets, privacy controls, and defenses against prompt injection.
Jargon Decoder
Every term from this guide (plus a few you'll meet on day one of any tutorial), in one table. Screenshot this.
| When you hear… | Think… |
|---|---|
| Weight / parameter | One importance dial. "7B model" = 7 billion dials. |
| Bias | A head start added before deciding. |
| Activation | The squash step that turns a total into a firing strength. |
| Layer | A team of decision-makers looking at the same input. |
| Forward pass / inference | Running data through the model to get an answer. "Inference" = using the model (vs. training it). |
| Loss | The "how wrong was I" score. Training only ever shrinks this. |
| Gradient descent | Walking downhill in fog toward "less wrong". |
| Learning rate | Step size. Too small = crawling, too big = kangaroo. |
| Backpropagation | Blame-tracing: which dial caused the miss? |
| Overfitting | Memorizing the practice exam instead of learning the subject. |
| Token | A LEGO brick of text, ~¾ of a word. AI is priced per brick. |
| Embedding | GPS coordinates on a map of meaning. |
| Attention | Each word glancing at the others to work out what matters. |
| Transformer | The architecture: group discussion + individual homework, stacked many times. |
| Context window | How much text fits on the model's desk at once. |
| Temperature | Adventurousness dial when picking the next word. |
| Pretraining | Learning broad statistical patterns from a large data mixture, often through next-token prediction. |
| Fine-tuning | The residency: extra training for a specific job. |
| LoRA / QLoRA | Sticky notes on a frozen (or compressed) textbook — cheap fine-tuning. |
| RAG | Open-book exam: search your documents, paste the pages into the prompt. |
| Hallucination | Confidently making things up. Reduced by RAG, never fully cured. |
| GPU / VRAM | The specialized chip that does the math / its onboard memory — the scarce resource everything is measured against. |
| Agent | A model in a loop that can use tools (search, code, your APIs) to finish multi-step tasks. |
References & Further Reading
The diagrams and analogies are original explanations informed by these foundational papers. They are starting points, not a claim that every current model uses the same implementation.
- Vaswani et al., Attention Is All You Need (2017).
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models (2021).
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs (2023).
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020).