Fine-tuning and LoRA adapters, learned by tossing paper into a bin. A real little neural network throws; real training, real LoRA, real physics with wind.
| on | adapter | params | file |
|---|
You The habits in your arm: how far you swing, when you let go.
The model Numbers that turn input into output. 1,282 here, billions in an LLM. Learning means changing them.
You You look at the bin (and the fan) and throw.
The model Room numbers go in, flow through the layers, a throw comes out. Also called inference or a prediction.
You You see how far you missed.
The model One number for how wrong the output was. Here: miss distance squared. Training means making it small.
You “Too long and a bit right” tells you which way to change.
The model For every weight: which direction lowers the loss, and how strongly. Backpropagation computes it.
You How much you change after a miss. Too much and you overshoot; too little and it takes all day.
The model The step size on the gradient. The single most important training setting.
You Throw, look, adjust: once.
The model One forward pass, loss, gradient, weight update.
You Your misses getting smaller over the afternoon.
The model Loss plotted against steps. The first thing to watch in any training run.
You Throw five, then adjust on all five, so one fluke doesn't fool you.
The model Average the gradient over several examples, then make one update.
You Working through every practice spot once.
The model One full pass over the training dataset.
You Long three times running? Then adjust more confidently.
The model Remembers recent gradients (momentum) and sizes each weight's step separately.
You A coach shows a good throw for each spot.
The model Examples paired with the right answer. Learning from them is supervised fine-tuning.
You Years of office throwing before today.
The model The huge first training run. What comes out is the base (pretrained) model.
You Someone switches on a fan. You start from your skill and practise the new condition.
The model Continue training the base model on a small, specific dataset.
You You practise only with a wet ball, and your normal throw gets worse.
The model New training overwrites skills the new data didn't cover.
You Keep your technique; add one small habit: lean into the wind.
The model Freeze every weight. Train a small low-rank add-on B·A beside them.
You Can you hit spots you never practised? If only the practice spots, you memorised them.
The model Score on held-out examples. A good training score with a bad held-out score is overfitting.
You How you practise: how many throws, how big your adjustments.
The model Settings you choose, not learned: learning rate, steps, batch size, rank.
Memory counts weights, gradients and Adam state only (16 bytes per trained parameter with mixed precision); activations and batch size add more. Layer shapes are from the models' published configs.
| The thrower's arm after 3,000 office tosses | A pretrained model: billions of weights learned from a huge general dataset |
| A room with a fan, or a wet paper ball | A new task or domain: your company's documents, a support tone, a JSON output format |
| Demonstration throws that land in the bin | The fine-tuning dataset: prompts paired with the answers you want |
| Retraining every muscle | Full fine-tuning: every weight gets a gradient and optimizer state |
| Forgetting how to throw in the office | Catastrophic forgetting: new training overwrites old skills |
| A small clip-on correction; the arm stays locked | A LoRA adapter: ΔW = B·A added beside the frozen weight W |
| The clip-on starts slack | B is initialised to zero, so training starts from the base model's behaviour |
| How many independent corrections the clip-on can make | Rank r: the adapter's inner width. Trainable numbers = r × (inputs + outputs) per matrix |
| How hard the clip-on pulls | α / r: the scale on B·A, so changing r doesn't change the update size |
| Which joints get a clip-on | Target modules: q_proj, v_proj, the MLP layers… |
| Swapping clip-ons per room | Serving many adapters on one base model; each is only megabytes |
| Gluing the clip-on in | Merging: W ← W + B·A, so inference costs nothing extra |
| A blurrier arm stored cheaply, plus a sharp clip-on | QLoRA: a 4-bit base with full-precision adapters trained on top |
| Hitting fresh rooms, not just the demo rooms | Evaluating on held-out data; low loss with few hits is overfitting |