Paper Toss LoRA

Paper Toss LoRA

Fine-tuning and LoRA adapters, learned by tossing paper into a bin. A real little neural network throws; real training, real LoRA, real physics with wind.

You throw

Look at the bin, set your aim and strength, and throw. Then adjust the sliders yourself and throw again.

Fine-tune on demonstration throws

Method
Clip adapters onto (target layers)

Adapter shelf

throwing with:
onadapterparamsfile

Your loss curve: how far each throw missed

you adjusted the slidersgradient descent adjusteddashed line: close enough to go in

What you just did, in fine-tuning words (cards light up as you play)

01Weights

You The habits in your arm: how far you swing, when you let go.

The model Numbers that turn input into output. 1,282 here, billions in an LLM. Learning means changing them.

02Forward pass

You You look at the bin (and the fan) and throw.

The model Room numbers go in, flow through the layers, a throw comes out. Also called inference or a prediction.

03Loss

You You see how far you missed.

The model One number for how wrong the output was. Here: miss distance squared. Training means making it small.

04Gradient

You “Too long and a bit right” tells you which way to change.

The model For every weight: which direction lowers the loss, and how strongly. Backpropagation computes it.

05Learning rate

You How much you change after a miss. Too much and you overshoot; too little and it takes all day.

The model The step size on the gradient. The single most important training setting.

06Training step

You Throw, look, adjust: once.

The model One forward pass, loss, gradient, weight update.

07Loss curve

You Your misses getting smaller over the afternoon.

The model Loss plotted against steps. The first thing to watch in any training run.

08Batch

You Throw five, then adjust on all five, so one fluke doesn't fool you.

The model Average the gradient over several examples, then make one update.

09Epoch

You Working through every practice spot once.

The model One full pass over the training dataset.

10Optimizer (Adam)

You Long three times running? Then adjust more confidently.

The model Remembers recent gradients (momentum) and sizes each weight's step separately.

11Training data

You A coach shows a good throw for each spot.

The model Examples paired with the right answer. Learning from them is supervised fine-tuning.

12Pretraining → base model

You Years of office throwing before today.

The model The huge first training run. What comes out is the base (pretrained) model.

13Fine-tuning

You Someone switches on a fan. You start from your skill and practise the new condition.

The model Continue training the base model on a small, specific dataset.

14Catastrophic forgetting

You You practise only with a wet ball, and your normal throw gets worse.

The model New training overwrites skills the new data didn't cover.

15LoRA adapter

You Keep your technique; add one small habit: lean into the wind.

The model Freeze every weight. Train a small low-rank add-on B·A beside them.

16Validation & overfitting

You Can you hit spots you never practised? If only the practice spots, you memorised them.

The model Score on held-out examples. A good training score with a bad held-out score is overfitting.

17Hyperparameters

You How you practise: how many throws, how big your adjustments.

The model Settings you choose, not learned: learning rate, steps, batch size, rank.

The thrower's brain: 4 inputs → 32 → 32 → 2 outputs (yaw, speed)

frozen weights (locked) being trained adapter B·A cell colour: negative · paper = 0 · positive

Training, live

Loss on the demo throws (log scale)
Hits on fresh rooms: this room vs the office

How much rank does the fan need?

LoRA on layers 1+2, α = 4: fan-room hits by rank, and trainable numbers
Full fine-tuning's own change to each layer, split into directions (singular values)

The same arithmetic on real LLMs

Memory counts weights, gradients and Adam state only (16 bytes per trained parameter with mixed precision); activations and batch size add more. Layer shapes are from the models' published configs.

Paper toss ↔ fine-tuning a language model

The thrower's arm after 3,000 office tossesA pretrained model: billions of weights learned from a huge general dataset
A room with a fan, or a wet paper ballA new task or domain: your company's documents, a support tone, a JSON output format
Demonstration throws that land in the binThe fine-tuning dataset: prompts paired with the answers you want
Retraining every muscleFull fine-tuning: every weight gets a gradient and optimizer state
Forgetting how to throw in the officeCatastrophic forgetting: new training overwrites old skills
A small clip-on correction; the arm stays lockedA LoRA adapter: ΔW = B·A added beside the frozen weight W
The clip-on starts slackB is initialised to zero, so training starts from the base model's behaviour
How many independent corrections the clip-on can makeRank r: the adapter's inner width. Trainable numbers = r × (inputs + outputs) per matrix
How hard the clip-on pullsα / r: the scale on B·A, so changing r doesn't change the update size
Which joints get a clip-onTarget modules: q_proj, v_proj, the MLP layers…
Swapping clip-ons per roomServing many adapters on one base model; each is only megabytes
Gluing the clip-on inMerging: W ← W + B·A, so inference costs nothing extra
A blurrier arm stored cheaply, plus a sharp clip-onQLoRA: a 4-bit base with full-precision adapters trained on top
Hitting fresh rooms, not just the demo roomsEvaluating on held-out data; low loss with few hits is overfitting