Multiplying two four-digit numbers is a strange test for a neural network. There is exactly one right answer, every digit of it depends on many digits of the input, and you can make as much practice data as you like. I trained a few hundred small transformers on it from scratch to see what they learn, and in particular what the way a model represents position has to do with it.
The setup
Each problem is a line of text such as 4721*3089=14583169. The model reads it one character at a time and learns to write the answer after the =. It is scored on exact match: the whole product right, on problems it never saw in training. Every setting was trained three times with different random seeds, and I report the median and the range across them.With three seeds no significance test means much. When the ranges of two settings overlap I treat the difference as noise.
There are two scales. A small model (two layers, 64 wide, about 100,000 parameters) trained on a laptop CPU, and a bigger one (six layers, 256 wide, 4.7 million parameters) trained on one RTX 4090, with 200,000 training problems and 30,000 steps of 256 problems each.The GPU runs used bf16 mixed precision. Each took 7 to 28 minutes, five at a time; the whole GPU part cost about $8 of rented time. Small models make the effects easy to see; the larger one shows which of them survive.
How far does a plain transformer get?
Not far. Writing the answer directly, the 4.7-million-parameter model gets 25% (12%–25%) of three-digit by three-digit problems right, 0.1% of four-digit ones and 0% of five-digit ones. The 3 × 3 curve was still rising when training stopped, so this is what 30,000 steps buy rather than a limit.
What is more interesting is where it fails. It is the same at every size:
The lowest digits of a product depend on one or two digit pairs (the units digit is just the units digits multiplied together), and the highest digits depend mostly on the rough size of the numbers. The model gets both nearly perfect. The middle columns are where the most digit pairs meet and the carries pile up, and there it sits at 10%, which is guessing. At 3 × 3 the hundreds digit alone is right 27% (13%–28%) of the time, and it accounts for nearly every miss.
This is also why per-token accuracy flatters these models. At 3 × 3 the model gets 88% of answer digits right, because most digits are easy; it needs all of them.
Does the way the answer is written matter?
Enormously, but not in the way I expected. Writing the digits least-significant first, the order in which you actually compute them, made no consistent difference here. What matters is breaking the problem into steps. If the model first writes the partial products (4721 × 9, 4721 × 80, …) and only then the total, the small CPU model goes from 0.7% to 99% (90%–100%) on 3 × 3.
That fix runs out at 5 × 5. With the same scratchpad the GPU model gets 0.1% (0.1%–0.2%) of 5 × 5 problems right, even though 96% of the tokens it writes are correct. The partial products come out fine. The last step, adding five ten-digit numbers in one go, does not, and once again it is the middle digits of that sum that sit at chance. A scratchpad that also wrote the running total would be the obvious next thing to try.
Why position matters
A transformer layer lets every token look at every earlier token and take a weighted mix of what it finds. The weights come from comparing contents, so on its own the mechanism has no idea which token came first: a sentence and its shuffled copy look alike. Something has to tell the model where each token sits.
For multiplication that information is the whole game. The k-th digit of a product is the sum of digit products with plus carries. To write it, the model has to find, among the digits in front of it, the ones in the right columns, and a digit’s column is a fact about its position: it is the units digit because it is the last digit of its number. Whatever the model uses to represent position is what it has to line columns up with.
I compared six ways of doing that:
| Scheme | What the model gets |
|---|---|
| Learned absolute | A trained vector for each position 0, 1, 2, …, added to the token. The original GPT. |
| Sinusoidal | A fixed pattern of sines and cosines for each position. The original Transformer. |
| RoPE | Queries and keys are rotated by an angle that grows with position, so attention depends on how far apart two tokens are. Llama and most current models. |
| ALiBi | No position vectors; attention to distant tokens is simply penalised, more in some heads than others. |
| NoPE | Nothing. A decoder can still count how many tokens came before, so this is less hopeless than it sounds. |
| Digit place-value | A trained vector for each digit’s column: units, tens, hundreds. Built for arithmetic, close to the “Abacus” embeddings below. |
Does the positional scheme matter?
It depends on how hard the task is, and the answer is more interesting than “scheme X is best”.
When the task is out of reach, nothing helps. Written directly, 4 × 4 stays below half a percent for every scheme, in every seed. When the model is part-way, the seeds are as loud as the schemes. At 3 × 3, also written directly, three RoPE runs that differ only in their random seed score anywhere from 8.7% to 43%. Writing the most significant digit first, there is still a real split: learned positions, RoPE and the digit place-value embedding (9–43% across their seeds) all beat sinusoidal, ALiBi and NoPE (1–10%). Writing the least significant digit first, every scheme’s range overlaps every other’s:
When the format makes the task learnable, the scheme decides how fast training finds the solution. With the partial-product scratchpad, 4 × 4 is within reach, and here the schemes separate sharply. RoPE got every held-out problem right in all three seeds. ALiBi did in two seeds of three and sinusoidal in one; learned absolute positions, NoPE and the digit place-value embedding never got past 20%:
The outcome is not “a bit better” or “a bit worse”. A run either finds the procedure and gets everything right or it does not and gets almost nothing. The training curves show when that happens. Every RoPE run was at 75–81% by step 2,000 and essentially perfect by step 8,000. The ALiBi and sinusoidal runs that succeeded stayed at or below 20% for more than 20,000 steps, then climbed to 90–100% within about 4,000 steps. The rest were still waiting when training stopped.
So at a fixed budget the positional scheme acts less like a dial on accuracy and more like a clock: it sets how long training wanders before it stumbles onto the procedure, and RoPE, which encodes how far apart two tokens are, cut that by about ten times. With longer training some of the stalled runs would probably have jumped too; at 30,000 steps, the difference looks like reliability.
When the columns move, learned absolute positions break. Train on a mix of widths (1 × 1 up to 3 × 3, with the scratchpad) and the same column lands at different positions in different problems. Learned absolute positions cannot cope and stay at 10% (4.4%–15%); sinusoidal, RoPE and ALiBi get essentially every problem right in every seed, the digit place-value embedding gets above 90% in two seeds of three, and NoPE reaches 100% in one:
The hollow dots are the other half of that experiment. 4 × 4 needs one more partial product than anything in training, and no scheme managed a single problem, including the ones that had 3 × 3 perfect. The same was true everywhere I tested one size up: none of the 108 GPU runs got any problem right at one digit wider than it was trained on. Getting a transformer to generalise to longer numbers takes more than a choice of position encoding.
Two tests of what the schemes actually encode
After training I fed each model its held-out problems again, changed in two ways that should not matter to someone who can multiply.
First, I added a constant to every position number, so the problem starts at “position 4” instead of 0. The tokens are untouched.
Learned and sinusoidal positions fail at a shift of one: they learned “the units digit is at position 4”, not “the units digit is the last one”. The other four do not care, because they only see relative distances or no positions at all. That much is true by construction, and it is a fair description of what “relative” buys you.
Second, I inserted extra start tokens before the problem. Now the problem really has moved, the way it would if it appeared after some other text.
Everything collapses, RoPE and ALiBi included. One extra token is enough: RoPE goes from 21% to 1.8% in the least-significant-first runs. Only the digit place-value embedding, which labels digits by their column and not their position, keeps most of its accuracy: 14% (12%–14%) with one extra token against 18% (15%–20%) without, in the most-significant-first runs. The small CPU models behave the same way. “Relative” position encoding makes a model indifferent to how positions are numbered, not to what else is in the context; these models learned to locate digits by counting from the start of the sequence.
Where the trained models look
The heatmaps below show, for one trained model per scheme, how much attention each digit of the answer pays to each token of the problem, averaged over the heads of the last layer. a0 is the units digit of the first number, b2 the hundreds digit of the second.
Try the trained models
Here are two of the 4 × 4 scratchpad models, one with RoPE and one with learned absolute positions, running in your browser. They have the same architecture, the same data and the same training budget.
Are the held-out problems really new?
Less than you would think. The original version of this study split problems into training and test sets by hashing the ordered pair, so 37 × 52 could be held out while 52 × 37 was trained on. With 6,000 training pairs, the mirror image of 73% of the held-out 2 × 2 problems was in the training set. Splitting so that a pair and its mirror always land together drops the 2 × 2 score from 35% (27%–38%) to 18% (12%–25%). Much of what looked like learning was recalling b × a.
Even with that fixed, “held out” means different things at different sizes. At 3 × 3, 200,000 training problems are a quarter of all 810,000 pairs and every held-out problem is one digit away from a training problem. At 5 × 5, most held-out problems are three or more digits away from anything in training.
What changes with Odia digits?
Odia writes its digits as ୦ ୧ ୨ ୩ ୪ ୫ ୬ ୭ ୮ ୯. Swapping them in, with one token per digit, changes nothing: 34% (32%–35%) against 35% (27%–38%) for ASCII on the small 2 × 2 model. A model trained from scratch has never seen what either set of symbols looks like.
What matters is how many tokens a digit becomes. In UTF-8 each Odia digit is three bytes, and the first two are the same for every digit. Most of the tokenizers in wide use fall back to those bytes:
Trained on bytes, a model sees sequences three times as long, two-thirds of them filler. That costs some positional schemes far more than others: NoPE drops from 21% (20%–21%) to 5.3% (4.7%–7.3%), the digit place-value embedding from 32% (32%–36%) to 21% (19%–21%), while learned positions barely move.
The good news is that arithmetic learned in one script carries over. Trained on 6,000 ASCII problems plus 1,000 Odia ones, the model gets 32% (26%–32%) of Odia problems right, about what 6,000 Odia-only problems give: 34% (32%–35%). With 1,000 Odia problems and nothing else it gets 6.7% (2.7%–6.7%).
What kind of extra data helps?
Since the data is free, I tried three ways of making more of it, on the small model.
More distinct problems helps up to a point. With the scratchpad, 3 × 3 needs a few thousand distinct problems: with 1,000 the model memorises them and gets 1.3% (0.3%–2.0%) of new ones right, with 10,000 it gets 96% (95%–97%), and beyond that the curve is flat. Problems whose second number contains a 7 were held out entirely, and on those no run at any data size got above 4.0%: the model learns the rows of the times table it sees.
Adding each training problem’s mirror image (b × a) helps a little, less than the same number of new problems. And spending a quarter of the training steps on the sub-skill, two-digit by one-digit products, made 2 × 2 worse: 35% (27%–38%) without it, 15% (15%–20%) with it. The sub-skill was never the bottleneck; combining partial products was.
What earlier work found
Most of this has been seen before, often more cleanly, and reading the papers after running the experiments was humbling. The ones I found most useful, grouped by what they explain:
Why the middle digits are hard
- Why Can’t Transformers Learn Multiplication? (Bai, Pres, Deng et al., 2025) is the closest match to the first chart above. A two-layer model trained directly on 4 × 4 gets about 1% exact match while getting about 81% of digits right; a twelve-layer model does no better. The outer digits are learned and the middle ones never are, because training settles into a solution with no long-range dependencies. They then reverse-engineer a model that does work (trained with an implicit chain of thought): it stores pairwise digit products in its first layer and reads them back in its second, with digits encoded in a Fourier basis. And they show that adding one auxiliary loss, predicting the running sum, lets the plain two-layer model reach 99%. That running sum is exactly what my 5 × 5 scratchpad was missing.
- Dissecting Multiplication in Transformers (Qiu et al., 2024) finds the same profile at 5 × 5: last digit 100%, top digit above 90%, middle digits around 10%, and accuracy falling as more partial products overlap at a position. Reversing the output alone did nothing for them either.
- Faith and Fate (Dziri et al., 2023) trains much bigger models (GPT-3 fine-tuned, GPT2-XL from scratch) and gets near-perfect in-distribution accuracy that collapses out of distribution, with or without a scratchpad. Their account of the easy digits matches mine: the leading digits of a product follow the leading digits of the inputs and the trailing digits the trailing ones, and models learn those shortcuts first.
Steps, written out or internalised
- Implicit chain of thought and stepwise internalisation (Deng et al., 2023; Deng, Choi and Shieber, 2024). A pretrained GPT-2 Small answering directly gets 29% on 4 × 4 and 1% on 5 × 5; with the partial products and running sums written out it gets 100%. Their striking result is that the written steps can then be removed gradually during training until the model does them silently: GPT-2 Small reaches 99% on 9 × 9 with no reasoning tokens at all.
- Teaching Arithmetic to Small Transformers (Lee et al., ICLR 2024) is the classic small-model study. Reversing the output gives a sharp phase transition for addition, but for multiplication “reverse is not particularly effective”, which matches my 3 × 3 results; detailed step-by-step data works best; and nothing in the paper generalises to longer numbers.
Position
- Positional Description Matters for Transformers Arithmetic (Shen, Bubeck, Eldan et al., 2023) puts its finger on the column problem. A GPT-2-small-sized model trained from scratch to answer directly gets 8% on 4 × 4; zero-padding the numbers, so that a column always sits at the same position, raises that to 73%; padding plus a reversed product reaches 100% through 5 × 5 and near-perfect accuracy up to 12 × 12. Make positions mean columns and the direct task becomes learnable.
- Transformers Can Do Arithmetic with the Right Embeddings (McLeish et al., 2024) introduces “Abacus” embeddings, which index each digit by its place within its own number, the idea behind my digit place-value scheme. Trained on 20-digit addition they reach 99% on 100 digits. Multiplication is near-perfect up to the 15 digits they trained on and collapses just beyond.
- Length Generalization in Arithmetic Transformers (Jelassi et al., 2023): relative position encodings take addition from 5 to 15 digits, but for multiplication every encoding they tried is near 100% in distribution and 0% one digit longer, as in my runs. What did work was “priming”: adding just 50 long examples, 1% of the training set.
- Position coupling (Cho et al., 2024, and a 2025 follow-up) gives digits of the same place value the same position ID. For general N × M multiplication nothing they tried worked without a scratchpad, even in distribution; with a scratchpad plus coupling they got two to three times the trained length. It is the one route past the length wall I found, and it combines both of this post’s main levers.
- Investigating the Limitations of Transformers with Simple Arithmetic Tasks (Nogueira, Jiang and Lin, 2021) is about addition, but it made the same point early: writing explicit place-value tokens (
3 10e1 2) made 60-digit numbers learnable, the surface form of the numbers matters more than model size, and results swung by 20–40 points between random seeds.
Small models, sudden jumps, and how small
- Grokking (Power et al., 2022) and Progress measures for grokking (Nanda et al., ICLR 2023) are about modular arithmetic, not multiplication, but they are the best account of the sudden jumps in the training curves above: a long stretch that looks like nothing is happening, during which the network builds the circuit that generalises. Nanda et al. show a one-layer model doing modular addition with Fourier features and trigonometric identities, the same kind of digit code Bai et al. find in multiplication.
- AdderBoard asks how small a transformer can be and still add two ten-digit numbers. Its leaderboard has hand-coded entries with six parameters and trained ones with a few dozen. Nothing in this post comes close to that kind of economy; a 4.7-million-parameter model is enormous for the task, and still it fails without the right format.
Among these papers I did not find one that compares how reliably each positional scheme trains on multiplication across seeds, which is the effect I saw most clearly.
What I would take away
- Exact match, not loss or per-token accuracy. The models in this post routinely got 85–96% of tokens right while getting almost no answers right.
- Format first. Breaking the problem into steps was the largest effect of anything I tried, and where it stopped working (adding five long numbers at once) the next step to break out was obvious.
- The positional scheme mattered most where the task was learnable but not easy, and mostly as the odds that training finds the solution at all. Three seeds per setting were barely enough to see that; one would have told a misleading story either way.
- None of the schemes gave length generalisation, and “relative” encodings did not make the models robust to text before the problem.
- Check the split. The single largest number I had to revise came from a test set that was not as held out as it looked.
Everything here is a fixed budget on a small model, and a bigger model or longer training would move many of these numbers. The point of the exercise is the shape of the effects, not the specific percentages.