The test we hope doesn't improve
Every trainer can show you a number going up. Almost none can tell you why. The cheapest way to find out is to measure something you expect to stay still — and then pay the price of not training it.
Suppose you train for three months and your matrix-reasoning score rises. What happened?
There are at least four answers, and they are not easy to separate:
- The training worked, and the ability it targets genuinely improved.
- You got better at that particular test — its format, its quirks, its rhythm.
- You got better at sitting tests in general. You are calmer, you know how long you have, you have stopped second-guessing.
- You were having a better day. You slept well. The first sitting was after a bad week.
Only the first is the one anyone is selling. The other three produce an identical rising line.
The cheapest instrument that separates them
You cannot run a control group on yourself. There is no version of your life where you skipped the training, and you cannot blind yourself to whether you did it.
But you can do something almost as useful: measure a second ability that the training deliberately does not touch, on the same schedule, under the same conditions, with the same repeated exposure.
Explanations 2, 3 and 4 all predict that both measures move. Familiarity with test formats, general test-taking composure, and simply having a good day are not specific to the trained ability — they lift everything in the room. Only explanation 1 predicts that the trained measures move and the untrained one does not.
So the pattern, not the number, carries the information:
- Everything rises together. The likeliest reading is that you got better at taking tests. Not nothing — but not what you trained for.
- The trained measures move, the untrained one doesn’t. That is the shape worth taking seriously.
Tanren Sensei’s control is 3D mental rotation: a reference figure built from unit cubes, and four candidates, of which exactly one is the same object turned. It is administered exactly like the three trained probes and reported below a divider under its own heading, never averaged into anything. A control folded into a composite stops being a control.
Making the control honest is most of the work
A control measure that can be beaten by a trick is worse than no control, because it produces a confident-looking number that means nothing. Two problems had to be solved before this one was worth reporting.
Chirality has to be decided exactly
The question “is this the same object turned, or its mirror image?” is only well posed for a figure that is genuinely chiral. If some rotation already maps a figure onto its own mirror image, then the mirror is the object turned, and the item has two correct answers.
This is decided with no tolerance and no heuristics. The cube’s symmetry group has 48 elements, 24 of them proper rotations. Taking the lexicographic minimum over each group gives a canonical form, which turns “is this a rotation of that?” into string equality. Achiral figures are rejected at generation rather than shipped and hoped about.
The verifier that does this was written and hand-tested before the generator existed — against a true rotation, a mirror image, a figure with the same cube count but different connectivity, and deliberately planted bad items with two valid answers and with none.
The option set must not leak the answer
Here is the trap that nearly got through. If the correct answer were the only figure of its handedness among three mirrors, then “pick the odd one out” solves every item with no mental rotation whatsoever.
So every option set is exactly {S, mirror S, T, mirror T} for two distinct
chiral shapes, each independently rotated and dealt to a slot by a uniform
permutation. Grouped by shape, the classes are {2, 2}; grouped by shape and
handedness, {1, 1, 1, 1}. Neither grouping favours a slot.
The two shapes also have to share an invariant fingerprint — cube count, bounding box, and neighbour-degree sequence, all of which survive rotation and reflection. That constraint exists because an automated attacker found the leak it closes: with unmatched bounding boxes, “pick the option whose box matches the reference’s” cuts four options to two for free.
An audit runs seven different attackers over 2,000 items. All of them score at chance — 24.4% to 25.5% against a 25% baseline — and the correct-slot χ² is 0.51 against a critical value of 16.27.
The floor is 50%, not 25%
Four options, and yet a score has to be read against half. This is structural, not a bug waiting to be fixed.
A solver who works out which two options share the reference’s shape, ignoring handedness, is left choosing between an object and its mirror. Can shape be made non-diagnostic? No, and the argument is short:
- An option sharing the reference’s shape and handedness is a rotation of the reference — that is what those words mean. So a shape group of size k contains exactly one correct option and k−1 mirrors.
- That group therefore has handedness ratio 1 : (k−1).
- For k ≥ 3 the ratio is lopsided, and a blind attacker who never even looks at the reference just picks the minority handedness in each group.
- Balanced handedness therefore forces k = 2.
- And k = 2 is a two-alternative choice for a shape-matching solver.
Blind-proofness and a 25% floor are incompatible. Blind-proofness wins: an exploit that requires no reasoning at all is worse than a floor higher than the option count implies. Measured over 2,000 items, a pure shape-matching attacker scores 48.9%.
The floor is conservative on purpose. Shape matching is not free for a human — telling two cube assemblies apart up to rotation is real spatial work, it just stops before resolving handedness. The true human floor is somewhere in [25%, 50%], and only human data could locate it. So scores are reported against the pessimistic end: the results bar runs 50–100% rather than 0–100%, and anything at or below 50% is labelled no signal rather than a decline.
What this control costs
Here is the part that is genuinely expensive, and it should be stated rather than buried.
Spatial training has the strongest transfer evidence of anything this app could plausibly train. Making mental rotation the control is a decision not to train the ability most likely to actually improve — in exchange for being able to tell whether the rest of the programme does anything at all.
That is a real cost, taken knowingly. It is also reversible: once the first comparison has produced an answer, spatial training can be added, with a different untrained ability serving as the control.
And one caveat that cuts the right way
The relational n-back uses 2D orientation as one of its stimulus attributes. Imagining an object turned in depth is a different operation from comparing planar orientations — but they are related, so the app does not train zero rotation-adjacent ability.
This makes the control conservative rather than invalid. If rotation rises partly because of orientation training, the gap between trained and control measures shrinks, so the comparison understates how specific any training effect is. It cannot manufacture specificity that is not there.
That sentence appears in the app too, next to the numbers, rather than only here.