One dial and one number
This is for people who already know what n-back is, have run the protocol, and have watched their n climb without ever finding out what that meant. The mechanism is sound. The instrumentation around it is what's thin.
If you have used Brain Workshop or one of the dual n-back apps, none of what follows will be an attack on them. Brain Workshop in particular is a well-regarded free tool that implements the Jaeggi protocol faithfully, and adaptive load at the edge of your ability is the mechanism with the best evidence behind it. That part is not in dispute here and is not something this app replaces.
The problem is what surrounds it.
What you learn from your n-level
Dual n-back gives you one dial and one number. If the number goes up, you know your n-back got better.
You do not know anything else. Specifically, you do not know whether you got better at holding items, or better at the task — its cadence, its interface, the particular way it cues you. And because the difficulty adapts to you, the number is partly a description of the staircase rather than of you.
That is not a flaw in n-back. It is the absence of an instrument. The training metric was never designed to answer the transfer question, and it cannot be made to.
Items are chunkable. Relations are less so.
Standard n-back asks: is this item the same as the item n back? A position, a letter, a colour.
The trouble is that a list of items is exactly the kind of thing people get good at compressing. Rehearsal strategies, verbal recoding, spatial grouping — what improves with practice is often the chunking, not the capacity. You get better at the representation, and the score reflects it.
Tanren Sensei’s relational n-back asks a different question: is the relation between the current stimulus and the previous one the same as the relation was n steps ago? Bigger-than. Left-of. Darker-than.
The strongest specific thing that can be said about this — and it is a claim about what the task demands, not about outcomes — is that matching a relation n steps back requires holding and updating a structure, not a list. Item identity is never the answer. An item that repeats is a deliberate trap, because a repeat is exactly what a familiarity heuristic fires on.
There is a second task in the same family, applied to reasoning material: does this premise share its structure with the one n back?
Seventeen dials, and you don’t get told which one is turning
In dual n-back, difficulty is n. That is the axis, and everyone knows it.
Here, difficulty is a point in seventeen-dimensional space across the roster. No single exercise carries all seventeen — each declares the ones it supports — and the axes load different things. The ten that do most of the work:
- n-level — how far back the comparison target sits
- Presentation time — truncates the encoding window
- Inter-stimulus gap — removes the rehearsal window
- Lure density — the proportion of trials planted as n±1 matches
- Relational complexity — how many relations must be held simultaneously
- Premise count — reading load, deliberately kept separate from the above
- Premise shuffle depth — how far premises depart from chain order
- Distractor subtlety — whether spotting the negation gets you anywhere
- Negation density — how much of the work is unpicking “not”
- Stimulus domains — how many parallel streams are interleaved
Only one or two are “hot” in any session, and which ones is seeded per session, so you cannot learn which axis is being pushed and pre-allocate effort to it. The rest are held or jittered slightly — not decoration, because if everything except the hot axis sat at exactly last session’s value, the hot axis would be identifiable by elimination within a few trials.
Two of those axes are worth pulling out.
Lure density is an explicit, adapted difficulty dimension rather than a fixed sprinkle. Lures — matches at the wrong distance — are what force genuine positional binding instead of a vague sense of “I’ve seen this”. Most implementations either omit them or fix them at a constant rate.
Premise count is adapted, but it is deliberately not the reasoning-load axis. Ten premises that chain linearly are easier than four that must be integrated pairwise, so relational complexity carries the reasoning and premise count carries the reading. Most trainers scale difficulty by adding items, which grows one and not the other.
A note on “quad”
Quad n-back means four simultaneous streams — more channels, same question per channel.
That is a legitimate way to make the task harder, and it is not the axis this app pushes. The direction here is to deepen the relation rather than multiply the number of things you are tracking. Both increase load. They do not load the same thing, and neither is obviously the better lever.
The part that is actually missing elsewhere
Everything above is a design argument, and design arguments are cheap. Any of it could be wrong.
Which is the real difference: this app ships four tests you never train on, at fixed difficulty, on a fixed schedule, never fed to the adaptive engine — and one of those four measures an ability the app deliberately does not train, so you can tell a specific training effect apart from simply getting better at sitting tests.
That is the piece the n-back ecosystem has never had. Not because anyone was hiding anything, but because a training tool’s natural output is a training metric, and nobody built the held-out instrument to sit beside it.
So the honest pitch is not “n-back is broken”. It is this:
Dual n-back gives you one dial and one number. If that number goes up, you know your n-back got better. You do not know anything else. This app keeps the mechanism — adaptive load at the edge of your ability — widens what is being loaded, and then hands you an independent way to check whether it did anything.
Whether it did do anything is not something anyone gets to claim in advance, including us. What to expect →