Tanren Sensei

Articles8 min read

Five ways the puzzles leaked their own answers

You do not need to notice a shortcut to use one. Hundreds of trials of implicit learning will find any signal that exists, and a score that rises because the generator has a tell is indistinguishable from a score that rises because you improved.

Verifying that a puzzle has exactly one correct answer stops it being broken. It does nothing to stop it being guessable.

These are different failures. A broken item has no defensible answer. A guessable item has exactly one — and also has some surface property that correlates with it. Count the premises, count the negations, notice where the longest word sits, and you can do better than chance without engaging with the problem at all.

The dangerous part is that nobody has to notice. Implicit learning is very good at picking up statistical regularities that never reach awareness. If a cue exists, hundreds of trials will find it, and the resulting improvement looks exactly like the improvement you wanted.

For an app whose entire claim is that it measures honestly, that is the failure mode that matters most.

Turning “is there a tell?” into a number

Every build measures the mutual information between each surface feature of a problem and its answer class. Features include premise count, negation count, distinct terms, term occurrence counts, relation-token matches, term positions, and chain depth. The measurement runs over at least 600 mixed-class items per family, with Miller–Madow bias correction.

If any single feature carries more than 0.05 bits, the build fails.

That threshold is not arbitrary. The answer classes sit at roughly 40/35/25, giving about 1.56 bits of total uncertainty, so 0.05 bits is around 3% of it. By Fano’s inequality, a cue that weak lifts even a perfect cue-follower only a few points above base rate — below the noise floor of a twenty-problem block. It also sits comfortably above the corrected estimator’s residual error at N=600, which means a failure is a real leak rather than sampling noise.

A separate test plants a deliberate leak and asserts the audit catches it. Without that, a green result would only prove the audit runs.

The five it has caught

1. The spare term that only appeared on one answer class — 0.81 bits

Five chain families shared a generator that introduced an extra, unused term in some items. It turned out that term appeared almost exclusively on items whose answer was indeterminate.

Nothing about the reasoning changed. The item just had one more noun in it, on the occasions when the answer was “can’t tell”. At 0.81 bits, a solver who noticed only that would have been most of the way to the answer.

2. Parity in exact-wording counts — 0.07 bits

In the same/opposite family, the way statements were counted separated the answer classes structurally, through parity. Odd and even counts fell differently across classes.

Small — 0.07 bits, barely over threshold — and completely invisible to inspection. You would not find this by reading items. You find it by measuring.

3. The existential premise that marked its own conclusion — 0.25 bits

In syllogisms, the presence of an existential premise marked the items whose correct answer was an entailed “some”. The grammar of the premise set was telling you the shape of the conclusion before you evaluated anything.

4. The answer, printed on the item — 0.93 bits

The worst one. Statement-level negation in the propositional family rendered the phrase “it is not the case that…” on exactly the contradicted conclusions.

0.93 bits out of 1.56. The item was, in a meaningful sense, labelled. Anyone who noticed the phrasing — and plenty of people would, eventually, without being able to say why — was reading the answer key.

5. A leak wearing a green checkmark — 0.0486 bits

This is the instructive one.

After the syllogism fix, the audit passed. But the worst feature sat at 0.0486 bits against a threshold of 0.05 — 97% of the way to failing. Somebody checked the margin instead of the pass/fail bit.

The cause was Aristotle’s quality rule: at least one premise must be affirmative, which constrains the premise-quantifier multiset in a way that correlates with the conclusion. The fix was per-conclusion-quantifier premise signatures. The worst feature now measures 0.004 bits.

Had nobody looked past the checkmark, that leak would have shipped, and the audit would have reported success every build thereafter.

Two rules that came out of it

Check the margin, not the pass bit. A feature sitting at 97% of threshold after a fix means the fix moved the leak, not closed it. The number that matters is how much headroom you have, not whether you are technically inside the line. This generalises well beyond puzzle generation.

When surface is semantics, the leak is intrinsic and the task needs redesign, not camouflage. Sometimes a feature correlates with the answer because of what the task fundamentally is — the wording cannot be changed without changing the problem. At that point, disguising the cue is dishonest engineering; the right move is to conclude the task family is not trainable in that form.

That criterion is now what decides whether a proposed new exercise gets built at all.

Why this is on a public page

Because five leaks is not a small number, and a page listing only the guarantees would be an advertisement.

Every one of these was in the software, and some of them shipped. The 0.93-bit one was handing away answers on a task built specifically to stop that.

What the audit gives us isn’t the absence of mistakes. It’s a way of finding them that doesn’t depend on anyone being clever on the day, plus a record of what it found. The threshold is a named constant, so it can be tightened rather than argued about.