Here is an uncomfortable fact that ought to be taught on the first day of every research methods course and is instead usually taught the hard way, by a retraction. If you take two groups that are genuinely identical — same treatment, same population, no real difference whatsoever — and you test enough things about them, you will eventually find a difference that is “statistically significant” at p<0.05. Not because anything is there, but because that is precisely what p<0.05 was built to allow. The threshold does not promise truth. It promises a one-in-twenty false alarm rate, and one in twenty is not rare if you keep rolling the dice.
It helps to say plainly what a p-value actually is, because almost everyone — including, on a tired day, the people computing them — quietly believes it means something grander. A p-value is the probability of seeing a result at least this striking if there were truly no effect at all. That is it. It is not the probability that your hypothesis is true. It is not the probability that you are wrong. It is a measure of how surprised pure chance would be to bump into your data, and chance, it turns out, is not easily surprised if you give it twenty tries.
The button below runs an honest little experiment in dishonesty. There is genuinely no effect in the simulated data — every group is drawn from the same distribution as every other. Each press does what a hopeful researcher does at 11 p.m. before a deadline: it tests one more variable, one more subgroup, one more way of slicing the same nothing. Watch the running tally, and watch the smallest p-value found so far creep downward until, with grim inevitability, one of them sneaks under the magic line and lights up green.
Notice the cruel arithmetic. Run the analysis once and you have roughly a one-in-twenty chance of a false positive — uncomfortable but survivable. Run it twenty times on the same dead data and your chance of at least one spurious “finding” climbs to about sixty-four per cent. The dots are not lying to you; each individual test is behaving exactly as designed. The lie is assembled later, when only the green dot gets written up and the grey ones are quietly forgotten — a practice with the deceptively cheerful name p-hacking, and its equally guilty twin, reporting the question you went looking for as though it were the question you set out to ask.
If you torture the data long enough, it will confess to anything. — attributed to Ronald Coase, and proven nightly in spreadsheets everywhere
How not to fool yourself
The defences are unglamorous and they work. Decide your primary outcome before you see the data, write it down somewhere you cannot later edit, and treat everything else as a hypothesis to be tested afresh rather than a result to be banked. Pre-register the study so the grey dots are on the record alongside the green one. When you do run many comparisons, correct for it — a Bonferroni adjustment is blunt but honest, and being made to share the 0.05 out among twenty tests is exactly the medicine the simulator above is warning you about. And treat a startling subgroup finding as a question for the next study, never an answer from this one.
None of this means p-values are worthless or that significance testing is a con. It means a p-value is a smoke alarm, not a verdict: useful, worth heeding, and prone to going off when you wave enough toast at it. The number tells you something genuinely worth knowing about how easily chance could have produced your data. It simply cannot tell you the one thing everyone desperately wants it to — whether you are right — and the surest way to be wrong is to keep pressing the button until it finally tells you what you came to hear.