We find evidence for what we already believe because we go looking for confirmation rather than contradiction, and confirmation is almost always available. This is not primarily a failure of honesty or intelligence. It is a failure of test design, running quietly, in a mind that has no particular reason to notice it is happening.
Peter Wason demonstrated the structure of the problem in 1960 with a task simple enough to explain in a sentence, and the demonstration has held up for more than sixty years.
The 2-4-6 task
Wason's paper, "On the failure to eliminate hypotheses in a conceptual task," appeared in the Quarterly Journal of Experimental Psychology and became one of the most cited articles that journal has published.
The setup: participants are shown the number triple 2, 4, 6, and told it conforms to a simple rule the experimenter has in mind. Their job is to work out the rule. They can propose as many triples of their own as they like, and after each one they are told only whether it fits the rule or not. They announce the rule when they are confident.
Almost everyone forms an immediate hypothesis. Numbers increasing by two. So they test 8, 10, 12. Fits. They test 20, 22, 24. Fits. They test 1, 3, 5. Fits. Three confirmations, high confidence, announce the rule.
The actual rule was any three numbers in ascending order. Every triple they generated fit it, because every triple they generated was designed to fit their own narrower hypothesis, which sat entirely inside the real one. The test that would have revealed the truth was a triple they expected to fail, something like 1, 2, 3 or 5, 19, 400. Being told that fit too would have exploded the hypothesis instantly. Nobody wanted to run that test.
Wason's numbers: of twenty-nine participants, six reached the correct rule without a previous incorrect announcement. Thirteen announced one incorrect rule first. Nine announced two or more. One never reached a conclusion. The people who did worst were not the ones with the wrong hypothesis. Everyone starts with the wrong hypothesis. They were the ones unable or unwilling to design a test that could have killed it.
Why confirming tests cannot work
This is the structural point, and it is worth stating in general terms because it applies far outside a number puzzle.
If your hypothesis is narrower than the truth, every confirming test succeeds and tells you nothing. A positive result is compatible with your hypothesis being right and equally compatible with a broader rule that happens to contain it. Only a test you predicted would fail can separate those two possibilities. Confirming evidence accumulates confidence without accumulating information, and confidence is the thing that feels like knowing.
What Nickerson's review established
Raymond Nickerson's 1998 article in Review of General Psychology remains the canonical survey of the field. He defined confirmation bias as the seeking or interpreting of evidence in ways that are partial to existing beliefs, expectations, or a hypothesis in hand, and traced it through medicine, law, science, politics, and everyday judgment. The full review is openly available and is worth reading in whole.
His most important claim is easy to miss. The bias is largely unintentional. It is not, in most cases, people knowingly ignoring inconvenient facts. It is a set of default habits that feel like ordinary careful thinking from the inside, which is precisely why it is so difficult to catch in yourself.
The four forms it takes
- Biased search. You generate the questions, queries, and tests most likely to return support. This is the 2-4-6 failure, and it survives contact with search engines intact.
- Biased interpretation. Ambiguous evidence gets read in the favorable direction. The same mixed result reads as encouraging to someone expecting success and inconclusive to someone expecting failure.
- Biased memory. Confirming instances are recalled more readily than disconfirming ones, so the retrospective tally is skewed before you even start counting.
- Biased stopping. The most quietly damaging one. You stop investigating at the moment you find support, so the search terminates precisely when it starts agreeing with you.
An honest complication
The clean story about Wason's task is that people are irrationally drawn to confirmation, but that reading has been contested for good reasons. Klayman and Ha argued in Psychological Review in 1987 that testing positive cases is often a reasonable default strategy under uncertainty. In many real environments, where the thing you are hypothesizing about is rare and the space of possibilities is enormous, checking cases you expect to fit is an efficient way to gather information. What makes it fail in the 2-4-6 task is a specific and somewhat unusual structure: a hypothesis nested entirely inside the true rule.
The practical upshot is not that positive testing is always wrong. It is that positive testing is unreliable exactly when your belief might be a special case of something broader, which describes most beliefs people hold about their own habits.
Why this matters if you are tracking a routine
Anyone evaluating whether a practice is working is running an experiment where they are the researcher, the subject, and the person who would prefer a particular answer. That is three roles that are supposed to be held by different people for a reason.
Start an affirmation practice and you will begin noticing moments that fit it. That noticing is real, and it is not a result. It is attention reallocating, which produces the frequency illusion, and then confirmation bias files each instance as evidence while never opening a file for the instances that did not occur. This combination is what makes almost any new routine feel effective in week one, including routines with no plausible mechanism at all. It is also the engine behind pseudo-explanations that feel self-verifying from the inside, like the claim that the reticular activating system filters reality to show you your goals.
Five ways to test your own routine more honestly
- Write down the outcome before you start. Name the specific thing you expect to change. A vague expectation cannot be disconfirmed, which is what makes it so comfortable.
- Name a falsifier. Decide in advance what you would have to observe to conclude it is not working. If you cannot name one, you are not running a test.
- Record before you interpret. Log the raw observation the same day, in the same format, whether or not it fits. Interpretation contaminates memory quickly.
- Count the misses. Keep a tally of the days nothing happened. The missing denominator is the single largest source of false confidence in self-tracking.
- Give the null result somewhere to land. Decide ahead of time what you will do if the answer is no. Without that, "no" has no path to being reached, and every observation becomes support.
This is roughly the standard applied in the research on which affirmations produce measurable effects and which produce nothing, and it is why that literature can say something useful. Controlled comparison, defined outcomes, and a real possibility of a null result are what separate a finding from a feeling.
Your mirror already has a time slot. Give it a better script.
Toothily laser-engraves a surprise affirmation into a Moso bamboo handle with soft bristles. Twice a day, it is in your hand before your phone is.
One-time or subscription. Keep the Brush Guarantee: not happy? Email within 30 days for a full refund, and keep the brushes.