What Brain Scans Can Tell Us, and What They Cannot

Updated

A functional brain scan can tell you where blood oxygenation changed while a person did one task rather than another. That is the measurement. It cannot tell you what the person was feeling, and the inference from "this region became more active" to "therefore they experienced X" is not valid reasoning, however often it appears under a picture of a glowing brain. This is the single most useful thing to understand about neuroimaging, because it is the step that converts a modest technical result into a marketing claim.

Start with what the machine actually records

Functional MRI does not measure neurons firing. It measures the blood oxygen level dependent signal, or BOLD, a chain of proxies: neural activity raises local metabolic demand, demand changes blood flow, the flow shift alters the ratio of oxygenated to deoxygenated hemoglobin, and that ratio has magnetic properties the scanner can detect. The response peaks several seconds after the neural event that caused it.

The spatial unit is a voxel, a few millimetres on a side, containing hundreds of thousands of neurons doing many different things at once. And the colored image is not a photograph. It is a statistical map of which voxels differed between two conditions by more than some threshold, painted onto an anatomical image for legibility. The color is a p-value, not a brightness.

None of that makes fMRI a bad tool. It makes it an indirect, slow, coarse, statistical tool, and everything that follows is a consequence of those four adjectives.

1. Reverse inference, the central error

Russell Poldrack named the problem in Trends in Cognitive Sciences in 2006, and it remains the most important concept in reading any neuroimaging claim.

Forward inference is what the experiment supports: assign a task, observe which regions become more active than in the control condition. Reverse inference runs the arrow backwards, observing activity in a region and concluding that the mental process associated with it was occurring. That move is not deductively valid, because a region being engaged by a process does not mean the process is the only thing engaging the region.

How much a reverse inference is worth depends on how selective the region is. If a structure activated for one process and nothing else, seeing it light up would be strong evidence. Real regions are nothing like that. The anterior cingulate, the insula and the amygdala each participate in a long and heterogeneous list of tasks, so their activation narrows the space of possible mental states far less than the confident caption suggests. Poldrack worked a quantitative example from a database of published activations and showed the evidential update is modest rather than decisive.

He proposed the constructive alternative in Neuron in 2011: rather than eyeballing a region and naming a feeling, build a decoder across large datasets and test whether it predicts the mental state in data it has never seen. That is a real path from activation to inference, and it demands the out-of-sample validation the informal version skips.

2. The dead salmon, and the arithmetic behind it

In 2009 Craig Bennett and colleagues placed a dead Atlantic salmon in a scanner and ran a standard emotion-recognition task on it. Analyzed with uncorrected statistics, the salmon showed a cluster of significant activation in its brain cavity. Corrected for multiple comparisons, nothing.

The joke carries real arithmetic. A whole-brain analysis tests tens of thousands of voxels at once. At an uncorrected threshold of p less than 0.001, roughly one in a thousand tests returns a false positive by chance alone, so a 40,000-voxel volume yields around 40 spurious hits before any biology is involved. Cluster them and some will look convincingly like a brain region. The salmon was an argument that correction is not statistical fussiness; it is the difference between a finding and a coincidence.

3. Voodoo correlations, or how to accidentally cheat

Vul, Harris, Winkielman and Pashler circulated the paper as "Voodoo Correlations in Social Neuroscience." It appeared in Perspectives on Psychological Science in 2009 under a milder title.

The error they identified is non-independence, sometimes called double dipping. A researcher searches the whole brain for voxels whose activity correlates with a personality or emotion measure, keeps the voxels that correlate most strongly, then reports the correlation within that selected set as the result. The selection used the same data that produced the number, so the correlation is inflated by construction. It is the statistical equivalent of drawing the target around the arrows.

The tell was that many published correlations exceeded what the reliability of the two measures could jointly support. A substantial share of the papers Vul and colleagues surveyed had used a non-independent analysis. The practice is far less common now, which is what a methodological critique working looks like.

4. Cluster failure, and a correction worth noticing

Eklund, Nichols and Knutsson ran what may be the most consequential audit of the field in PNAS in 2016. They analyzed resting-state scans, which contain no task effect, with the same software pipelines used for real experiments. Any "significant" result was by definition a false positive.

The common parametric methods for cluster-extent inference in the major packages produced familywise error rates as high as roughly 70 percent, against a nominal 5 percent, with the inflation worst at permissive cluster-forming thresholds. The team also found a longstanding bug in one package's simulation routine.

Then the part that belongs in this article on principle. The original paper stated that the results question the validity of some 40,000 fMRI studies, a line that traveled worldwide within days. The authors subsequently issued a formal correction revising it to question the validity of a number of studies and to note the largest impact on weakly significant results. The narrower claim is the accurate one. A cluster of articles about the gap between what a study found and what got said about it should apply that standard to the papers it likes too.

Underneath all of it sits statistical power. Button and colleagues estimated in Nature Reviews Neuroscience in 2013 that median power in neuroscience was around 21 percent, which lowers the chance a real effect is detected and raises the chance a detected one is exaggerated.

Six questions to ask when someone shows you a brain scan

  1. What was the comparison condition? Every activation is a difference between two states. A scan with no stated contrast is not reporting anything.
  2. Is the claim forward or reverse? "People doing X showed more activity in region Y" is supported. "Region Y was active, so they felt X" is not.
  3. Was the analysis corrected for multiple comparisons, and how? This is the salmon question, and it has a right answer.
  4. Were the voxels selected using the same data that produced the reported effect? If yes, the number is inflated by construction.
  5. How many participants? Neuroimaging samples have often been small, and small samples produce unstable, overestimated effects.
  6. Has it been replicated, ideally by someone else? A single striking scan is a hypothesis.

What scans are genuinely good for

Presurgical mapping. Testing mechanistic hypotheses where a theory predicts a particular contrast in advance. Measuring structural change in training studies, with the caveats that carries. Building and validating decoders on large, shared datasets. This is a productive tool, used carefully by a field that has spent two decades publicly auditing itself, which is more than most fields can say.

What it is not is a mind reader, and that is the capacity implied whenever a scan appears next to a product. The claim that a mental technique reshapes the brain almost always rests on an image and a caption performing a reverse inference, the same structure that props up folk neuroscience like the reticular activating system explanation. The discipline extends to research whose conclusions you like: the imaging work in the self-affirmation literature deserves to be taken seriously and is still fMRI, subject to every limit above. So is the neuroscience behind memory reconsolidation. Applying the standard only to claims you dislike is not skepticism.

Your mirror already has a time slot. Give it a better script.

Toothily laser-engraves a surprise affirmation into a Moso bamboo handle with soft bristles. Twice a day, it is in your hand before your phone is.

Shop Toothily

One-time or subscription. Keep the Brush Guarantee: not happy? Email within 30 days for a full refund, and keep the brushes.

Back to blog