A Bayesian Analysis of the Bem Precognition Experiments
Also known as: Wagenmakers et al. 2011 Bem critique, Bem Bayesian critique, Why psychologists must change the way they analyze their data: The case of psi
This dossier is a research synthesis sourced using AI, not documentary evidence. Use the reference leads to check important claims.
This dossier concerns a 2011 methodological intervention in the controversy surrounding Daryl J. Bem’s reported experimental evidence for precognition. Bem’s study program used ordinary laboratory tasks and statistical comparisons to argue that participants sometimes appeared to respond to future events, including future exposure to erotic stimuli or later practice opportunities. Wagenmakers and colleagues did not present a competing paranormal experiment. Instead, they used Bayesian reasoning to ask how much Bem’s reported data should change belief in a highly implausible claim. Their central remembered position was that conventional statistically significant results, particularly when treated without the prior improbability of precognition and without full regard to analytic flexibility, did not amount to persuasive evidence for psi under plausible prior assumptions. The dispute therefore belongs as much to the reform of psychological inference as to parapsychology. It became a durable example in arguments about p-values, Bayes factors, prior distributions, optional stopping, replication, publication bias, and the meaning of extraordinary evidence. The case should not be read as proving either that precognition exists or that it is impossible. Its importance lies in displaying how different inferential frameworks can yield sharply different assessments of the same reported results, and how later reproducibility concerns changed the perceived evidential context. The cited works below are leads for later verification, not evidence consulted for this synthesis.
- Words
- 1,996
- Observations
- 10
- Reference leads
- 4
- Validation score
- 100/100
Chronology and documentary frame
The relevant sequence begins with Bem’s experimental reports, which treated several laboratory paradigms as tests of anomalous retroactive influence or precognition. The best-known public controversy intensified in 2011 after publication and wide discussion of the reported positive findings. Wagenmakers, Wetzels, Borsboom, and van der Maas then offered a Bayesian critique focused on the inferential implications of the reported evidence rather than on a new direct observation of a paranormal event.
The critique entered an already developing methodological debate in psychology about whether null-hypothesis significance testing could overstate evidential strength. Subsequent replication attempts, registered-report practices, and discussions of researcher degrees of freedom broadened the case from a dispute about a single psi claim into a teaching example about reproducibility. Exact publication order, response dates, and the status of each later replication should be checked against the relevant papers and records.
People, organisations, and setting
Daryl J. Bem was the psychologist whose precognition experiments supplied the data under debate. Eric-Jan Wagenmakers, Ruud Wetzels, Denny Borsboom, and Han L. J. van der Maas are associated with the Bayesian critique. Their work connected Dutch methodological and psychometric scholarship with a controversy centered on experimental psychology in the United States, making the setting international rather than a single haunted site, laboratory, or community tradition.
The operative setting was a research and publication environment: university psychology laboratories, journals, peer review, conference and online discussion, and later replication laboratories. The physical tasks reportedly involved commonplace computer-mediated or behavioral experiments, not spectacular sensory apparitions. That mundane setting mattered rhetorically because apparently ordinary procedures could generate results framed as extraordinary, while also leaving room for conventional concerns about design, analysis, selective reporting, and chance.
Reported phenomena and experimental character
The phenomenon at issue was not a witness report of a ghost, vision, or physical manifestation. It was a pattern of group-level statistical deviations that Bem interpreted as compatible with information or influence traveling from future events to present responses. Reported task experiences were generally ordinary and sensory: participants viewed screens, made choices, recalled or identified stimuli, completed computer tasks, and in some paradigms encountered emotionally or erotically salient images. The claimed anomaly resided in aggregate scoring and timing, not in an individually documented moment of unmistakable foresight.
Behaviourally, the studies reportedly assessed response accuracy, preference, priming-like effects, or performance changes relative to chance or a control condition. Participants could experience anticipation, uncertainty, attention to visual material, reward expectations, embarrassment around sexual imagery, fatigue, guessing, or ordinary learning effects. None of those subjective reactions independently establishes precognition. The Bayesian critique addressed the evidential interpretation of the final numerical pattern, and it did not itself validate a particular sensory experience as anomalous.
For cross-case purposes, the important motif is an extraordinary causal interpretation attached to small effects in familiar laboratory behavior. Such cases require separation of the raw response record, the pre-specified hypothesis, the analysis choices, and the narrative explanation offered after statistical significance is obtained.
Investigation history and Bayesian approach
The recalled critique reanalyzed or evaluated Bem’s reported results through Bayesian methods, especially the question of how observed data alter prior odds between an anomalous hypothesis and more conventional accounts. In this framing, an evidential measure such as a Bayes factor is not identical to a p-value. A small p-value measures incompatibility with a specified null model under particular assumptions, whereas Bayesian assessment compares predictive performance across models and incorporates prior assumptions. The distinction was central to the controversy.
The authors’ remembered conclusion was that the data did not justify strong confidence in precognition when plausible skeptical priors and reasonable alternative models were considered. This conclusion was conditional rather than a mathematical proof that all psi hypotheses are false. It depended on model specification, priors, and treatment of the reported experiments as a body of evidence. Critics of the critique could reasonably focus on those choices, while proponents of stronger claims would need to explain why more favorable priors or models are warranted.
The later investigation history includes attempts to repeat relevant paradigms and broad reforms aimed at making psychological evidence more diagnostic. Preregistration, transparent reporting of exclusions and stopping rules, sharing materials and data where possible, adequately powered independent replication, and prospective meta-analysis are all more informative than retrospective argument alone. The available context does not establish the complete outcome of every subsequent replication, so a researcher should verify each one separately.
Disputes, inferential disagreements, and ordinary alternatives
The principal disagreement concerned what counts as adequate evidence for a proposition that conflicts with established expectations about causation and time. Wagenmakers and colleagues emphasized that an extraordinary prior claim needs correspondingly strong diagnostic evidence. Bem and sympathetic readers could contend that data should be judged by experimental results rather than dismissed in advance, or that priors should not be assigned so skeptically. The disagreement is therefore partly philosophical and methodological, not merely a disagreement about arithmetic.
Mundane explanations include sampling variation, multiple testing, flexible analysis decisions, undisclosed exclusions, optional stopping, selective publication, imperfect randomization, experimenter effects, expectation effects, data-processing errors, and unrecognized correlations among studies. Some possibilities may be constrained by the original procedures, but their relative importance cannot be resolved from recalled summaries alone. A Bayesian analysis may represent some alternatives more explicitly than a conventional significance test, yet it also introduces debate about the appropriate alternative model and prior distribution.
It would be inaccurate to portray the critique as a direct demonstration of fraud, participant deception, or a laboratory malfunction. It would likewise be inaccurate to treat statistically significant outcomes as verified evidence that future events influenced present behavior. The most defensible characterization is a contested evidential case whose conclusions shift with assumptions and with the credibility assigned to the research process.
Transmission, retelling, and commercial influences
The case circulated through a scientific paper, methodological commentary, teaching materials, blogs, social-media discussion, journalism, and later reproducibility debates. Its compact storyline made it especially transmissible: a respected social psychologist reported effects suggestive of precognition, and methodologists argued that standard significance testing had made the evidence look stronger than it was. Retellings often simplify this into either a proof that psychology endorsed ESP or a proof that statistics debunked it, neither of which captures the conditional and technical nature of the dispute.
Commercial and reputational incentives may have shaped the wider transmission without determining the truth of the findings. Psi claims attract attention because they are surprising, while controversy over them can increase citations, media interest, speaking opportunities, classroom use, and public visibility for authors, journals, and commentators. Conversely, skepticism and replication reform also have professional incentives and institutional stakes. These influences are context for interpretation, not evidence of misconduct by any named individual.
Later accounts should distinguish the original reported experiments, the Bayesian critique, responses by participants in the debate, later replication studies, and popular summaries. Collapsing these layers can create false certainty about what was directly observed, what was statistically inferred, and what was later claimed about the episode.
Cross-case connections and comparison motifs
The strongest comparison motifs are extraordinary claims requiring extraordinary evidence, anomalous effects inferred from aggregate statistics, ordinary sensory tasks given nonordinary causal interpretations, and disagreement between frequentist and Bayesian standards. These motifs connect this case to replication controversies and disputed anomalous-cognition research, but thematic resemblance alone does not make another case a duplicate.
A second motif is the evidential role of research transparency. In cases where an effect is small and the theory is contentious, preregistered prediction, independent replication, full disclosure of analysis paths, and clear stopping rules can matter as much as a single nominally significant result. A third motif is transmission distortion: public summaries tend to turn a conditional methodological argument into a binary verdict about the paranormal.
Comparison should remain careful. A Bayesian critique of Bem’s experiments is not interchangeable with reports of séance phenomena, spontaneous haunt experiences, or military remote-viewing programs. Each has different data types, opportunities for sensory leakage, cultural genres, and relevant base-rate assumptions.
Limits of this dossier
This is an unverified recalled synthesis prepared from bounded context and general model knowledge, not a source-checked literature review. It does not establish exact numerical results, effect sizes, Bayes factors, sample sizes, model priors, journal dates, author responses, or the outcomes of all replication attempts. Those details should be verified from the original publications and subsequent methodological literature before use in research, teaching, or argument.
Bayesian conclusions are sensitive to choices about prior odds, parameter priors, and the models compared. Frequentist conclusions are likewise sensitive to design, stopping rules, multiplicity, and the definition of the tested null. Neither framework automatically answers every scientific question, and neither removes the need for transparent data, robust design, and independent replication.
No paranormal claim is treated here as verified. The dossier records a controversy about the evidential status of reported experimental results and the standards used to assess them. It should be used as a map for verification and cross-case indexing, not as a substitute for primary documents or a final adjudication of psi.
Chronology
Bem develops and reports experimental precognition paradigms.
Bem’s research program used controlled psychology tasks to test whether future events appeared statistically related to present responses.
approximateBem’s precognition findings become a major public methodological controversy.
Reported positive findings drew attention because they were framed as challenging ordinary causal expectations.
documentedWagenmakers and colleagues publish or circulate a Bayesian critique.
The critique argued, under its stated Bayesian assumptions, that the reported evidence was not persuasive for a highly extraordinary hypothesis.
documentedResponses and methodological debate expand.
Participants in the controversy disputed appropriate priors, inferential standards, and the implications of the case for psychological science.
reportedReplication and reproducibility discussions reshape the case’s reception.
The episode became a recurrent example in debates about preregistration, analytic flexibility, publication bias, and independent replication.
approximatePeople and roles
Daryl J. Bem
Psychologist and author associated with the original precognition experiments.His reported experimental results provided the empirical material assessed in the Bayesian controversy.
Eric-Jan Wagenmakers
Methodologist and coauthor of the Bayesian critique.He is associated with arguments for Bayesian evidence assessment in this case.
Ruud Wetzels
Methodologist and coauthor of the Bayesian critique.He participated in the critique of the evidential strength of the reported results.
Denny Borsboom
Psychometrician and coauthor of the Bayesian critique.He was among the authors linking the case to broader methodological concerns.
Han L. J. van der Maas
Psychologist and coauthor of the Bayesian critique.He was among the authors of the remembered 2011 analysis.
Psychological science journals and replication laboratories
Organisational and institutional setting.They served as channels for publication, debate, replication, and methodological reform.
Connections to explore
Extraordinary claims and prior plausibility.
Compare cases in which the same numerical result is interpreted differently because investigators assign different prior credibility to the proposed mechanism.
Suggested search: Bayesian extraordinary claims prior odds anomalous cognition methodology.P-values versus Bayes factors.
Compare methodological controversies where null-hypothesis significance and Bayesian model comparison give different impressions of evidential strength.
Suggested search: psychology p values Bayes factors replication controversy.Researcher degrees of freedom.
Compare small-effect literatures vulnerable to optional stopping, multiple outcomes, exclusions, and selective reporting.
Suggested search: optional stopping analytic flexibility publication bias psychology.Ordinary laboratory behaviour with anomalous interpretation.
Compare psi experiments with routine cognitive tasks whose claimed anomaly appears only in aggregate data.
Suggested search: anomalous cognition laboratory task aggregate statistical effect.Replication and transmission distortion.
Compare cases where popular retellings obscure the difference between an original experiment, a critique, and later replication evidence.
Suggested search: Bem precognition replication public retelling reproducibility.Unretrieved reference leads
A Bayesian Analysis of the Bem Precognition Experiments.
Eric-Jan Wagenmakers, Ruud Wetzels, Denny Borsboom, and Han L. J. van der Maas. · Methodological critique.
This is the central suggested lead for verifying the Bayesian argument, its assumptions, and its exact conclusions.
Suggested search: Wagenmakers Wetzels Borsboom van der Maas 2011 Why psychologists must change the way they analyze their data Bem.Feeling the Future: Experimental Evidence for Anomalous Retroactive Influences on Cognition and Affect.
Daryl J. Bem. · Original experimental report.
This is the suggested primary lead for checking the reported paradigms, outcomes, and original interpretation.
Suggested search: Daryl Bem Feeling the Future anomalous retroactive influences cognition affect.Methodological discussions and responses concerning the Bem precognition debate.
Various authors. · Response and methodological literature.
This lead can help reconstruct disagreements about priors, Bayes factors, significance testing, and the scope of the critique.
Suggested search: Bem precognition Bayesian critique response priors replication methodology.Later independent replications and registered studies of Bem-style paradigms.
Various research groups. · Replication literature.
This lead is needed to assess how later prospective evidence affected reception of the original claims.
Suggested search: independent replication Bem precognition experiments preregistered.