AGANOMALY GRAPHLINK ANALYSIS← All dossiers sourced using AI
SOURCED USING AIpsi_consciousness

The Reproducibility Project: Psychology

Large-scale replication project · 2015 · Multiple laboratories worldwide · International

Also known as: Open Science Collaboration 2015, OSC reproducibility project, Estimating the reproducibility of psychological science, Many psychology replications

WHAT THIS LABEL MEANS

This dossier is a research synthesis sourced using AI, not documentary evidence. Use the reference leads to check important claims.

The Reproducibility Project: Psychology was a coordinated, international effort led under the name Open Science Collaboration to conduct direct or close replications of 100 experimental and correlational findings drawn from psychology papers published in 2008. Its principal report appeared in Science in 2015 as “Estimating the reproducibility of psychological science.” The project did not test telepathy, precognition, psychokinesis, mediumship, or any other psi claim as a dedicated research program. It is nevertheless highly relevant to cross-domain evaluation of psi and consciousness claims because many such claims depend on small statistical effects, flexible analytic choices, selective publication, and the distinction between an original result and an independent attempt to reproduce it. The project is best understood as evidence about the reliability of a sampled body of conventional psychology literature and about the practical value of transparent, multi-laboratory replication methods, rather than as evidence that all psychological findings, or all controversial findings, are false. Its central comparison lesson is methodological: apparent evidential strength changes when protocols, outcomes, samples, and analyses are independently revisited under more open conditions.

Words
1,911
Observations
10
Reference leads
5
Validation score
100/100

Chronology and project formation

In 2011, the Open Science Collaboration began organizing a distributed replication initiative in the context of widening concern about publication bias, underpowered studies, undisclosed analytic flexibility, and the difficulty of checking influential findings in psychology. The collaboration used online coordination and invited numerous researchers to contribute replication work, making it an early conspicuous example of networked, crowd-sourced meta-research.

The final sample consisted of 100 studies selected from papers published in 2008 in Psychological Science, Journal of Personality and Social Psychology, and Journal of Experimental Psychology: Learning, Memory, and Cognition. The sampled articles covered several areas of psychology rather than a single unified theory or method, so the eventual aggregate findings cannot automatically be applied to every subfield or every type of claim.

Replication teams developed study materials and procedures, generally sought input from original authors where feasible, and arranged review of proposed protocols before collecting data. The report was published in 2015 and compared original and replication results through several metrics, including statistical significance, effect-size estimates, subjective assessments of replication success, and meta-analytic combinations of original and replication evidence.

People, organizations, and research setting

The named collective author was the Open Science Collaboration, an organized network rather than a conventional single-laboratory team. Brian A. Nosek is widely associated with the project and with the Center for Open Science, an organization that promoted infrastructure and norms for preregistration, data sharing, registered reports, and reproducible workflows.

The work occurred across multiple laboratories and countries, with protocols implemented by many independent teams. This distributed setting matters because it reduced reliance on a single laboratory’s personnel and local conditions, while also introducing variation in experimenter skill, participant pools, languages, recruitment systems, equipment, and exact procedural interpretation.

The original-study authors were an important stakeholder group. Their involvement in answering questions or commenting on methods was intended to improve fidelity, but subsequent debate shows that consultation does not eliminate disagreements over what counts as the essential features of an experimental manipulation or a sufficiently direct replication.

Reported results and comparison variables

The headline result was that the replication studies yielded less favorable evidence for the original findings by several predeclared or reported indicators. A commonly repeated summary is that 97 of the original studies had statistically significant results whereas roughly 36 of the replications did, using the conventional threshold applied in the project report. This contrast is a descriptive result of the selected sample and specified analyses, not a universal probability that any given psychology result will fail.

Replication effect sizes were reported as substantially smaller on average than the original effect sizes, with a frequently cited aggregate comparison placing the replication estimates at approximately half the magnitude of the originals. In many individual cases, replication confidence intervals and effect estimates left room for partial consistency, sampling variation, or context dependence, even where a conventional significance test did not reproduce the original result.

The project also examined whether experts’ expectations tracked success. Researchers’ prior assessments were associated with replication outcomes to a degree, but prediction was not perfect. This is useful for controversial consciousness research because reputation, theoretical plausibility, and intuitive confidence may inform triage decisions without replacing transparent outcome data and independent testing.

Methods and evidential value

The project’s investigation model combined protocol review, larger or newly calculated samples, independent data collection, and public-facing research artifacts to a greater extent than was customary in much earlier psychology. Its practical contribution was to make replication a planned research product rather than an informal, often unpublished, response to an unexpected claim.

For comparison with alleged anomalous cognition, the most transferable features are advance specification of hypotheses and primary outcomes, separation of exploratory from confirmatory analysis, sufficiently powered designs, retention and disclosure of exclusions, standardized materials, and independent replication by groups with differing incentives. A psi experiment may require additional protections, such as automated randomization, secure time-stamped data capture, blinding, fraud-resistant procedures, and strong controls against sensory leakage, but those measures complement rather than replace open-science practices.

The project did not establish a simple binary rule for scientific truth. A replication can be affected by differences in population, language, cultural setting, task implementation, experimenter behavior, incentives, timing, measurement quality, or the original result’s sampling error. Conversely, an original positive result can be amplified by publication selection and researcher degrees of freedom. The appropriate unit of assessment is therefore the full evidential record for a particular claim and protocol family.

Disagreements, qualifications, and alternative accounts

Critics argued that some replication attempts were not sufficiently faithful to the original studies, that supposedly direct replications may have altered meaningful contextual features, and that judging success chiefly through a significance threshold can be misleading. Some original authors disputed particular replication outcomes or maintained that their effects were theory-dependent and not expected under the replication conditions.

The Open Science Collaboration and allied commentators responded that no single metric had been treated as decisive and that converging indicators generally pointed in the same direction: replication estimates were smaller and positive evidence was less frequent than in the original literature. Debate also addressed whether lower replication success was better explained by questionable research practices and publication bias, by replication quality, by genuine contextual heterogeneity, or by combinations of these mechanisms.

For extraordinary-claim assessment, neither side licenses an easy conclusion. The project supports caution about isolated, dramatic, small-sample results, particularly when they are reported after multiple analytic choices or have not been independently repeated. It does not demonstrate that low prior-probability claims are impossible, nor does it prove that a null or attenuated replication necessarily identifies the mechanism behind an original positive report.

Transmission, public interpretation, and commercial context

The Science article became a major public symbol of psychology’s broader “replication crisis,” a label that is rhetorically powerful but can conceal meaningful variation among fields, methods, and claims. Press accounts often compressed a multi-metric, heterogeneous study into the simpler statement that most psychology findings were false or failed, which goes beyond what the project itself could strictly show.

Its methods and vocabulary spread through university training, funder discussions, journals, preprint culture, and open-science organizations. Terms such as preregistration, reproducibility, replication, direct replication, effect size, and publication bias consequently became common reference points for researchers studying both mainstream behavioral findings and fringe or anomalistic topics.

Commercial and career incentives are relevant background conditions. Novel positive findings can attract attention, publication, grant prospects, media coverage, book sales, and institutional prestige, while replications have historically been viewed as less original and harder to publish. The project partly countered that imbalance by making coordinated replication visible and citable, but it did not remove incentives that can shape research agendas and public messaging.

Cross-case motifs and comparison use

This case connects to any research controversy where an initially impressive claim rests on a modest sample, a marginal statistical threshold, multiple outcomes, or a laboratory-specific protocol. It supplies a framework for asking whether later teams used an adequately similar procedure, whether a null result reflects failure of a theory or variation in implementation, and whether effect sizes remain stable as samples grow.

It is particularly relevant to psi and consciousness dossiers as a methodological comparator rather than as direct evidence. Claims of precognition, presentiment, ganzfeld information transfer, meditation-related anomalous perception, or altered-state effects should be evaluated on their own data and controls, while also being mapped against the same recurrent risks: selective reporting, insufficient blinding, optional stopping, weak randomization, sensory cues, expectancy effects, and publication asymmetry.

Scope limits and research cautions

The 100-study corpus was a purposive sample from three high-profile journals and one publication year, not a random census of all psychological research. It does not directly estimate reproducibility in parapsychology, clinical medicine, neuroscience, qualitative consciousness studies, industrial research, or disciplines with materially different methods and publication cultures.

“Reproducibility” has several meanings that should not be conflated. Computational reproducibility concerns whether the same data and code yield the reported analysis; direct replication concerns whether a near-match study produces compatible evidence; conceptual replication tests related predictions with a different design; and generalizability asks whether an effect survives changes of setting or population. The project chiefly concerned empirical replication and compatibility of effects, not all of these questions at once.

Researchers using this project as a benchmark should inspect the individual study records, protocol correspondence, exclusions, sample sizes, outcomes, and later scholarly exchanges. A headline aggregate is valuable for orientation, but it cannot adjudicate a particular controversial experiment without case-specific evidence.

Chronology

2011

Collaboration is organized.

The Open Science Collaboration reportedly began coordinating distributed replications amid increasing debate over transparency and reliability in psychology.

reported
2008

Source articles are published.

The project sampled studies from articles published in three psychology journals during 2008.

documented
2011–2014

Protocols and replications are developed.

Multiple teams reportedly prepared, reviewed, and implemented replication studies, often with attempted communication with original authors.

approximate
2015

Main report appears in Science.

The collaboration published “Estimating the reproducibility of psychological science,” reporting results from 100 replications.

documented
2015 onward

Methods and interpretation are debated.

Commentary challenged aspects of fidelity and interpretation, while proponents emphasized convergent evidence from multiple outcome metrics.

documented

People and roles

Open Science Collaboration

Coordinating authorship collective.

A distributed network credited as the author of the principal 2015 report.

Brian A. Nosek

Open-science researcher and prominent project-associated figure.

Commonly associated with the collaboration and the Center for Open Science; exact responsibilities should be checked in the article and project materials.

Center for Open Science

Open-science infrastructure and advocacy organization.

Relevant organizational context for the project’s transparency and coordination practices.

Original-study authors

Authors of the 100 selected findings.

Some provided clarification or feedback, and some later contested the adequacy or interpretation of specific replications.

Replication teams

Independent or semi-independent data-collection groups.

Their multi-site participation was a principal design feature, though teams necessarily varied in local context and implementation.

Connections to explore

Independent multi-laboratory replication.

A claim can be tested by teams outside the originating laboratory to reduce dependence on one investigator, participant pool, or local procedural culture.

Suggested search: multi-laboratory replication preregistration psychology anomalous cognition.

Effect-size shrinkage.

Early reports may overestimate an effect because of sampling variation, publication selection, flexible analysis, or contextual conditions, making later estimates smaller without proving literal error or fraud.

Suggested search: effect size shrinkage replication publication bias psychology.

Protocol fidelity versus contextual moderation.

A disputed replication may reflect a flawed reproduction, a real but context-sensitive effect, or an unstable original finding; case comparison should document the exact differences rather than assume one explanation.

Suggested search: direct replication fidelity contextual sensitivity original authors criticism.

Preregistration and outcome transparency.

Advance specification of primary analyses reduces ambiguity created by optional stopping, multiple outcomes, and post hoc hypotheses, all of which are important in small-effect research.

Suggested search: preregistration optional stopping small effects psi research.

Headline simplification of nuanced aggregate evidence.

Public narratives may turn heterogeneous replication results into categorical claims that an entire field is true, false, broken, or vindicated.

Suggested search: replication crisis media framing Open Science Collaboration 2015.

Unretrieved reference leads

LEADS, NOT CITATIONS These suggestions have not been retrieved or verified. They are starting points for source checking.
  1. Estimating the reproducibility of psychological science.

    Open Science Collaboration. · Science research article.

    This is the principal 2015 project report and should be checked for the sample, protocols, metrics, data-access statements, and wording of conclusions.

    Suggested search: Open Science Collaboration Estimating the reproducibility of psychological science Science 2015.
  2. Center for Open Science project materials for the Reproducibility Project: Psychology.

    Center for Open Science and Open Science Collaboration. · Project collection or institutional web materials.

    Potential source for study-level protocols, data, implementation details, and project-history context.

    Suggested search: Center for Open Science Reproducibility Project Psychology materials protocols data.
  3. Comment on “Estimating the reproducibility of psychological science”.

    Daniel T. Gilbert and collaborators. · Scholarly comment or correspondence.

    A prominent critical response concerning replication quality and interpretation that can anchor the dispute record.

    Suggested search: Gilbert comment Estimating the reproducibility of psychological science replication quality.
  4. Response to Comment on “Estimating the reproducibility of psychological science”.

    Open Science Collaboration. · Scholarly response or correspondence.

    Relevant for the collaboration’s reply to methodological criticisms and its defense of multi-metric interpretation.

    Suggested search: Open Science Collaboration response comment Estimating reproducibility psychological science.
  5. Making sense of replications.

    Alexander Etz and Joachim Vandekerckhove. · Methodological commentary.

    A useful lead for debate over inferential frameworks and how replication outcomes should be interpreted beyond a binary significance criterion.

    Suggested search: Etz Vandekerckhove Making sense of replications reproducibility project psychology.