SIGNIFICANCE SYNDROME: A MIXED METHODS EXPLORATION OF P VALUE OVERRELIANCE IN MEDICAL SCIENCE

Main Article Content

Dr. Ravi Shastri
Dr. Neeraj Sidana
Dr. Rama Shankar
Dr. Saurabh Singh
Dr. B.N Singh

Keywords

p‑Value; Mixed‑Methods; Bibliometric Analysis; Effect Size; Confidence Interval; Bayesian Inference; Reproducibility.

Abstract

Introduction: Over the past three decades, the p‑value has become the dominant metric for judging statistical significance in biomedical research, yet its binary interpretation (significant vs. non‑significant) has contributed to publication bias, irreproducible findings, and an undue focus on arbitrary thresholds.


Aim and Objectives: This mixed‑methods study was designed to examine longitudinal trends in p‑value reporting and to explore researchers’ attitudes toward p‑value usage and alternatives. Specifically, we aimed to (1) quantify the proportion of articles reporting the p‑value alone versus those also including effect sizes, confidence intervals, or alternative inferential metrics at four timepoints (1990, 2000, 2010, and 2020–2024); (2) identify study and journal characteristics associated with richer statistical reporting; and (3) assess researchers’ perceived pressure to achieve p<0.05 and their familiarity with non‑p‑value approaches.


Methods: We conducted a bibliometric analysis of 1,042 original research articles sampled from five high‑impact journals (JAMA, The Lancet, NEJM, BMJ, and PLOS Medicine) at the four specified timepoints. Each article was classified by reporting style—p‑value only; p‑value + effect size; p‑value + confidence interval; or alternative metrics—and data were extracted on study design, journal impact factor, funding source, author count, and geographic region. Concurrently, we administered a cross‑sectional online survey to 217 active biomedical researchers, capturing perceived publication pressure, statistical training background, and familiarity with effect sizes, confidence intervals, and Bayesian methods. Trends were analyzed using Cochran–Armitage tests, and multivariable logistic regression identified independent predictors of enriched reporting.


Results: Exclusive p‑value reporting declined from 76.3% in 1990 to 37.9% in 2020–2024. Articles in higher‑impact journals and randomized controlled trials were significantly more likely to include effect sizes and confidence intervals. Survey respondents reported high familiarity with effect sizes (92%) and confidence intervals (95%) but lower familiarity with Bayesian methods (42%), and 78% felt pressured to meet p<0.05 thresholds, particularly early‑career investigators without advanced statistical training.


Conclusion: Although editorial guidelines have encouraged more informative reporting, entrenched incentives and gaps in statistical literacy continue to drive “significance syndrome.” We recommend mandating effect‑size and interval reporting, integrating advanced inferential training into research curricula, and promoting preregistration and open data sharing to enhance transparency and reproducibility.

Abstract 0 | Pdf Downloads 0

References

1. Wasserstein, R. L., & Lazar, N. A. (2016). The ASA Statement on p-Values: Context, Process, and Purpose. The American Statistician, 70(2), 129–133.https://doi.org/10.1080/00031305.2016.1154108
2. Ioannidis JPA (2005) Why Most Published Research Findings Are False. PLoS Med 2(8): e124. https://doi.org/10.1371/journal.pmed.0020124
3. Goodman SN. Toward evidence-based medical statistics. 1: The P value fallacy. Ann Intern Med. 1999 Jun 15;130(12):995-1004. doi: 10.7326/0003-4819-130-12-199906150-00008. PMID: 10383371.
4. Benjamin DJ, Berger JO, Johannesson M, Nosek BA, Wagenmakers EJ, Berk R, et al. Redefine statistical significance. Nat Hum Behav. 2018 Jan;2(1):6–10. doi:10.1038/s41562-017-0189-z. PubMed PMID: 30980045.
5. Amrhein V, Greenland S, McShane B. Scientists rise up against statistical significance. Nature. 2019 Mar;567(7748):305-307. doi: 10.1038/d41586-019-00857-9. PMID: 30894741.
6. Halsey LG, Curran-Everett D, Vowler SL, Drummond GB. The fickle P value generates irreproducible results. Nat Methods. 2015 Mar;12(3):179-85. doi: 10.1038/nmeth.3288. PMID: 25719825.
7. Sterne JA, Davey Smith G. Sifting the evidence-what's wrong with significance tests? BMJ. 2001 Jan 27;322(7280):226-31. doi: 10.1136/bmj.322.7280.226. PMID: 11159626; PMCID: PMC1119478.
8. Wasserstein RL, Schirm AL, Lazar NA. Moving to a world beyond “p < 0.05”. Am Stat. 2019;73(Suppl 1):1–19. doi:10.1080/00031305.2019.1583913.
9. Trafimow, D. and Marks, M. (2015) Editorial. Basic and Applied Social Psychology, 37, 1-2.
https://doi.org/10.1080/01973533.2015.1012991
10. Kass RE, Raftery AE. Bayes factors. J Am Stat Assoc. 1995;90(430):773–795. doi:10.1080/01621459.1995.10476572.
11. Gelman, Andrew & Stern, Hal. (2006). The Difference Between “Significant” and “Not Significant” Is Not Itself Statistically Significant. The American Statistician. 60. 328-331. 10.1198/000313006X152649.
12. Cumming G. The new statistics: why and how. Psychol Sci. 2014 Jan;25(1):7-29. doi: 10.1177/0956797613504966. Epub 2013 Nov 12. PMID: 24220629.
13. Greenland S, Senn SJ, Rothman KJ, Carlin JB, Poole C, Goodman SN, Altman DG. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. Eur J Epidemiol. 2016 Apr;31(4):337-50. doi: 10.1007/s10654-016-0149-3. Epub 2016 May 21. PMID: 27209009; PMCID: PMC4877414.
14. Gelman A, Carlin J. Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors. Perspect Psychol Sci. 2014 Nov;9(6):641-51. doi: 10.1177/1745691614551642. PMID: 26186114.
15. Smaldino PE, McElreath R. The natural selection of bad science. R Soc Open Sci. 2016 Sep 21;3(9):160384. doi: 10.1098/rsos.160384. Erratum in: R Soc Open Sci. 2023 Sep 6;10(9):rsos231026. doi: 10.1098/rsos.231026. PMID: 27703703; PMCID: PMC5043322.
16. Munafò, M., Nosek, B., Bishop, D. et al. A manifesto for reproducible science. Nat Hum Behav 1, 0021 (2017). https://doi.org/10.1038/s41562-016-0021
17. Begley CG, Ioannidis JP. Reproducibility in science: improving the standard for basic and preclinical research. Circ Res. 2015 Jan 2;116(1):116-26. doi: 10.1161/CIRCRESAHA.114.303819. PMID: 25552691.
18. Collins, F., Tabak, L. Policy: NIH plans to enhance reproducibility. Nature 505, 612–613 (2014). https://doi.org/10.1038/505612a
19. Ioannidis JPA (2018) Meta-research: Why research on research matters. PLoS Biol 16(3): e2005468. https://doi.org/10.1371/journal.pbio.2005468
20. Cumming G, Finch S. Inference by eye: confidence intervals and how to read pictures of data. Am Psychol. 2005 Feb-Mar;60(2):170-80. doi: 10.1037/0003-066X.60.2.170. PMID: 15740449.

Most read articles by the same author(s)