THE EVIDENCE BASE FOR IMPORTANCE OF ARTIFICIAL INTELLIGENCE IN GENERAL SURGERY

Main Article Content

Dr A.Swathi Reddy
Dr Yanala Sheker
Dr Janipalli Mounica

Keywords

artificial intelligence, machine learning, general surgery, clinical trials as topic, translational research, research reporting

Abstract

Background:  Publication output on artificial intelligence (AI) in general surgery has grown steeply, and it is widely assumed that prospective clinical evaluation has failed to keep pace. That assumption has not been tested. We quantified the relationship between published output and registered prospective evaluation over eleven years.
Methods: Cross-sectional analysis of Clinical Trials.gov and PubMed covering complete calendar years 2015 to 2025. A pre-specified vocabulary of 21 AI terms, 48 general surgery terms and 24 exclusion terms was applied identically to both sources so that numerator and denominator were matched. Annual percentage change was estimated by Poisson regression with heteroscedasticity-consistent standard errors. The pre-specified primary hypothesis, that publications were growing faster than registered trials, was tested as a year-by-source interaction, the exponentiated coefficient giving the ratio of growth rates. A stricter exclusion screen was applied as a sensitivity analysis.
Results: Of 447 registry records retrieved, 145 met criteria (sensitivity analysis, 135). Over the period there were 6,818 publications, 145 registrations, 38 interventional trials and 24 randomised trials, a ratio of 179 publications Per interventional trial and 284 per randomized trial. No trial was registered before2018,by which time 476 papers had appeared. The primary hypothesis was not supported :publications grew by 37.3% per year (95%CI34.8to39.8)and interventional trials by 43.7% per year (95%CI27.4to62.2) ,giving a growth -rate Ratio of 0.955(95%CI0.845to1.079;p=0.461).Registered studies were predominantly observational (107/145, 73.8%),single-centre (125/145,86.2%) andnon-industry (142/145,97.9%); four(2.8%) recruited in more than one country and two of 38 interventional trials (5.3%) declared a phase. Median planned enrolment in interventional trials was 110 (IQR 50–266). Of 49 completed studies, one (2.0%) had posted results, at a median of 25 months after completion (IQR 19–31); 39 registrations (26.9%) had lapsed to unknown status.
Conclusion: The gap between publication and prospective evaluation in surgical AI is not widening. It is fixed at an extreme level and has been for a decade. Registered evaluation grew in proportion to the literature, but from a base so small, and of such uniform design weakness, that eleven years produced 24 randomised trials, none of which has reported results. The constraint on this field is not the volume of research but that almost none of it is designed to change practice, and almost none of what finishes is ever reported.

Abstract 0 | PDF Downloads 0

References

1. Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. 2019;25(1):44–56.
2. Hashimoto DA, Rosman G, RusD ,Meireles OR. Artificial intelligence in surgery: promises and perils. Ann Surg. 2018;268(1):70–6.
3. Maier-Hein L, Vedula SS, Speide lS, etal. Surgical data science for next-generation interventions. Nat Biomed Eng. 2017;1:691–6.
4. Nagendran M, Chen Y, Lovejoy CA,etal. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ. 2020;368:m689.
5. WuE, WuK, Danesh jouR, OuyangD, HoDE, ZouJ. How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals. Nat Med. 2021;27(4):582–4.
6. Collins GS, Moons KGM, Dhiman P,etal. TRIPOD + AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378.
7. Vasey B, Nagendran M, Campbell B,etal. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. 2022;28(5):924–33.
8. Liu X, Cruz Rivera S, Moher D, Calvert MJ, Denniston AK. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med.2020;26(9):1364–74.
9. Cruz Rivera S, Liu X, Chan AW, Denniston AK, Calvert MJ. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med.2020;26(9):1351–63.
10. Mascagni P, Vardazaryan A, Alapatt D,etal. Artificial intelligence for surgical safety: automatic assessment of the critical view of safety in laparoscopic cholecystectomy using deep learning. Ann Surg. 2022;275(5):955–61.
11. Meara JG, Leather AJM, Hagander L,etal. Global Surgery 2030: evidence and solutions for achieving health, welfare, and economic development. Lancet. 2015;386(9993):569–624.
12. Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. BMJ. 2015;350:g7594.
13. von Elm E, Altman DG, Egger M, Pocock SJ, Gøtzsche PC, Vandenbroucke JP. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement. Lancet.2007;370(9596):1453–7.
14. Ioannidis JPA. Why most published research findings are false. PLoS Med. 2005;2(8):e124.