UTILITY OF LARGE LANGUAGE MODELS (LLMS) IN PREOPERATIVE RISK STRATIFICATION: A COMPARATIVE STUDY WITH HUMAN ANESTHESIOLOGISTS
Main Article Content
Keywords
Artificial Intelligence, Perioperative Medicine, Large Language Models, Risk Stratification, Patient Safety.
Abstract
Background: Preoperative risk stratification is critical for mitigating perioperative complications, yet it remains subject to inter-observer variability among clinicians. Large Language Models (LLMs) have demonstrated sophisticated reasoning in medical contexts, but their utility in synthesizing complex patient data for anesthesia risk assessment remains insufficiently validated against clinical expertise. This study evaluates the accuracy and consistency of LLMs compared to board-certified anesthesiologists in assigning risk scores and identifying perioperative red flags.
Methods: A comparative diagnostic study was conducted using 100 high-fidelity simulated patient vignettes representing diverse surgical subspecialties. Each vignette included medical history, physical examination findings, and laboratory data. Three state-of-the-art LLMs ( GPT-4, Claude 3.5, and Gemini Pro) and five board-certified anesthesiologists independently assigned American Society of Anesthesiologists (ASA) Physical Status scores and identified top-tier perioperative risks. The primary outcome was the level of agreement (Cohen’s Kappa) between LLMs and the "gold standard" expert consensus. Secondary outcomes included the sensitivity of LLMs in detecting "critical" contraindications (severe aortic stenosis).
Results: Preliminary analysis indicates that top-tier LLMs demonstrate "substantial" agreement (k > 0.70) with human experts in ASA classification. While LLMs excelled in extracting data from unstructured text, they exhibited a higher rate of "over-stratification" in low-risk cases compared to human clinicians. However, LLMs demonstrated a 98% sensitivity in identifying objective contraindications, occasionally outperforming humans in detecting minor laboratory abnormalities.
Conclusion: LLMs serve as a potent adjunct for preoperative screening, offering high sensitivity for risk factor identification. While they do not replace the clinical nuance of an anesthesiologist, their integration into Electronic Health Records could provide a valuable "safety net" for identifying high-risk patients and standardizing preoperative care.
References
2. Bohr, A., & Memarzadeh, K. The Rise of Artificial Intelligence in Healthcare Applications. Academic Press.
3. Cascella, M., et al. The Role of ChatGPT and Large Language Models in Anesthesiology and Perioperative Medicine: A Systematic Review. Journal of Clinical Medicine, 13(4), 1102.
4. Chen, J. S., et al. Comparative Accuracy of GPT-4o and Claude 3.5 in Preoperative Triage: A Multicenter Evaluation. Digital Health & Anesthesia, 12(1), 45-58.
5. Deyrup, A. T., & Gupta, P. Artificial Intelligence in Clinical Decision Support: Assessing the Nuance of Surgical Risk. The Lancet Digital Health, 6(2), e112-e120.
6. Emanuel, E. J., & Wachter, R. M. Artificial Intelligence in Health Care: Will the Human-in-the-Loop Model Prevail? JAMA, 331(14), 1181-1182.
7. Feuerriegel, S., et al. Generative AI in Healthcare: Opportunities and Challenges. Nature Medicine, 30, 15–26.
8. Glance, L. G., et al. The ASA Physical Status Classification: Re-evaluating Inter-rater Reliability in the Age of Digital Assistants. Anesthesiology, 139(3), 310-322.
9. Gombar, S., et al. Scaling AI Deployment in Perioperative Pathways: Lessons from the EHR Integration. NEJM AI, 1(3), 100-115.
10. He, J., et al. Zero-Shot vs. Few-Shot Prompting for Medical Diagnosis: An Empirical Study with Large Language Models. NPJ Digital Medicine, 8, 22.
11. Kannan, S., et al. Hallucination Rates in Medical LLMs: A Comparative Analysis of GPT-4, Claude, and Gemini. Journal of Medical Systems, 48, 114.
12. Lee, P., et al. The AI Revolution in Medicine: GPT-4 and Beyond. Pearson Education.
13. Mesko, B., & Topol, E. J. The Evolution of the Digital Anesthesiologist: From Monitoring to Mentoring. The Lancet, 405(10482), 882-885.
14. Neylan, C. J., et al. Predictive Analytics for Postoperative Complications: A Comparison of Machine Learning and Clinical Intuition. Annals of Surgery, 279(1), 14-22.
15. Rajpurkar, P., & Kohane, I. S. AI in Health Care: A Comprehensive Guide. MIT Press.
16. Singhal, K., et al. Large Language Models Encode Clinical Knowledge. Nature, 620, 172–180.
17. Wornow, M., et al. The Health System Ecosystem for AI: Preoperative Triage and Resource Allocation. Journal of Biomedical Informatics, 150, 104592.
18. Zhang, Y., et al. Ethical Constraints and Legal Liability of AI-Assisted Anesthesia Planning. British Journal of Anaesthesia, 134(5), 602-610.
19. American Society of Anesthesiologists. ASA Physical Status Classification System: Update on Comorbidity Guidelines. ASA Publications.

