Usability of Large Language Models for Developing Exercise-Based Rehabilitation Programs in Patients With Stable Angina
INTERNATIONAL JOURNAL OF CLINICAL PRACTICE, cilt.2026, sa.1, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 2026 Sayı: 1
- Basım Tarihi: 2026
- Doi Numarası: 10.1155/ijcp/1434301
- Dergi Adı: INTERNATIONAL JOURNAL OF CLINICAL PRACTICE
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, Directory of Open Access Journals, Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
- Hatay Mustafa Kemal Üniversitesi Adresli: Evet
Özet
Background: While artificial intelligence (AI) and large language models (LLMs) are increasingly used in exercise-based cardiac rehabilitation (CR), the clinical reliability and adherence to guidelines of AI-generated prescriptions require systematic comparison with expert-designed protocols. Objective: This study aimed to evaluate the heterogeneity between therapeutic exercise-based rehabilitation programs generated by LLMs (ChatGPT and Gemini Advanced) and those developed by healthcare professionals for individuals with stable angina and chronic coronary syndrome (CCS). Methods: Exercise-based rehabilitation prescriptions generated by ChatGPT and Gemini Advanced were compared with those developed by an expert panel across key clinical domains, including dietary recommendations, resistance-training prescription method, lower-extremity strengthening decisions, and aerobic exercise intensity selection. Results: Significant heterogeneity was identified between AI-generated and expert panel-prescribed rehabilitation programs across all evaluated domains. Statistically significant differences were observed for dietary recommendations (chi(2) = 80.00, p < 0.001), resistance-training prescription method (chi(2) = 20.16, p < 0.001), lower-extremity strengthening decisions (chi(2) = 54.07, p < 0.001), and aerobic exercise intensity selection (chi(2) = 49.24, p < 0.001). Agreement analyses demonstrated poor-to-absent concordance for most clinically relevant decisions. Fleiss' kappa values indicated significant disagreement for dietary recommendations (kappa = -0.500, p < 0.001), flexibility training (kappa = -0.500, p < 0.001), respiratory muscle training (RMT) (kappa = -0.500, p < 0.001), and rating of perceived exertion (RPE) based safety warnings (kappa = -0.364, p = 0.001). Pairwise analyses showed perfect agreement only for selected isolated parameters (diet recommendations between the healthcare Professional and ChatGPT; flexibility training and RMT between the expert panel and Gemini Advanced), whereas no systematic agreement was observed for lower-extremity strengthening, depression alerts, or RPE warnings. Overall, LLMs generated structurally consistent outputs in some domains but demonstrated substantial variability in exercise prescription and safety-related clinical decisions. Conclusion: Current LLMs may serve as supportive tools for structuring rehabilitation recommendations; however, their outputs showed limited agreement with expert-designed programs and inconsistent adherence across key components of comprehensive CR. Given the marked variability in exercise prescription, safety alerts, and guideline-concordant recommendations, LLM-generated programs should be interpreted within a human-in-the-loop framework and not used as autonomous prescribing systems. Further development, guideline integration, and prospective clinical validation are required before routine implementation in CR practice.