TY - JOUR
T1 - Transforming Systematic Reviews
T2 - Evaluating a Fine-Tuned Large Language Model for Abstract Screening in Uveitis and Retinal Vasculitis
AU - Cifuentes-González, Carlos
AU - Singer, Maxwell B.
AU - Rojas-Carabali, William
AU - Mejía-Salgado, Germán
AU - Cicinelli, Maria Vittoria
AU - Biswas, Jyotirmay
AU - Gangaputra, Sapna
AU - de-la-Torre, Alejandra
AU - Gupta, Vishali
AU - Pulido, Jose S.
AU - Agrawal, Rupesh
N1 - Publisher Copyright:
© 2026 American Academy of Ophthalmology, Inc. Published by Elsevier Inc. This is an open access article under the CC BY-NC-ND license. http://creativecommons.org/licenses/by-nc-nd/4.0/
PY - 2026/9
Y1 - 2026/9
N2 - Purpose: To evaluate the classification performance of UveAItis, a domain-specific large language model (LLM) fine-tuned for automated title and abstract screening in systematic reviews, using retinal vasculitis as a prototype. Design: Comparative evaluation study embedded within a registered systematic review and meta-analysis (PROSPERO: CRD42023489232). Subjects: A total of 1030 randomly selected articles from an initial search of 5533 records related to retinal vasculitis. Methods: Articles were independently screened by 2 uveitis experts (gold standard), final-year medical students, and 3 LLMs: UveAItis (fine-tuned Generative Pre-trained Transformer [GPT]-4o), base GPT-4o, and Claude Sonnet 3.5. Screening followed a 2-question binary logic regarding human subjects and primary empirical research design. Discrepancies were resolved through expert adjudication. Main Outcome Measures: Classification accuracy, sensitivity, specificity, area under the receiver operating characteristic curve, and Cohen Kappa coefficient for inter-rater agreement. Results: UveAItis achieved the highest performance with an accuracy of 93.3%, area under the curve (AUC) of 0.887, and Kappa of 0.77. It significantly outperformed base GPT-4o (AUC: 0.805, P = 0.021), Claude Sonnet 3.5 (AUC: 0.669, P < 0.0001), and medical students (AUC: 0.585, P < 0.00001). The fine-tuned model correctly identified 65.4% of expert-included articles postconsensus, whereas students only identified 22.3%. UveAItis also demonstrated the lowest rate of ambiguous “Need Consensus” outputs (3.4%) compared to experts (16.7%). Conclusions: UveAItis demonstrated expert-level performance, significantly outperforming general-purpose LLMs and nonexpert human reviewers. These findings validate the potential of domain-specific fine-tuning to enhance the efficiency, scalability, and reproducibility of evidence synthesis in specialized medical fields like ophthalmology. Financial Disclosure(s): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
AB - Purpose: To evaluate the classification performance of UveAItis, a domain-specific large language model (LLM) fine-tuned for automated title and abstract screening in systematic reviews, using retinal vasculitis as a prototype. Design: Comparative evaluation study embedded within a registered systematic review and meta-analysis (PROSPERO: CRD42023489232). Subjects: A total of 1030 randomly selected articles from an initial search of 5533 records related to retinal vasculitis. Methods: Articles were independently screened by 2 uveitis experts (gold standard), final-year medical students, and 3 LLMs: UveAItis (fine-tuned Generative Pre-trained Transformer [GPT]-4o), base GPT-4o, and Claude Sonnet 3.5. Screening followed a 2-question binary logic regarding human subjects and primary empirical research design. Discrepancies were resolved through expert adjudication. Main Outcome Measures: Classification accuracy, sensitivity, specificity, area under the receiver operating characteristic curve, and Cohen Kappa coefficient for inter-rater agreement. Results: UveAItis achieved the highest performance with an accuracy of 93.3%, area under the curve (AUC) of 0.887, and Kappa of 0.77. It significantly outperformed base GPT-4o (AUC: 0.805, P = 0.021), Claude Sonnet 3.5 (AUC: 0.669, P < 0.0001), and medical students (AUC: 0.585, P < 0.00001). The fine-tuned model correctly identified 65.4% of expert-included articles postconsensus, whereas students only identified 22.3%. UveAItis also demonstrated the lowest rate of ambiguous “Need Consensus” outputs (3.4%) compared to experts (16.7%). Conclusions: UveAItis demonstrated expert-level performance, significantly outperforming general-purpose LLMs and nonexpert human reviewers. These findings validate the potential of domain-specific fine-tuning to enhance the efficiency, scalability, and reproducibility of evidence synthesis in specialized medical fields like ophthalmology. Financial Disclosure(s): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
UR - https://www.scopus.com/pages/publications/105046576296
UR - https://www.scopus.com/pages/publications/105046576296#tab=citedBy
U2 - 10.1016/j.xops.2026.101296
DO - 10.1016/j.xops.2026.101296
M3 - Research Article
AN - SCOPUS:105046576296
SN - 2666-9145
VL - 6
JO - Ophthalmology Science
JF - Ophthalmology Science
IS - 9
M1 - 101296
ER -