AI-Driven Treatment Prediction for Diabetic Retinopathy: A Novel Deep Learning Approach with Cross-Population Validation
Amritha Bharath Ram
Dougherty Valley High School, San Ramon, United States of America
Publication date: July 10, 2026
Dougherty Valley High School, San Ramon, United States of America
Publication date: July 10, 2026
DOI: http://doi.org/10.34614/JIYRC2026I12
ABSTRACT
Diabetic retinopathy affects roughly 103 million people and is the leading cause of blindness in working-age adults, with the heaviest burden in low-income regions where specialists are scarce. Whereas most AI systems output a severity grade, this study predicts actionable treatment categories derived from severity grades to support frontline healthcare workers. A ResNet50 model, trained via transfer learning on the Mexican OCT-and-Eye-Fundus dataset, achieved 96.0% internal accuracy and a 0.97 weighted F1-score. On external validation using an independent Indian dataset (3,662 APTOS 2019 images), overall accuracy fell to 83.2%, with strong Monitoring recall (93%) but low recall (0.10, 0.20) for the treatment-requiring classes. This drop may partly reflect differences in disease presentation, grading, and treatment practices across the two regions. These results demonstrate a promising decision-support framework that requires further training and validation on larger, balanced, multi-institutional datasets before clinical use.
Diabetic retinopathy affects roughly 103 million people and is the leading cause of blindness in working-age adults, with the heaviest burden in low-income regions where specialists are scarce. Whereas most AI systems output a severity grade, this study predicts actionable treatment categories derived from severity grades to support frontline healthcare workers. A ResNet50 model, trained via transfer learning on the Mexican OCT-and-Eye-Fundus dataset, achieved 96.0% internal accuracy and a 0.97 weighted F1-score. On external validation using an independent Indian dataset (3,662 APTOS 2019 images), overall accuracy fell to 83.2%, with strong Monitoring recall (93%) but low recall (0.10, 0.20) for the treatment-requiring classes. This drop may partly reflect differences in disease presentation, grading, and treatment practices across the two regions. These results demonstrate a promising decision-support framework that requires further training and validation on larger, balanced, multi-institutional datasets before clinical use.