Peer-reviewed large-scale, independent head-to-head study of seven AI algorithms for DR detection on over 20,000 consecutive patient visits from the Veterans Affairs (VA) Health Care System (HCS)
Overview
This independent multicenter, noninterventional comparison study of multiple AI was conducted on a total of 311,604 retinal images from 23,724 veterans who presented for teleretinal DR screening at the Veterans Affairs (VA) Puget Sound Health Care System (HCS) or Atlanta VA HCS from 2006 to 2018. Among 23 invited companies with automated AI-based DR screening systems, only five (5) companies agreed to participate and submitted seven algorithms for evaluation. Eyenuk submitted its EyeArt system which was identified as Algorithm G in the study.
Methods
The study analyzed the seven AI algorithms under the same analytical protocol by using the same retinal images. The sensitivity/specificity of each algorithm when classifying images as referable DR or not were compared with original VA teleretinal grades and a regraded arbitrated data set (7,379 images from 735 encounters). Value per encounter was estimated.
Results
The comparison of multiple AI algorithms using the same protocol on the same retinal images, identified the best performing algorithm (Algorithm G, which was the EyeArt technology from Eyenuk).
- Algorithm G was the only algorithm that was statistically indistinguishable from the standard of care
- Most notably, Algorithm G did not miss a single case of moderate or severe non-proliferative or proliferative DR in the arbitration set, achieving 100% sensitivity for each (Figure 2 from the publication)
- Algorithm G also enabled the highest amount of cost savings to the VA teleretinal screening program
Conclusion
Authors attempt to make the following conclusions from the study in the publication.
- “The DR screening algorithms showed significant performance differences… Although some algorithms in our study performed well from a screening perspective, others would pose safety concerns.”: We agree that all AI is not created equal. FDA cleared systems have gone through rigorous prospective validation against gold standard (ETDRS) clinical reference standard in intended use settings and are expected to perform well
- “These results argue for rigorous testing of all such algorithms on real-world data before clinical implementation.”: Authors are making this general conclusion based on poor performance of some algorithms, but our view is that for systems that have FDA clearance after already going through more rigorous prospective clinical trial validation, additional testing is not necessary
Click here to read a detailed analysis by Eyenuk.
Link to full study: https://diabetesjournals.org/care/article/44/5/1168/138752/Multicenter-Head-to-Head-Real-World-Validation