Automated and generalized system for detecting diabetic retinopathy in ultra-widefield retinal images across diverse ethnicities and regions

The automated UWF-SLO system with machine learning algorithms addresses interpretability and generalizability issues, enabling efficient diabetic retinopathy detection and categorization across diverse populations with high accuracy and user-friendly integration into clinical workflows.

WO2026046321A1PCT designated stage Publication Date: 2026-03-05THE HONG KONG UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/117683
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-29
Filing Date
2025-08-29
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Current automated systems for diabetic retinopathy detection face challenges such as lack of interpretability, limited generalizability across diverse populations, and impractical integration into clinical workflows, particularly due to the use of conventional fundus photography that misses peripheral lesions and lacks user-friendly interfaces.

Method used

An automated system utilizing ultra-widefield scanning laser ophthalmoscopy (UWF-SLO) for high-resolution retinal imaging, combined with machine learning algorithms and visualization tools, capable of categorizing diabetic retinopathy severity and providing interpretable outputs, which can be integrated into clinical workflows without pharmacological pupil dilation.

Benefits of technology

The system achieves robust, real-time diabetic retinopathy detection and categorization across diverse populations, enhancing early detection and triage of retinal abnormalities with high accuracy and user-friendly integration into clinical settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025117683_05032026_PF_FP_ABST
    Figure CN2025117683_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Described herein is an automated diabetic retinopathy (DR) detection and categorization system is developed using machine learning models trained on ultra-widefield scanning laser ophthalmoscopy (UWF-SLO) images. The system adopts a cascaded structure comprising three ResNet50-based models to sequentially assess image gradability, detect referable DR (RDR), and identify vision-threatening DR (VTDR). The present model incorporates transfer learning, multilayer perceptron (MLP) output layers, and optimized down-sampling and augmentation techniques to enhance classification accuracy. A graphical user interface (GUI) platform is also developed to streamline clinical deployment, providing interpretable predictions and confidence scores. This system demonstrates high generalizability across eye images from diverse ethnicities and geographical regions, showcasing its practical utility for real-world DR screening, particularly in non-mydriatic settings using UWF-SLO imaging, thereby supporting early detection and timely referral in diabetic eye care.
Need to check novelty before this filing date? Find Prior Art

Description

AUTOMATED AND GENERALIZED SYSTEM FOR DETECTING DIABETIC RETINOPATHY IN ULTRA-WIDEFIELD RETINAL IMAGES ACROSS DIVERSE ETHNICITIES AND REGIONSInventors: Ahmed Abdul QUADEER; Yim Lui Carol CHEUNG; Dawei YANG; Muhammad Saqib SOHAIL; and Peifeng WUCross-Application to Relevant Applications:

[0001] The present application claims priority from U.S. Provisional Patent application number 63 / 688,856 filed August 29, 2024; the disclosure of which is incorporated by reference herein in its entirety.Field of the Invention:

[0002] The present invention is generally related to diabetic retinopathy system. More specifically, the present invention is related to automated systems for diabetic retinopathy detection and categorization.Background of the Invention:

[0003] Diabetic retinopathy (DR) poses a critical global healthcare burden, being one of the most common microvascular complications associated with diabetes mellitus. As of 2022, diabetes affected estimated 537 million individuals worldwide, with projections indicating further increases in the coming years. Approximately 27%of those with diabetes experienced some form of DR, a condition that can progress to irreversible vision impairment if not identified and managed in a timely manner.

[0004] The conventional approach for evaluating DR involves manual inspection of retinal images to identify characteristic features such as hemorrhages, microaneurysms, non-perfusion, neovascularization, and retinal edema. This diagnostic process requires highly specialized training and is typically carried out by ophthalmologists or retinal specialists. However, the availability of such professionals is limited and unevenly distributed, especially in regions where the burden of diabetes is greatest, including large parts of Asia and Africa. The resulting gap in early detection and monitoring leads to increased risk of preventable vision loss in affected populations.

[0005] In response, automated image analysis techniques have been explored as a means to support retinal image interpretation. Early methods relied on classical machine learning algorithms such as support vector machines and random forests, which necessitated the laborious design of hand-crafted features and often demonstrated limited generalizability.

[0006] More recently, machine learning (ML) methods, particularly those that utilize convolutional neural networks (CNNs) , have shown promising results by learning discriminative image features directly from training data without manual feature extraction. These ML-based methods have reported high accuracy in classifying retinal images for clinical deployment.

[0007] Nevertheless, significant obstacles remain, hindering the widespread adoption of machine learning systems in routine clinical workflows. One of the major challenges is the lack of interpretability often associated with ML models, commonly referred to as the “black-box” problem. Medical practitioners are typically hesitant to rely on predictions from opaque algorithms without insight into the rationale behind their outputs. To address this, recent research has incorporated model explainability tools, such as saliency maps and class activation visualizations, to highlight image regions most relevant to model decision.

[0008] Another shortcoming is that many existing ML systems have been developed using conventional fundus photography, which captures only a 30–50 degree field of the retina, often missing peripheral lesions that are critical for comprehensive assessment of retinal disease.

[0009] A newer imaging modality, ultra-widefield scanning laser ophthalmoscopy (UWF-SLO) , enables visualization of up to 200 degrees of the retina in a single image, capturing peripheral pathology that may otherwise go undetected. UWF-SLO also offers high-resolution and high-contrast images without the need for pupil dilation, increasing patient comfort and compliance. These advantages render UWF-SLO a promising modality for enhancing image-based retinal analysis, particularly in detecting abnormalities extending into the far periphery.

[0010] Despite these technological advances, current image analysis systems still lack applicability in practical clinical settings. Many models demonstrate reduced performance when applied to datasets from different geographic or ethnic populations, limiting their generalizability. Also, most systems are designed without user-friendly interfaces or real-time decision support capabilities, thereby rendering them less practical if integrated into existing clinical workflows.

[0011] As such, there remains a need for an image analysis system that robustly processes UWF-SLO images, provides interpretable and visualized outputs that can be exported and integrated into existing clinical workflows, and generalizes well across diverse populations. The present invention addresses this need.Summary of Invention:

[0012] In the first aspect of the present invention, a system for automated classification and visual interpretation of ultra-widefield retinal images of a subject is provided herewith. The system comprises an ultra-widefield scanning laser ophthalmoscopy imaging device configured to capture a retinal image of at least 150 degrees field of view of the subject’s eye, and a computing device comprising at least one processor and one non-transitory memory.

[0013] Specifically, the non-transitory memory of the system stores instructions prompting the process to execute a series of commands. These commands include: (i) pre-processing a retinal image of the subject’s eye captured by the ultra-widefield scanning laser ophthalmoscopy imaging device to readjust resolution and intensity values and applying data augmentations, (ii) determining whether the retinal image is gradable for diagnostic purposes based on image quality features, (iii) categorizing the retinal image into one of four diabetic retinopathy severity categories by detection and analysis of retinal lesion regions by one or more classification neural networks, (iv) generating a visual activation map highlighting pixels relating to the retinal lesion regions, and (v) outputting a diagnostic result comprising a categorization label and confidence score to a graphical user interface.

[0014] In an embodiment of the first aspect of the present invention, the four diabetic retinopathy severity categories are (i) no diabetic retinopathy (no DR) , (ii) non-referable diabetic retinopathy (non-DR) , (iii) referable diabetic retinopathy (RDR) , and (iv) vision-threatening diabetic retinopathy (VTDR) .

[0015] In another embodiment, the retinal image of the subject is captured without requiring pharmacological pupil dilation administered to the subject.

[0016] In another embodiment, the ultra-widefield scanning laser ophthalmoscopy imaging device captures the retinal image of the subject’s eye at a resolution of at least 2600 × 2048 pixels.

[0017] In one embodiment, the pre-processing of the retinal image comprises resizing the retinal image to a resolution in a range of 500 × 500 pixels to 700 × 700 pixels.

[0018] In another embodiment, the data augmentations comprise one or more of random horizontal flipping, Gaussian blurring and affine transformation.

[0019] In yet another embodiment, the determining of gradability of the retinal image is performed by a convolutional neural network trained to identify blurring, occlusion, dark artifacts and image coverage of a standardized retinal grid.

[0020] In another embodiment, the categorizing the retinal image into one of four diabetic retinopathy severity categories is performed by a cascaded ML architecture comprising a first neural network configured to detect VTDR, and a second neural network configured to detect RDR from non-DR images.

[0021] In a further embodiment, each of the first neural network and the second neural network comprises a convoluted neural network backbone pre-trained on a general-purpose image dataset and fine-tuned on retinal image data, and a multilayer perceptron having a rectified linear unit-activated hidden layer, a dropout layer, and a final output layer generating class probabilities.

[0022] In another embodiment, the graphical user interface is configured to display the original retinal image, the visual activation map highlighting pixels relating to the retinal lesion regions as an overlay, diagnostic result comprising a categorization label and confidence score, and at least one export option for exporting and integrating the original retinal image, the visual activation map and the diagnostic result to an external clinical information system.

[0023] In another embodiment, the confidence score is based on the softmax probability output from a final convolutional layer of the one or more classification neural networks and is displayed as a numeric value between 0 and 1.

[0024] In yet another embodiment, the generating a visual activation map comprises applying a gradient-weighted class activation mapping (Grad-CAM) technique to a final convolutional layer of the one or more classification neural networks.

[0025] In yet other embodiment, the non-transitory memory further stores an instruction to export diagnostic results to an external electronic medical record system.

[0026] In another embodiment, the retinal lesions comprise hemorrhages, microaneurysms, non-perfusion and neovascularization.

[0027] In another embodiment, the computing device is further configured to perform real-time inference within 5 seconds of image acquisition.

[0028] In a second aspect of the present invention, a method of automated classification and visual interpretation of an ultra-widefield retinal image is also provided herewith. The method comprises: (i) capturing a retinal image of a subject using an ultra-widefield scanning laser ophthalmoscopy (UWF-SLO) imaging device, (ii) pre-processing the retinal image by readjusting resolution and intensity values of the image and applying data augmentations, (iii) determining whether the retinal image is gradable for diagnostic purposes based on image quality features, (iv) categorizing the retinal image into one of four diabetic retinopathy severity categories by detection and analysis of retinal lesion regions by one or more classification neural networks, (v) generating a visual activation map highlighting pixels relating to the retinal lesion region, and (vi) providing a diagnostic result comprising the original retinal image, the visual activation map, a categorization label and a confidence score to a graphical user interface.

[0029] In an embodiment of the second aspect of the present invention, the four diabetic retinopathy severity categories are (i) no diabetic retinopathy (no DR) , (ii) non-referable diabetic retinopathy (non-DR) , (iii) referable diabetic retinopathy (RDR) , and (iv) vision-threatening diabetic retinopathy (VTDR) .

[0030] In another embodiment, the retinal image of the subject is captured without requiring pharmacological pupil dilation administered to the subject.

[0031] In one embodiment, the ultra-widefield scanning laser ophthalmoscopy imaging device captures the retinal image of the subject’s eye at a resolution of at least 2600 × 2048 pixels.

[0032] In another embodiment, the pre-processing of the retinal image comprises resizing the retinal image to a resolution in a range of 500 × 500 pixels to 700 × 700 pixels.

[0033] In yet another embodiment, the data augmentations comprise one or more of random horizontal flipping, Gaussian blurring and affine transformation.

[0034] In another embodiment, the determining of gradability of the retinal image is performed by a convolutional neural network trained to identify blurring, occlusion, dark artifacts and image coverage of a standardized retinal grid.

[0035] In another embodiment, the categorizing the retinal image into one of four diabetic retinopathy severity categories is performed by a cascaded ML architecture comprising a first neural network configured to detect VTDR, and a second neural network configured to detect RDR from non-DR images.

[0036] In a further embodiment, each of the first neural network and the second neural network comprises a convoluted neural network backbone pre-trained on a general-purpose image dataset and fine-tuned on retinal image data, and a multilayer perceptron having a rectified linear unit-activated hidden layer, a dropout layer, and a final output layer generating class probabilities.

[0037] In another embodiment, the confidence score is based on the softmax probability output from a final convolutional layer of the one or more classification neural networks and is displayed as a numeric value between 0 and 1.

[0038] In yet another embodiment, the generating a visual activation map comprises applying a gradient-weighted class activation mapping (Grad-CAM) technique to a final convolutional layer of the one or more classification neural networks.

[0039] In another embodiment, the retinal lesions comprise hemorrhages, microaneurysms, non-perfusion and neovascularization.

[0040] In another embodiment, the computing device is further configured to perform real-time inference within 5 seconds of image acquisition.Brief Description of the Drawings:

[0041] FIG. 1 tabulates the imaging device and image resolution for each dataset.

[0042] FIG. 2 provides an overview of all data used for training and testing the ML system for diabetic retinopathy diagnosis in the present invention.

[0043] FIG. 3 provides a schematic overview of the diabetic retinopathy detection process.

[0044] FIG. 4 provides a schematic overview of the multilayer perception (MLP) output layer used in the RestNet50 model.

[0045] FIG. 5 tabulates the gradability performance of the method in the preset invention in both primary and external datasets with an image size 512*640 using MLP ResNet50.

[0046] FIG. 6 tabulates the RDR performance of the method in the preset invention in both primary and external datasets with an image size 512*640 using MLP ResNet50.

[0047] FIG. 7 tabulates the Vision-threatening Diabetic Retinopathy (VTDR) performance of the method in the preset invention in both primary and external datasets with an image size 512*640 using MLP ResNet50.

[0048] FIG. 8 tabulates the average Area Under the Receiver Operating Characteristic Curve (AUROC) scores comparison of models developed with different image sizes. For each image size, an average AUROC score was calculated across all testing datasets to measure the overall performance of each task (gradability, RDR, and VTDR) .

[0049] FIG. 9 tabulates the average Area Under the Precision-Recall Curve (AUPRC) scores comparison of models developed with different image sizes. For each image size, an average AUROC score was calculated across all testing datasets to measure the overall performance of each task (gradability, RDR, and VTDR) .

[0050] FIG. 10 tabulates the average AUROC comparison of models developed with different augmentations. For each set of augmentation, an average AUROC score was calculated across all testing datasets to measure the overall performance of each task (gradability, RDR, and VTDR) .

[0051] FIG. 11 tabulates the average AUPRC comparison of models developed with different augmentations. For each augmentation, an average AUPRC score was calculated across all testing datasets to measure the overall performance of each task (gradability, RDR, and VTDR) .

[0052] FIG. 12 tabulates the average AUROC comparison of models developed with different backbone structures. For each backbone structure, an average AUROC score was calculated across all testing datasets to measure the overall performance of each task (gradability, RDR, and VTDR) .

[0053] FIG. 13 tabulates the average AUPRC comparison of models developed with different backbone structures. For each backbone structure, an average AUPRC score was calculated across all testing datasets to measure the overall performance of each task (gradability, RDR, and VTDR) .

[0054] FIG. 14 provides a schematic illustration of the workflow of the system of the present invention.

[0055] FIG. 15 is a schematic depiction of the computation of the confidence score.

[0056] FIGs. 16A to 16C demonstrate the probability calibration results of the model output with two methods, namely logistic regression and isotonic regression. FIG. 16A shows the calibration curves of the gradability primary test dataset. FIGs. 16B and 16C are empirical probability density plots for isotonic regression and logistic regression respectively post-calibration.

[0057] FIG. 17 demonstrates the confidence score box plots for the gradability external 1 dataset before and after calibration, respectively.

[0058] FIG. 18 shows an exemplary layout of the GUI-based software platform of the present invention.Detailed Description:

[0059] The present invention addresses the above-mentioned need for an image analysis system that satisfies all of: (i) robust processing and analysis of UWF-SLO images for diabetic retinopathy categorization, (ii) providing interpretable and visualized outputs which are exportable and subsequently integrated into existing clinical workflows, and (iii) generalizing well across diverse populations.

[0060] The present invention is capable of achieving the above through a technical framework that integrates ML-based image classification with visualization tools and a graphical user interface. In accordance with the various embodiments, the systems and methods provided by the present invention enable reliable, automated analysis of UWF-SLO images, with outputs including classification results, confidence scores, and region-specific visual explanations. They are designed for practical deployment in clinical settings and has been evaluated across multiple datasets representing varied ethnic and regional backgrounds.

[0061] By combining high-resolution widefield imaging, advanced machine learning, and explainability-enhanced outputs in a unified software platform, the present invention offers a practical and scalable tool to support retinal image interpretation, facilitating early detection and triage of retinal abnormalities in diverse healthcare environments.

[0062] Automated methods for detecting diabetic retinopathy (DR) rely on properly labeled images. DR is classified into five stages based on the International Clinical Diabetic Retinopathy Disease Severity Scale: no DR, mild, moderate, severe non-proliferative diabetic retinopathy (NPDR) , and proliferative diabetic retinopathy (PDR) .

[0063] Referable diabetic retinopathy (RDR) refers to cases that need to be evaluated by an eye specialist. This includes images showing moderate NPDR or more severe stages, or signs of diabetic macular edema (DME) .

[0064] Vision-threatening diabetic retinopathy (VTDR) is a more serious condition that can lead to significant vision loss or require urgent treatment. VTDR includes severe NPDR, PDR, or the presence of DME.

[0065] DME occurs when fluid leaks from damaged blood vessels and causes swelling in the macula, which is central part of the retina. It may happen at any stage of DR and affects the part of the eye responsible for sharp central vision.

[0066] Image gradability refers to the ability of an image to be evaluated by a doctor or an expert, which is essential for diagnosing eye diseases such as DR. An image may have high quality, with clear and focused arcades and vessels, but it may not be gradable if it does not show some key anatomical features, such as the optic nerve and the fovea.

[0067] The systems and methods in accordance with the embodiments of the present invention for analysis and detection of DR in UWF-SLO images employ a cascade of three ML neural networks (NN) , including, but not limited to, Convoluted Neural Networks (CNNs) , each with a specific function. The first NN assesses image gradability, the subsequent one detects VTDR, and the final NN identifies RDR. Images are sequentially processed by the NNs until classified as ungradable, VTDR, RDR, or non-RDR.

[0068] An additional advantage of the present invention is its integration of ultra-widefield scanning laser ophthalmoscopy (UWF-SLO) imaging, which is able to capture a high-resolution and wide-angle retinal image of at least 150 degrees field of view of the subject’s eye without the need of pupil dilation induction to the subject.

[0069] Traditional retinal imaging involves mydriasis (i.e., pupil dilation) through the administration of special dilation drops. These dilation drops (i.e., tropicamide and phenylephrine) , through temporarily paralyzing the sphincter muscle of the iris or stimulating the dilator muscle of the iris, induces mydriasis in the patient.

[0070] However, the mydriasis effect takes time (typically 15-30 minutes) to show but nonetheless may subsequently last for hours, which may bring discomfort and short-term visual impairment to subjects. Rare cases have also been reported that these dilation drops may cause allergic reactions, increased intraocular pressure and systemic side-effects from anticholinergic agents such as tropicamide. It should also be taken into account that, for the elderly people who are more susceptible to diabetes and require diabetic retinopathy examination, they may experience longer recovery times as their bodies generally display slower drug clearance.

[0071] As such, the present invention is highly efficient as retinal images are taken from patients without the need to wait for the onset time of the mydriasis, and the image processing and analysis are configured to be performed within seconds with high precisions.

[0072] Below are further elaborations of the embodiments of the present invention, which are specifically designed to detect DR in UWF-SLO images.

[0073] EXAMPLES

[0074] Performance Metrics

[0075] The performance of the ML system is rigorously validated using several key performance metrics:

[0076] 1. Accuracy –it measures the proportion of correctly predicted cases out of the total cases. Its mathematical expression is:

[0077] 2. Sensitivity (recall) –it measures the proportion of the actual positive cases correctly predicted by the model. Its mathematical expression is:

[0078] 3. Specificity –it measures the proportion of the actual negative cases correctly predicted by the model. Its mathematical expression is:

[0079] 4. AUROC (Area Under the Receiver Operating Characteristic Curve) –it quantifies the model’s ability to distinguish between positive and negative cases across different probability thresholds. Its mathematical expression is:

[0080] 5. F1 score –it measures the balance between the precision and the recall. It’s a harmonic mean of the precision and the recall. Its mathematical expression is:

[0081] TP, TN, FN, FP denotes true positive, true negative, false negative, false positive, respectively.

[0082] Data sources and labelling

[0083] The ML system is developed using UWF-SLO images obtained from the Chinese University of Hong Kong Sight Threatening Diabetic Retinopathy study conducted between 2015 and 2019. This dataset, denoted as the primary dataset, consists of images collected from subjects with diabetic mellitus and no other eye diseases at the Hong Kong Eye Hospital. A total of 4, 083 images from 413 patients with diabetes are enrolled in this study. The primary dataset is divided into the development set and the validation set, with a ratio of 5: 1.

[0084] In addition to the primary dataset, the ML system is tested on four External datasets of UWF-SLO images from diverse sources (see FIG. 1) . External Dataset 1 contained 1, 320 images collected from the Moorfields Biomedical Research Centre in London, UK. Although this dataset lacks DR severity information, it includes image quality labels, making it suitable for gradability model assessment. External Dataset 2 contains 3, 876 images collected from the retina clinic at the Sankara Nethralaya Hospital in Chennai, India. External Dataset 3 contains 217 images collected from the Giridhar Eye Institute in Kochi, India. External Dataset 4 contains 95 images collected from the Diagnostic Ophthalmology Centre in Buenos Aires, Argentina.

[0085] For assessing the performance of the referable diabetic retinopathy (RDR) and vision-threatening diabetic retinopathy (VTDR) models, only External Datasets 2, 3, and 4 are used, as External Dataset 1 does not provide DR severity information. All External datasets utilize Optos devices with different models and resolutions: External Datasets 1 and 2 used Daytona Plus P200Tx with 2,600 x 2,048 pixels, while External Datasets 3 and 4 used Daytona with 3,698 x 2,448 pixels.

[0086] Referring to FIG. 2, a training set of 3,319 images is used to train the gradability model, 2,592 gradable images are used to train the RDR and VTDR presence detection network. A primary test set and four External test sets are used to evaluate model performance. All 764 images in the primary test set are used to evaluate gradability model, and 572 gradable images are used to test the RDR and VTDR model. All 1,320 images in External dataset 1 are used to evaluate gradability model. All 3,876 images in External dataset 2 are used to evaluate gradability model and 3,862 gradable images were used to test the RDR and VTDR model. All 217 images in External dataset 3 are used to test gradability model and 209 gradable images are used to test RDR and VTDR model. All 95 images in External dataset 4 are used to test gradability model and 78 gradable images are used to test RDR and VTDR model.

[0087] Each image is labelled with three classifications through fundus examination to denote the image gradability, RDR and VTDR. Gradability indicates whether the image is suitable for analysis. An image is considered ungradable if it lacked a clear retina surface. Such ungradable images typically suffer from blurriness, darkness, or artifacts. To categorize gradability, retina experts employ the standard 7-field Early Treatment Diabetic Retinopathy Study (ETDRS) grid. The severity of DR is graded into five classes by the International Clinical Diabetic Retinopathy Disease Severity Scale: No DR, mild non-proliferative DR (mild NPDR) , moderate NPDR, severe NPDR, and proliferative DR (PDR) . Referable DR (RDR) is defined as moderate NPDR or worse, while Vision-threatening DR (VTDR) is defined as severe NPDR or worse.

[0088] System design

[0089] Referring to FIG. 3 for a schematic overview of the DR detection process executed by the ML system. The suite of generalizable models in the ML system is tested for detecting the varying degrees of DR severity. This suite comprises three specialized models: one dedicated to detecting RDR, another focused on VTDR and a third for assessing gradability. The NNs executing these models are structured in a cascaded manner. First, the gradability assessment model NN 401 determines whether the images were gradable or not. Subsequently, the gradable images are fed to the VTDR model NN 402 to detect the presence of VTDR. Since a VTDR image must also be an RDR image, the non-VTDR images that are either DR absent or moderate NPDR are then fed to the RDR model NN 403 to detect the presence of moderate NPDR. Each UWF-SLO image is processed by the ML system until its class was identified (i.e., ungradable, VTDR, RDR, or non-RDR) .

[0090] For training purposes, all 3, 319 images in the primary training dataset are used to train the gradability assessment model, which performed binary classification to determine whether an image is gradable or not. Among the 2, 592 gradable images in the primary training dataset, 2, 289 are labelled with RDR and VTDR information. Consequently, only these 2, 289 images are utilized to train the RDR and VTDR models, both of which perform binary classification to detect the presence of RDR or VTDR.

[0091] Each NN in the ML system utilizes a ResNet50 architecture. To enhance the performance of the ResNet50 model, an output layer comprising a multilayer perceptron (MLP) with dropout is appended. This layer extracts high-level features, reduces overfitting, and increases generalization performance. It has a fully-connected layer with a ReLU activation, a dropout function with a 0.2 rate, and another fully-connected layer that produces a two-element vector. The vector indicates the probabilities of the negative and positive classes, respectively. This output format is chosen to apply a cross-entropy loss with manual class weights as the loss function, which can address the imbalance class issue. An overview of the MLP output layer is shown in FIG. 4.

[0092] Specifically, in this investigation, transfer learning is applied using the world’s largest object detection dataset, ImageNet, which contains over 14 million images.

[0093] Data processing

[0094] Raw UWF-SLO images, which are unprocessed and uncompressed data with high resolutions, are very large and complex. To speed up training and save storage space, each UWF-SLO image is resized to 512 x 640 before being fed to the system. To prevent size mismatch in the average pooling layer, adaptive average pooling is used to output a feature with 1 unit. The image also undergoes augmentations such as horizontal flipping, affine transformation, and Gaussian blur.

[0095] Three ResNet50 models are trained using the following specifications:

[0096] Data pre-processing: · Image Resizing: Each RGB image is resized to 512x640 · Normalization: the default ImageNet weight normalization is applied, with  the mean set to (0.485, 0.456, 0.406) and the standard deviation set to (0.229, 0.224, 0.225)

[0097] Data augmentation: · Random Horizontal Flipping: the images are randomly flipped horizontally  with a probability of 0.5 · Gaussian Blur: Images underwent Gaussian blur with sigma values of (0.1,  2.0)

[0098] DR detection results

[0099] The results of the DR detection process executed by the ML system are presented, in which a ResNet50 architecture with an MLP layer as the output layer instead of a standard fully connected layer is used. An input image size of 512×640 and a combination of training augmentation techniques (Horizontal flipping, Affine transformation and Gaussian blur) that differ from Phoom et al. ’s setting are also used. In the gradability assessment performance as described below, RDR detection performance and VTDR detection performance of the ML system are shown, respectively. The gradability assessment network is tested using all 6, 727 images in the Primary test dataset and the External datasets.

[0100] Referring to FIG. 5, the results of the gradability assessment model validation are summarized. An AUROC of 0.839, Sensitivity of 0.859, Specificity of 0.818, F1 score of 0.899, and Accuracy of 0.850 in the Primary dataset are achieved by the model. In the External dataset 1, an AUROC of 0.760, Sensitivity of 0.717, Specificity of 0.802, F1 score of 0.770, and Accuracy of 0.754 are achieved by the model. In the External 2 dataset, an AUROC of 0.923, Sensitivity of 0.917, Specificity of 0.928, F1 score of 0.956, and Accuracy of 0.917 are achieved by the model. In the External 3 dataset, an AUROC of 0.906, Sensitivity of 0.937, Specificity of 0.875, F1 score of 0.965, and Accuracy of 0.935 are achieved by the model. In the External 4 dataset, an AUROC of 0.822, Sensitivity of 0.802, Specificity of 0.842, F1 score of 0.871, and Accuracy of 0.810 are achieved by the model.

[0101] The RDR model is tested using all 4, 952 images in Primary test, External 2, External 3 and External 4 datasets. In FIG. 6, the results of the RDR model validation are summarized. An AUROC of 0.909, Sensitivity of 0.941, Specificity of 0.877, F1 score of 0.945, and Accuracy of 0.923 in the Primary dataset are achieved by the model. In the External 2 dataset, an AUROC of 0.839, Sensitivity of 0.868, Specificity of 0.809, F1 score of 0.917, and Accuracy of 0.862 are achieved by the model. In the External 3 dataset, an AUROC of 0.885, Sensitivity of 0.854, Specificity of 0.916, F1 score of 0.915, and Accuracy of 0.861 are achieved by the model. In the External 4 dataset, an AUROC of 0.855, Sensitivity of 0.826, Specificity of 0.884, F1 score of 0.877, and Accuracy of 0.846 are achieved by the model.

[0102] The VTDR model is tested using all 4, 952 images in Primary test, External 2, External 3 and External 4 datasets. In FIG. 7, the results of the VTDR model validation are summarized. An AUROC of 0.866, Sensitivity of 0.824, Specificity of 0.909, F1 score of 0.819, and Accuracy of 0.881 in the Primary dataset are achieved by the model. In the External 2 dataset, an AUROC of 0.823, Sensitivity of 0.845, Specificity of 0.800, F1 score of 0.867, and Accuracy of 0.830 are achieved by the model. In the External 3 dataset, an AUROC of 0.826, Sensitivity of 0.819, Specificity of 0.833, F1 score of 0.868, and Accuracy of 0.823 are achieved by the model. In the External 4 dataset, an AUROC of 0.889, Sensitivity of 0.944, Specificity of 0.833, F1 score of 0.883, and Accuracy of 0.885 are achieved by the model.

[0103] Overall, over 80%sensitivity and specificity on the primary test dataset are achieved by all of the gradability assessment, RDR, and VTDR models, indicating a good primary validation result for each model. Further validation is conducted for the RDR and VTDR models on External 2, 3, 4 datasets and for the gradability assessment model on External datasets 1, 2, 3, 4. In the External gradability assessment, only one External dataset (External dataset 1) with sensitivity below 80%(0.760) is observed. In the External RDR and VTDR model assessment, over 80%specificity and sensitivity for all External datasets are shown. Good generalization ability of the trained models is indicated by the External validation results.

[0104] Performance using different input image sizes

[0105] The performance of image classification models is significantly impacted by image size. More details and information are offered by larger images, enabling intricate features and patterns to be learned by the model. Better generalization to unseen data is aided, particularly when complex and variable datasets are dealt with. Conversely, noise and redundancy are minimized by smaller images, essential features are emphasized, and overfitting is prevented, especially when limited or noisy data is worked with. The performance of the gradability assessment, RDR, and VTDR models with different input image sizes are compared. For the model settings, the same model structures, training parameters, and augmentation methods for all input sizes are kept. Five input sizes with the same aspect ratio are used: 384×480, 448×560, 512×640, 544×680, and 576×720. The models are tested on three tasks: gradability assessment, RDR detection, and VTDR detection, respectively. For each task, five testing datasets (Primary test, and External datasets 1, 2, 3, 4) for gradability assessment and four testing datasets (Primary test, External 2, External 3, and External 4) for RDR and VTDR detection are used. The AUROC and AUPRC values for each testing dataset and each input size are calculated. The average AUROC and average AUPRC over all testing datasets as the performance metrics for each input size in FIGs. 8 and 9 are reported. The input size of 512×640 achieves the best AUROC values for all tasks. The best AUPRC values for RDR and VTDR detection are also achieved.

[0106] Performance using different augmentations

[0107] The size and diversity of the training dataset is increased, overfitting is prevented, and the issue of imbalanced data is tackled by image augmentation. The model’s robustness and accuracy can also be enhanced. Different training augmentations are experimented with to improve the performance of the present model. Horizontal flipping (HP) , Affine transformation (AT) , and Gaussian blur (GB) are used as augmentation methods. The data size is increased by horizontal flipping and affine transformation by new images with different orientations and perspectives being created. Noise are be reduced by Gaussian blur by the image pixels being smoothed. After augmentations are performed, the data are be standardized to a common scale by normalization, which allows more consistent and effective learning by the model. Two augmentation settings are compared: HP only and HP + AT + GB. Using present model setting, the same model structures, training parameters, and input image size for both settings are kept. Again, the gradability assessment, RDR, and VTDR models are tested on three tasks: gradability assessment, RDR detection, and VTDR detection, respectively. The same testing datasets as in the previous section are used. The AUROC and AUPRC values for each testing dataset and each augmentation setting are calculated. The average AUROC and AUPRC over all testing datasets as the performance metrics in FIGs. 10 and 11 are reported.

[0108] It is observed that for the average AUROC values, HP + AT + GB is outperformed by HP in all tasks and for the average AUPRC values, HP + AT + GB is outperformed by HP in gradability assessment and RDR detection. It is indicated that the model’s performance is enhanced by a combination of transformations (AT + GB) being applied to the original data (HP) , rather than HP being used alone.

[0109] Performance using different model architectures

[0110] The choice of MLP ResNet50 as the model architecture is justified, based on the comparison with other models. The standard ResNet50, as well as Inception v3, Densenet121, and EfficientNetV2_small, which have similar memory requirements as ResNet50, are considered. The present model additionally adopts an MLP layer with dropout as the output layer, instead of a fully connected layer. The gradability assessment, RDR, and VTDR models are again tested on three tasks: gradability assessment, RDR detection, and VTDR detection, respectively. Using the present model setting, all training parameters are kept constant. The input image size and the training augmentation (HP + AT + GB) for all the models are fixed, and only the model architecture is changed. Five testing datasets are used for gradability assessment and four testing datasets are used for RDR and VTDR detection. The performance of the models is measured using the average AUROC and AUPRC values for each testing dataset. The results are shown in FIGs. 12 and 13.

[0111] The average AUROC and AUPRC values show that MLP ResNet50 is performed better than all other models on all tasks, which justifies the choice of MLP ResNet50 as the model structure for the present model.

[0112] GUI software platform

[0113] The GUI-based software platform in accordance with an embodiment of the present invention is meticulously designed to be user-friendly and highly adaptable. An exemplary layout of the GUI-based software platform is shown in FIG. 18.

[0114] As illustrated in FIG. 14, when an input image is provided, the platform is first used to assess its gradability. If the image is determined to be gradable, the referable DR (RDR) prediction is then generated by the platform. For images classified as RDR-positive, the vision-threatening DR (VTDR) prediction is further provided.

[0115] For each task (gradability assessment, RDR detection and VTDR detection) , a binary prediction accompanied by a confidence score is generated by the platform. The confidence score is used to reflect the level of certainty in the predictions of the present model. High confidence scores for positive predictions are indicated to reflect strong confidence in the positive class, while similarly high scores for negative predictions are shown to signify confidence in the negative class.

[0116] A confidence score is devised to quantify the certainty of model predictions. For each model, a value between 0 and 1 is initially produced. The classification decision is then made by comparing this model output to a selected threshold. Subsequently, through probability calibration, a calibrated model output (see later) is obtained that more accurately reflects the true likelihood of class predictions. Given a calibrated model output s and the calibrated classification threshold t, it is assumed that the confidence score follows a uniform distribution from t to s.

[0117] A schematic illustration of the computation of the confidence score is shown in FIG. 15. If s > t, the confidence score is defined as:

[0118] Otherwise, if t > s, the confidence score is defined as:

[0119] Probability calibration is defined as the process of aligning predicted probabilities with the true likelihood. In an ideally calibrated model, a predicted probability of p is expected to mean that, on average, p of instances with a predicted probability p actually belongs to that particular class. For each task (gradability, RDR and VTDR) , the model output is calibrated using the validation dataset and an optimal calibration method is selected among two methods: logistic regression and isotonic regression.

[0120] FIG. 16A displays the calibration curves for the gradability primary test dataset. Two calibration methods are trained: isotonic regression and logistic regression. It is evident that the calibration curve more closely resembling the line of the perfectly calibrated model is produced by the isotonic regression model. Additionally, FIGS. 16B and 16C are used to depict the empirical probability density plots for each calibration method after calibration.

[0121] The confidence score is used to measure the level of certainty in the predictions of the present model. When a prediction is made with a high confidence score, greater trust is generally placed in it. The confidence score is calculated both before and after probability calibration. FIG. 17 displays the confidence score box plot for the gradability external datasets 1, grouped by correct predictions and false predictions.

[0122] In both cases, a higher mean confidence score than the false predictions (false positives and false negatives) are exhibited by the correct predictions (true positives and true negatives) , which is aligned with expectations. Furthermore, after calibration, the correct predictions are generally observed to have a confidence score higher than 0.5, while false predictions tend to have a confidence score lower than 0.5. This calibration is shown to enhance the meaningfulness of the confidence score.

[0123] Several embodiments of the present disclosure and features of details are briefly described above. The embodiments described in the present disclosure may be easily used as a basis for designing or modifying other processes and structures for realizing the same or similar objectives and / or obtaining the same or similar advantages introduced in the embodiments of the present disclosure. Such equivalent construction does not depart from the spirit and scope of the present disclosure, and various variations, replacements, and modifications can be made without departing from the spirit and scope of the present disclosure.

[0124] As used herein, terms "approximately" , "basically" , "substantially" , and "about" are used for describing and explaining a small variation. When being used in combination with an event or circumstance, the term may refer to a case in which the event or circumstance occurs precisely, and a case in which the event or circumstance occurs approximately. As used herein with respect to a given value or range, the term "about" generally means in the range of ±10%, ±5%, ±1%, or ±0.5%of the given value or range. The range may be indicated herein as from one endpoint to another endpoint or between two endpoints. Unless otherwise specified, all the ranges disclosed in the present disclosure include endpoints. The term "substantially coplanar" may refer to two surfaces within a few micrometers (μm) positioned along the same plane, for example, within 10 μm, within 5 μm, within 1 μm, or within 0.5 μm located along the same plane. When reference is made to "substantially" the same numerical value or characteristic, the term may refer to a value within ±10%, ±5%, ±1%, or ±0.5%of the average of the values.

Claims

A system for automated classification and visual interpretation of ultra-widefield retinal images of a subject, comprising:an ultra-widefield scanning laser ophthalmoscopy imaging device configured to capture a retinal image of at least 150 degrees field of view of the subject’s eye; anda computing device comprising at least one processor and one non-transitory memory;wherein the non-transitory memory stores machine instructions prompting a process for executing commands comprising:pre-processing a retinal image of the subject’s eye captured by the ultra-widefield scanning laser ophthalmoscopy imaging device to readjust resolution and intensity values and applying data augmentations;determining whether the retinal image is gradable for diagnostic purposes based on image quality features;categorizing the retinal image into one of four diabetic retinopathy severity categories by detection and analysis of retinal lesion regions by one or more classification neural networks;generating a visual activation map highlighting pixels relating to the retinal lesion regions; andoutputting a diagnostic result comprising a categorization label and confidence score to a graphical user interface.The system of claim 1, wherein the four diabetic retinopathy severity categories are:no diabetic retinopathy (no DR) ;non-referable diabetic retinopathy (non-DR) ;referable diabetic retinopathy (RDR) ; andvision-threatening diabetic retinopathy (VTDR) .The system of claim 1, wherein the retinal image of the subject is captured without requiring pharmacological pupil dilation administered to the subject.The system of claim 1, wherein the ultra-widefield scanning laser ophthalmoscopy imaging device captures the retinal image of the subject’s eye at a resolution of at least 2600 × 2048 pixels.The system of claim 1, wherein the pre-processing of the retinal image comprises resizing the retinal image to a resolution in a range of 500 × 500 pixels to 700 × 700 pixels.The system of claim 1, wherein the data augmentations comprise one or more of random horizontal flipping, Gaussian blurring and affine transformation.The system of claim 1, wherein the determining of gradability of the retinal image is performed by a convolutional neural network trained to identify blurring, occlusion, dark artifacts and image coverage of a standardized retinal grid.The system of claim 1, wherein the categorizing the retinal image into one of four diabetic retinopathy severity categories is performed by a cascaded machine learning architecture comprising:a first neural network configured to detect VTDR; anda second neural network configured to detect RDR from non-DR images.The system of claim 8, wherein each of the first neural network and the second neural network comprises a convoluted neural network backbone pre-trained on a general-purpose image dataset and fine-tuned on retinal image data, and a multilayer perceptron having a rectified linear unit-activated hidden layer, a dropout layer, and a final output layer generating class probabilities.The system of claim 1, wherein the graphical user interface is configured to display:the original retinal image;the visual activation map highlighting pixels relating to the retinal lesion regions as an overlay; anddiagnostic result comprising a categorization label and confidence score; andat least one export option for exporting and integrating the original retinal image, the visual activation map and the diagnostic result to an external clinical information system.The system of claim 1, wherein the confidence score is based on the softmax probability output from a final convolutional layer of the one or more classification neural networks and is displayed as a numeric value between 0 and 1.The system of claim 1, wherein the generating a visual activation map comprises applying a gradient-weighted class activation mapping (Grad-CAM) technique to a final convolutional layer of the one or more classification neural networks.The system of claim 1, wherein the non-transitory memory further stores machine instructions for commanding an export of diagnostic results to an external electronic medical record system.The system of claim 1, wherein the retinal lesions comprise hemorrhages, microaneurysms, non-perfusion and neovascularization.The system of claim 1, wherein the computing device is further configured to perform real-time inference within 5 seconds of image acquisition.A method for automated classification and visual interpretation of an ultra-widefield retinal image, comprising:capturing a retinal image of a subject using an ultra-widefield scanning laser ophthalmoscopy (UWF-SLO) imaging device;pre-processing the retinal image by readjusting resolution and intensity values of the image and applying data augmentations;determining whether the retinal image is gradable for diagnostic purposes based on image quality features;categorizing the retinal image into one of four diabetic retinopathy severity categories by detection and analysis of retinal lesion regions by one or more classification neural networks;generating a visual activation map highlighting pixels relating to the retinal lesion region; andproviding a diagnostic result comprising the original retinal image, the visual activation map, a categorization label and a confidence score to a graphical user interface.The method of claim 16, wherein the four diabetic retinopathy severity categories are:no diabetic retinopathy (no DR) ;non-referable diabetic retinopathy (non-DR) ;referable diabetic retinopathy (RDR) ; andvision-threatening diabetic retinopathy (VTDR) .The method of claim 16, wherein retinal image of the subject is captured without requiring pharmacological pupil dilation administered to the subject.The method of claim 16, wherein the ultra-widefield scanning laser ophthalmoscopy imaging device captures the retinal image of the subject’s eye at a resolution of at least 2600 × 2048 pixels.The method of claim 16, wherein the pre-processing of the retinal image comprises resizing the retinal image to a resolution in a range of 500 × 500 pixels to 700 × 700 pixels.The method of claim 16, wherein the data augmentations comprise one or more of random horizontal flipping, Gaussian blurring and affine transformation.The method of claim 16, wherein the determining of gradability of the retinal image is performed by a convolutional neural network trained to identify blurring, occlusion, dark artifacts and image coverage of a standardized retinal grid.The method of claim 16, wherein the categorizing the retinal image into one of four diabetic retinopathy severity categories is performed by a cascaded machine learning architecture comprising:a first neural network configured to detect VTDR; anda second neural network configured to detect RDR from non-DR images.The method of claim 23, wherein each of the first neural network and the second neural network comprises a convoluted neural network backbone pre-trained on a general-purpose image dataset and fine-tuned on retinal image data, and a multilayer perceptron having a rectified linear unit-activated hidden layer, a dropout layer, and a final output layer generating class probabilities.The method of claim 16, wherein the generating a visual activation map comprises applying a gradient-weighted class activation mapping (Grad-CAM) technique to a final convolutional layer of the one or more classification neural networks.The method of claim 16, wherein the confidence score is based on the softmax probability output from a final convolutional layer of the one or more classification neural networks and is displayed as a numeric value between 0 and 1.The method of claim 16, wherein the retinal lesions comprise hemorrhages, microaneurysms, non-perfusion and neovascularization.The method of claim 16, wherein the computing device is further configured to perform real-time inference within 5 seconds of image acquisition.

Citation Information

Patent Citations

  • Diabetic retina eye ground image pathology detection method

    CN108470359A

  • Diabetic retinopathy automatic evaluation system based on focus perception modeling

    CN112967227A

  • Detection of Pathologies in Ocular Images

    US20200196853A1

  • Predicting a pathological condition from a medical image

    US20200320692A1