Systems and methods for deep learning to detect retinal vasculitis on color fundus photographs of patients with uveitis
Patent Information
- Application Number
- EP2024710954
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-01
- Filing Date
- 2024-01-31
- Publication Date
- 2025-12-10
AI Technical Summary
Current methods for diagnosing retinal vasculitis in patients with uveitis, such as fluorescein angiography, are invasive, time-consuming, and costly, and may not be feasible in resource-limited settings, and manual evaluation is challenging due to subtle signs of inflammation.
A deep learning system using a convolutional neural network (CNN) classifier and ensemble learning model is trained to detect retinal vasculitis from ultrawide-field color fundus photographs, which can identify retinal vascular leakage by analyzing original, cropped, and region-of-interest images, providing a non-invasive and efficient diagnostic tool.
The system demonstrates high accuracy in differentiating between images with and without retinal vasculitis, offering a robust and efficient alternative to traditional invasive methods, with improved sensitivity and specificity, facilitating non-invasive monitoring of posterior segment uveitis.
Smart Images

Figure US2024013833_08082024_PF_FP
Abstract
Description
SYSTEMS AND METHODS FOR DEEP LEARNING TO DETECT RETINAL VASCULITIS ON COLOR FUNDUS PHOTOGRAPHS OF PATIENTS WITH UVEITISFIELD
[0001] The present disclosure generally relates to machine learning algorithms for detecting conditions of the eye based on photographic evaluation, and particularly to systems and methods for deep learning to detect retinal vasculitis exhibited in color fundus photographs of patients with uveitis.BACKGROUND
[0002] Machine learning is a system of artificial intelligence (Al) that enables computers to learn and detect predictive features from medical images / information without specified programming or rules. As data availability has increased, machine learning has shown promising results in interpretation of medical imaging / information and diagnosis of diseases. In the field of ophthalmology, research studies have demonstrated the machine learning’s capability in diagnosing several common ocular diseases including diabetic retinopathy, age-related macular degeneration, retinal vein occlusion, and glaucoma. The efforts led to the Food and Drug Administration’s approval of the IDx-DR system for diabetic retinopathy diagnosis.
[0003] Uveitis is a heterogenous group of inflammatory intraocular diseases accounting for up to 15% of cases of blindness in Western countries. Uveitis can be classified into anterior, intermediate, posterior, or panuveitis based on anatomical involvement. Retinal vasculitis refers to the inflammation of the retinal vessels and is commonly associated with posterior segment uveitic diseases. Fundus examination of patients with retinal vasculitis may also reveal sheathing and cuffing of the blood vessels, vascular occlusion, telangiectasis, microaneurysms, and ischemia-induced neovascularization; however, the signs of retinal vasculitis are usually not apparent on fundus examination. Therefore, much of the diagnosis, monitoring and management are based on fluorescein angiography, and signs of retinal vasculitis on fluorescein angiography including retinal vascular leakage and staining can help to assess the severity of the disease. However, fluorescein angiography is an invasive procedure with risk of adverse effects (although low).Additionally, the imaging procedure is time-consuming and may impose economic burdens on the patients as well as the physicians, especially in resource limited settings.
[0004] It is with these observations in mind, among others, that various aspects of the present disclosure were conceived and developed.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] FIG. 1 is a simplified block diagram showing a system for a trained neural network classifier to evaluate ultrawide-field color fundus photographs and identify those ultrawide-field color fundus photographs that exhibit retinal vascular leakage (RVL).
[0006] FIG. 2 illustrates a method for cropping ultrawide-field color fundus photographs from original ultrawide-field color fundus photographs according to one aspect of the system.
[0007] FIG. 3A shows ultrawide-field color fundus photographs without RVL and FIG. 3B shows ultrawide-field color fundus photographs exhibiting retinal vascular leakage RVL.
[0008] FIG. 4A shows cropped ultrawide-field color fundus photographs without RVL and FIG. 4B shows cropped ultrawide-field color fundus photographs that exhibit RVL.
[0009] FIG. 5 shows region of interest (ROI) ultrawide-field color fundus photographs in which the ROI has been highlighted for each image size of the eye.
[0010] FIG. 6 is an illustration showing the architecture of a convolutional neural network (CNN) classifier trained to identify ultrawide-field color fundus photographs that exhibit RVL.
[0011] FIG. 7 is a simplified illustration showing the architecture of an ensemble learning (EL) model based on three CNN classifiers used to evaluate whether various kinds of ultrawide-field color fundus photographs exhibit RVL.
[0012] FIG. 8 is a simplified diagram showing an exemplary computing system for implementation of the system of FIG. 6 or FIG. 7.
[0013] FIG. 9 is a process flow diagram illustrating the method for preparing cropped ultrawide-field color fundus photographs from original ultrawide- field color fundus photographs.
[0014] FIG. 10 is a simplified block diagram showing an example neural network architecture model used in the classification operation of FIG. 6 for the implementation of the system of FIG. 1.
[0015] Corresponding reference characters indicate corresponding elements among the view of the drawings. The headings used in the figures do not limit the scope of the claims.DETAILED DESCRIPTION
[0016] The present disclosure discloses systems and methods for a machine learning algorithm 290 of a deep neural network 120 trained to detect retinal vasculitis exhibited in ultrawide-field color fundus photographs 110 of patients with uveitis. In one aspect, the ultrawide-field color fundus photographs used to train and later evaluated by the deep neural network 120 may be original ultrawide-field fundus photographs 110, cropped ultrawide-field color fundus photographs 112 based on an estimation of the original ultrawide-field color fundus photographs 112, or region of interest (ROI) ultrawide-field color fundus photographs 114 that highlight the area of the eye around the optic disc. In a further aspect, the deep neural network 120 may be a Convolutional Neural Network (CNN) classifier 120 that is operable for feature extraction of the evaluated ultrawide-field color fundus photographs 110 / 112 / 114 and then classification of the extracted features of the CNN classifier 120 to determine whether the evaluated ultrawide-field color fundus photograph 110 / 112 / 114 exhibits RVL. In another aspect, an ensemble learning model is disclosed that combines the output of a plurality of CNN classifiers 120 for determining through a hard voting method (e.g., majority voting) whether RVL is exhibited in any evaluated ultrawide-field color fundus photographs 110 / 112 / 114. In addition, a study is disclosed herein that was conducted in compliance with the Declaration of Helsinki and approved by the Institutional Review Board of the National Institutes of Health demonstrating the efficacy of the trained deep neural network model 120 in identifying ultrawide-field color fundus photographs 110 / 112 / 114 that exhibit retinal vasculitis leakage.Trained Neural Network
[0017] FIG. 1 illustrates a system 100 having a deep neural network 120 trained to evaluate a plurality of different types of inputted ultrawide-field colorfundus photographs 110 / 112 / 114 and detect any of those ultrawide-field color fundus photographs 110 / 112 / 114 that display retinal vasculitis leakage in the image. In particular, the trained deep neural network 120 has a retinal vasculitis detection algorithm 290 trained to detect those inputted ultrawide-field color fundus photographs 110 / 112 / 114 that exhibit retinal vasculitis leakage as shall be discussed in greater detail below.Ultra-widefield Fluorescein Angiography Acquisition and Assessment
[0018] Prior to training the deep neural network 120, the ultrawide-field color fundus photographs 110 / 112 / 114 were captured. The original ultrawide-field color fundus photographs 110 were captured with the Optos 200Tx (Optos PLC, Dunfermline, UK) ultra-widefield retinal imaging system (not shown). After pupil dilation, 5 ml of 10% sodium fluorescein was administered via the antecubital vein. The eye of patients with more severe symptoms as determined by the uveitis specialist was chosen as the transit eye. Standard fluorescein angiographic techniques and image acquisition consisted of early-phase (15-45 seconds) images of the eye of primary interest and images of the fellow eye at 45-60 seconds, if available. In addition to those standard images, late-phase images were obtained at 5 to 10 minutes for both eyes.
[0019] The original ultrawide-field color fundus photographs 110 were graded for retinal vascular leakage (RVL), which was classified as absent or present by an expert. After the completion of the grading, the corresponding ultrawide-field color fundus photographs 110 of each group were used for the training the retinal vasculitis detection algorithm 290 of a Convolutional Neural Network (CNN) classifier 120 to train the CNN classifier 120 to identify those ultrawide-field color fundus photographs 110 that exhibit RVL.Original Ultrawide-Field Color Fundus Photograph Dataset
[0020] The photograph dataset consisted of 438 original retinal ultrawide-field color fundus photographs 110. Among them, 172 original ultrawide- field color fundus photographs 110 exhibited retinal vascular leakage (RVL) and 266 of the original ultrawide-field color fundus photographs 110 did not exhibit RVL. The original ultrawide-field color fundus photographs 110 were classified according to the presence or absence of RVL by evaluation of the corresponding fluoresceinangiography images. In one aspect, the size of the original ultrawide-field color fundus photographs 110 used in the dataset was 4,000x4,000x3.Cropped Ultrawide-Field Color Fundus Photo Estimation from Original Ultrawide- Field Color Fundus Photographs
[0021] As noted above, three different types of retinal ultrawide-field color fundus photographs 110 / 112 / 114 were used as input for the CNN classifier 120. The first type of input were original ultrawide-field color fundus photographs 110 of the retina as described above, the second type of input were cropped ultrawide- field color fundus photographs 112 estimated from the original ultrawide-field color fundus photographs 110 of the retina, and the third type of input used were region of interest (ROI) ultrawide-field color fundus photographs 114 in which the a particular area around the optic disc of the original ultrawide-field color fundus photographs 110 was highlighted as a region of interest in the image. Cropped and ROI ultrawide- field color fundus photographs 112 / 114 were found to provide a better image of the retina for detecting RVL in in the eye when evaluated by the retinal vasculitis detection algorithm 290, although the original ultrawide-field color fundus photographs 110 without any cropping were found to be sufficient as input to the CNN classifier 120 for training the retinal vasculitis detection algorithm 290.
[0022] One method 200 for estimating cropped ultrawide-field color fundus photographs 112 from the original ultrawide-field color fundus photographs 110 is shown in FIG. 2 and the process flow of method 200 is illustrated in FIG. 9. Referring specifically to FIG. 9, at block 400 the method 200 collects a green channel from each of a plurality of original ultrawide-field color fundus photographs 110. At block 402, the method 200 applies a CLAHE operator (with tile grid size = 0.9 x width of the inputted original ultrawide-field color fundus photograph 110) to the green channel of each inputted original ultrawide-field color fundus photograph 110 (see Image enhancement). At block 404, the method 200 applies gaussian blurring to the output image 117 of CLAHE operation using 5x5 operator to produce a blurred image. At block 406, the method 200 binarizes the blurred image using Otsu operator to produce a binarized image 119. At block 408, the method 200 applies morphological erosion (with 3x3 window) fifty times and morphological dilation (with 3x3 window) fifty times to improve the binarized image 119. At block 410, the method 200 inverts the binarized image 119 of the morphological operations andselects the largest blob in the inverted binarized image 121 (see Image Inversion). At block 412, the method 200 estimates a rectangle 116 (top left and bottom right of ultrawide-field color fundus photograph) that covers the largest blob having white intensity in the inverted binarized image 121 (see Crop retina area from the ultrawide-field color fundus photograph). At block 414, the method 200 crops the retina area from the original fundus photo 110 using the rectangle 116 to produce the cropped ultrawide-field color fundus photograph 112 (see cropped ultrawide-field color fundus photograph 112 in FIG. 2). FIGS. 3A and 3B show original ultrawide- field color fundus photographs 110, while FIGS. 4A and 4B show corresponding cropped ultrawide-field color fundus photographs in FIGS. 3A and 3B.Region of Interest Photo Estimation from Ultrawide-Field Color Fundus Photographs
[0023] Macular and optic disc (optic nerve head) are key functional areas related to vision and retinal diseases. Among these, the optic disc is the area of the eye where blood vessels and optic nerve fibers enter and exit the eye. Since uveitis is related to leakage in the vessels, the region of interest (ROI) area near the optic disc was cropped in the ROI ultrawide-field color fundus photograph 114 to test if the trained CNN classifier 120 can classify RVL from the optic disc area when evaluating the image. The process estimated the area of the optic disc from the original ultrawide-field color fundus photographs 110 and then generated five different sizes of ultrawide-field color fundus photographs 110 that include the optic disc in the middle of the image. FIG. 5 shows the five different sizes of ROI ultrawide-field color fundus photographs 114 highlighted from the original ultrawide- field color fundus photographs 110. In particular, the first row of FIG. 5 is the sizes of ROI ultrawide-field fundus photographs 114:1536 x 1536, 1024 x 1024, 512 x 512, 256 x 256, and 128 x 128. The second row of FIG. 5 shows region of interest boxes of ROI ultrawide-field color fundus photographs 114 in the original ultrawide-field color fundus photographs 100 and the third row of FIG. 5 shows ROI ultrawide-field color fundus photographs 114 cropped from original ultrawide-field color fundus photographs 100.Convolutional Neural Network (CNN) Classifier
[0024] FIG. 6 shows the architecture of the deep neural network 120, such as a CNN classifier 120, having a feature extraction block 122 operable forextracting features from the ultrawide-field color fundus photographs 110 / 112 / 114 inputted into the CNN classifier 120 and a classification block 124 operable for classifying the ultrawide-field color fundus photographs 110 / 112 / 114 that exhibit RVL from the extracted features 308.
[0025] In one embodiment, six different CNN classifiers 120A-120F, such as VGG16 120A (https: / / arxiv.org / abs / 1409.1556), VGG19 (https: / / arxiv.org / abs / 1409.1556) 120B, ResNet50 120C (https: / / arxiv.org / abs / 1512.03385), ResNet152 120D (https: / / arxiv.org / abs / 1512.03385), DenseNet121 120E (https: / / arxiv.org / abs / 1608.06993), and Densenet169 120F were used as feature extractors for each of the CNN classifiers 120A-120F, although any number of CNN classifiers 120 may be used.
[0026] As noted above, to perform the feature extraction and classification operations, each CNN classifier 120 includes a feature extraction block 122 for performing feature extraction from each of the inputted ultrawide-field color fundus photographs 110 / 112 / 114 and a classification block 124 for performing the classification of the extracted features to determine whether the extracted features exhibit RVL. In one embodiment, the classification block 124 may include a global average pooling layer 126 that receives the extracted features 308 (FIG. 10) and is in communication with a first fully connected layer 140 having a first dense layer and a first dropout layer, which communicates with a second fully connected layer 142 having a second dense layer and a second dropout layer. The second fully connected layer 142 communicates with an output layer 128. The output layer 128 is operable for assigning a binary class 1 if retinal vascular leakage is identified in the extracted features 308 by the classification block 124 and a binary class 0 if retinal vascular leakage is not detected in the extracted features 308 by the classification block 124. In one aspect, the extracted features 308 may include one or more physical features of the eye image captured in the color fundus photograph 110 / 112 / 114 that the feature extraction block 122 extracts and communicates to the classification block 124 for classification to determine whether the one or more extracted features 308 classify a particular color fundus photograph 110 / 112 / 114 as exhibiting RVL. In one embodiment, the size of the first and second dense layers was 1 x1 ,024, and dropout ratio = 0.25 was used for the first and second dropout layers to resolve overfitting issues. Pre-trained weights of the ImageNet were usedas initial weights for feature extraction and train all layers including the layers in the feature extraction block 122 for performing feature extraction. In one embodiment, all the inputted ultrawide-field color fundus photographs 110 / 112 / 114 were resized to 512x512x3 prior to inputting into the CNN classifier 120.
[0027] Python with Tensorflow Keras was used to implement the CNN classifier 120. To train the CNN classifier 120 to identify inputted ultrawide-field color fundus photographs 110 / 112 / 114 that exhibit RVL, a stochastic gradient descent (SGD) was used with learning rate = 0.001 , decay=1e-6, momentum = 0.9, nesterov momentum = True, and epochs = 100. The hardware configuration used for this experiment was 2 x intel Xeon Gold 5218 processors 2.3 GHz, 64 hyper-thread processors, 8 x RTX 2080 Ti, and Red Hat Enterprise Linux 7.
[0028] During the feature extraction operation, the feature extraction block 122 of the CNN classifier 120 was trained to extract features from each evaluated ultrawide-field color fundus photograph 110 / 112 / 114 and then sending those extracted features to the classification block 124 of the CNN classifier 120 to perform the classification of the extracted features so that the output layer 128 may conclude whether the extracted features from a particular ultrawide-field color fundus photograph 110 / 112 / 114 exhibits RVL by assigning a class 0 for non-leakage (no RVL being exhibited) or a class 1 for leakage (RVL is being exhibited).Ensemble Learning (EL) Model based on Three CNN Classifiers
[0029] The present system 100 further includes an Ensemble Learning Model 118 operable to combine the results of each respective output layer 128 from a plurality of CNN classifiers 120 to generate more stable and robust output for identifying evaluated ultrawide-field color fundus photographs 110 / 112 / 114 that exhibit RVL. FIG. 7 shows an example architecture of the Ensemble Learning model 118 based on three different CNN classifiers 120, among the CNN classifiers 120A- 120F, that are arranged in parallel operation. After training multiple CNN classifiers 120A-120F to identify those ultrawide-field color fundus photographs 110 / 112 / 114 that exhibit RVL, three of the best CNN classifiers 120 among the CNN classifiers 120A-120F were selected for the Ensemble Learning Model 118. In this arrangement, each CNN classifier 120 conducted its own feature extraction 122 and then generated its own classification 124 result at a respective output layer 128 (e.g., output layers 128A-128C to provide either an output (Class 0 or Class 1 ) and allthese respective output results of each of the CNN classifiers 120 were aggregated based on a hard voting method to make the final decision on whether a particular ultrawide-field color fundus photograph 110 / 112 / 114 truly exhibited RVL in the combined aggregated output 130 of the CNN classifiers 120. The hard voting (majority voting) of the combined output 130 involves the act of tallying up the predictions (Class 0 - no RVL or Class 1 - RVL exhibited) made by each of the respective CNN classifiers 120 such that the Ensemble Learning Model 118 selects the classification output that receives the most votes as the final prediction on whether the RVL is being exhibited by the evaluated ultrawide-field color fundus photographs 110 / 112 / 114.Experimental Results
[0030] A five-fold cross validation method was used to evaluate the ability of CNN classifier 120 to identify ultrawide-field color fundus photographs 110 / 112 / 114 that exhibit RVL. For example, five training datasets and five test datasets were generated from a dataset of ultrawide-field color fundus photographs 110 / 112 / 114, and the CNN classifier 120 was trained and tested five times to evaluate its performance. In each dataset, 80% of the ultrawide-field color fundus photographs 110 / 112 / 114 were used for training the CNN classifier 120 to identify ultrawide-field color fundus photographs 110 / 112 / 114 that exhibit RVL, while 20% of the ultrawide-fundus color fundus photographs 110 / 112 / 114 were used for testing the ability of the CNN classifier 120 to identify those ultrawide-field color fundus photographs 110112 / 114 being evaluated that exhibited RVL. The training ultrawide- field color fundus photographs 110 / 112 / 114 were augmented to six-fold by flipping each of the ultrawide-field color fundus photographs 110 / 112 / 114 along the vertical axis and then rotating the ultrawide-field color fundus photographs ± 10 degree to increase the size of the training set.
[0031] Table I showed the estimation results of the CNN classifier 120. The performances of the six different backbone feature extractors 122 of the CNN classifier 120 were as shown in columns three to eight. The third and fourth rows of Table I show the results using original ultrawide-field color fundus photographs 110 and the cropped ultrawide-field color fundus photographs 112, respectively. In the case of original ultrawide-field color fundus photographs 110, the CNN classifier 120 using VGG19 showed the best performance with 0.7235 accuracy, 0.8308sensitivity, and 0.5581 specificity. In the case of cropped ultrawide-field color fundus photographs 112, the CNN classifier 120 using DenseNet169 showed the best performance with 0.7898 accuracy, 0.8532 sensitivity, and 0.6908 specificity.Overall, DenseNetl 69-based CNN classifier 120 using cropped ultrawide-field color fundus photographs 112 showed the best performance.Table I: Performance of the Proposed Binary CNN Classifiers
[0032] Table II shows estimation results of the Ensemble Learning Model 118. The performances of the two different types of Ensemble Learning Models 118 were as shown in columns two and three. Cropped ultrawide-field color fundus photographs 112 were used as inputs. ResNet152, DenseNetl 21 , and DenseNetl 69 classifiers 120 were used for the first Ensemble Learning Model 118 (second column) and VGG19, DenseNetl 21 , and DenseNetl 69 classifiers 120 were used for the second Ensemble Learning Model 118 (third column). Among them, the second Ensemble Learning Model 118 showed the best performance with 0.8012 accuracy, 0.8759 sensitivity, and 0.7135 specificity.Table II: Performance of the Ensemble Learning Model.
[0033] Table III showed estimation results of the CNN classifiers 120 using ROI ultrawide-field color fundus photographs 114 as input. DenseNet169 was used as backbone feature extractor 122 since it showed the best performance in the cropped ultrawide-field color fundus photographs 112. Five different sizes of ROIs were used for this experiment. Among them the CNN classifiers 120 using 1536 x 1536 ROI size showed the best performance with 0.7761 accuracy, 0.8382 sensitivity, and 0.7484 specificity. In addition, the CNN classifier 120 using 128 x 128 ROI size showed a meaningful performance with 0.7535 accuracy, 0.8011 sensitivity, and 0.6785 specificity. Since 128 x 128 ROI includes optic disc in most area as shown in FIG. 5, it is assumed that the CNN classifier 120 found key evidence of RVL from optic disc area of the eye in the inputted ultrawide-field color fundus photographs 110 / 112 / 114.Table III: Performance of the Binary CNN Classifiers for different ROI size inputs.System
[0034] As shown in FIG. 1 , the present disclosure provides a system 100 that uses one or more CNN classifiers 120 trained to detect retinal vasculitis leakage in ultrawide-field color fundus photographs 110A / 112A / 114A based solely on the evaluation of the ultrawide-field color fundus photographs 110 / 112 / 114 inputted into and evaluated by each of the trained CNN classifiers 120. As noted above, the CNN classifier 120 is trained using training sets of ultrawide-field color fundus photographs 110 / 112 / 114 that train a retinal vasculitis detection algorithm 290 (FIG. 10) of the CNN classifier 120 to detect retinal vasculitis identified in ultrawide-field color fundus photographs 110 / 112 / 114 in patients with uveitis
[0035] FIG. 10 is a schematic block diagram of an example deep neural network 120 that may be used with one or more embodiments described herein, e.g., as a component of system 100 shown in FIG. 1, and particularly as a component of the classification block 124 of CNN classifier 120. Possible implementations of theclassification block 124 can be used by the system 100 to identify those ultrawide- field color fundus photographs 110 / 112 / 114 that exhibit retinal vasculitis based on the evaluation of the features extracted by the feature extraction block 122 of the CNN classifier 120.
[0036] In particular, the classification block 124 of CNN classifier 120 is disclosed by an example neural network description 301 in an engine model (neural controller) 330. The neural network description 301 can include a full specification of the classification block 124 and a part of the specification of the CNN classifier 120. For example, the neural network description 301 can include a description or specification of the CNN classifier 120 (e.g., the layers, layer interconnections, number of nodes in each layer, etc.); an input and output description which indicates how the input and output are formed or processed; an indication of the activation functions in the neural network, the operations or filters in the CNN classifier 120, etc.; neural network parameters such as weights, biases, etc.; and so forth.
[0037] The classification block 124 of CNN classifier 120 includes a global average pooling layer 126 that functions as an input layer, which receives the extracted features 308 extracted by the feature extraction block 122 from one or more ultrawide-field color fundus photographs 110 / 112 / 114.
[0038] The CNN classifier 120 includes hidden layers 304A through 304 V (collectively “304” hereinafter). The hidden layers 304 can include n number of hidden layers, where n is an integer greater than or equal to one. The number of hidden layers can include as many nodes as needed for a desired processing outcome and / or rendering intent. The CNN classifier 120 further includes an output layer 128 that provides an output (e.g., the set of depth estimation data) resulting from the processing performed by the hidden layers 304. In an illustrative example, the output layer 128 can output those ultrawide-field color fundus photographs 110A / 112A / 114A that exhibit RVL.
[0039] The CNN classifier 120 in this example includes a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated with the nodes is shared among the different layers and each layer retains information as information is processed. In some cases, the CNN classifier 120 can include a feed-forward neural network, in which case there are no feedback connections where outputs of the CNN classifier 120 are fed back into itself. In other cases, the CNN classifier 120 can include a recurrentneural network, which can have loops that allow information to be carried across nodes while reading in input.
[0040] Information can be exchanged between nodes through node-to- node interconnections between the various layers. Nodes of the input layer 123 can activate a set of nodes in the first hidden layer 304A. For example, as shown, each of the input nodes of the input layer 123 is connected to each of the nodes of the first hidden layer 304A. The nodes of the hidden layer 304A can transform the information of each input node by applying activation functions to the information. The information derived from the transformation can then be passed to and can activate the nodes of the next hidden layer (e.g., 304B), which can perform their own designated functions. The output of the hidden layer (e.g., 304B) can then activate nodes of the next hidden layer (e.g., 304W), and so on. The output of the last hidden layer can activate one or more nodes of the output layer 128, at which point an output is provided.
[0041] In some cases, each node or interconnection between nodes can have a weight that is a set of parameters derived from training the CNN classifier 120. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a numeric weight that can be tuned (e.g., based on a training dataset), allowing the CNN classifier 120 to be adaptive to inputs and able to learn as more data is processed.
[0042] The CNN classifier 120 can be pre-trained to process the features from the data in the input layer 123 using the different hidden layers 304 to provide the output through the output layer 128. The CNN classifier 120 can learn to identify RVL in ultrawide-field color fundus photographs 110 / 112 / 114 and can be trained using training data that includes sets of ultrawide-field color fundus photographs 110A / 112A / 114A that exhibit RVL. For instance, training data is inputted into the CNN classifier 120, which can be processed by the CNN classifier 120 to generate outputs which can be used to tune one or more aspects of the CNN classifier 120, such as weights, biases, etc. for identifying ultrawide-field color fundus photographs 110A / 112A / 114A that exhibit RVL.
[0043] In some cases, the CNN classifier 120 can adjust weights of nodes using a training process called backpropagation. Backpropagation can include a forward pass, a loss function, a backward pass, and a weight update. The forwardpass, loss function, backward pass, and parameter update is performed for one training iteration. The process can be repeated for a certain number of iterations for each set of training data of ultrawide-field color fundus photographs 110A / 112A / 114A exhibiting RVL until the weights of the layers are accurately tuned.
[0044] For a first training iteration for the CNN classifier 120, the output can include values that do not give preference to any particular class due to the weights being randomly selected at initialization. For example, if the output is a vectorwith probabilities that the ultrawide-field color fundus photograph 110 / 112 / 114 exhibits RVL, the probability value for each of the different ultrawide-field color fundus photographs 110 / 112 / 114 may be equal or at least very similar. With the initial weights, the CNN classifier 120 is unable to determine whether RVL is being exhibited in a particular ultrawide-field color fundus photograph 110 / 112 / 114 and thus cannot make an accurate determination. A loss function can be used to analyze errors in the output. Any suitable loss function definition can be used.
[0045] The loss (or error) can be high for the first training dataset process (iteration) since the actual values will be different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output comports with a target or ideal output that accurately identifies an ultrawide-field color fundus photographs 110A / 112A / 114A that exhibit RVL. The CNN classifier 120 can perform a backward pass by determining which inputs (weights) most contributed to the loss of the CNN classifier 120 and can adjust the weights so that the loss decreases and is eventually minimized.
[0046] A derivative of the loss with respect to the weights can be computed to determine the weights that contributed most to the loss of the CNN classifier 120. After the derivative is computed, a weight update can be performed by updating the weights of the filters. For example, the weights can be updated so that they change in the opposite direction of the gradient. A learning rate can be set to any suitable value, with a high learning rate including larger weight updates and a lower value indicating smaller weight updates.
[0047] The CNN classifier 120, for example the feature extraction block 122 and the classification block 124, can include any suitable neural or deep learning network. One example includes a convolutional neural network (CNN), which includes an input layer and an output layer, with multiple hidden layers between the input and out layers. The CNN include a series of convolutional, nonlinear, pooling(for downsampling), and fully connected layers. In other examples, the CNN classifier 120 can represent any other neural or deep learning network, such as an autoencoder, a deep belief nets (DBNs), and recurrent neural networks (RNNs), generative adversarial networks (GANs), and capsule network (CapsNets) etc.
[0048] The CNN classifier 120 can include any suitable machine learning algorithms in the classification block 124 such as Support Vector Machines (SVM), Multi-layer Perception (MLP), Random Forest Classifier (RFC), XGBoost (XGB), and Light GBM (LGBM).Computer-implemented System
[0049] FIG. 8 is a schematic block diagram of an example computer- related apparatus 102 that may be used with one or more embodiments described herein, e.g., as a component of system 100 having one or more trained CNN classifiers 120 that identifies ultrawide-field fundus photographs 110 / 112 / 114 exhibiting retinal vascular leakage based on an evaluation of a plurality of ultrawide- field fundus photographs 110 / 112 / 114 inputted into one or more trained CNN classifiers 120.
[0050] In one embodiment, the computer-related apparatus 102 comprises one or more network interfaces 210 (e.g., wired, wireless, PLC, etc.), at least one processor 220, and a memory 240 interconnected by a system bus 250, as well as a power supply 260 (e.g., battery, plug-in, etc.). The computer-related apparatus 102 can further include display device 230 in communication with the processor 220 that displays information to a user.
[0051] Network interface(s) 210 include the mechanical, electrical, and signaling circuitry for communicating data over the communication links coupled to a communication network. Network interfaces 210 are configured to transmit and / or receive data using a variety of different communication protocols. As illustrated, the box representing network interfaces 210 is shown for simplicity, and it is appreciated that such interfaces may represent different types of network connections such as wireless and wired (physical) connections. Network interfaces 210 are shown separately from power supply 260, however it is appreciated that the interfaces that support PLC protocols may communicate through power supply 260 and / or may be an integral component coupled to power supply 260.
[0052] Memory 240 includes a plurality of storage locations that are addressable by processor 220 and network interfaces 210 for storing software programs and data structures associated with the embodiments described herein. In some embodiments, computer-related apparatus 102 may have limited memory or no memory (e.g., no memory for storage other than for programs / processes operating on the device and associated caches). Memory 240 can include instructions executable by the processor 220 that, when executed by the processor 220, cause the processor 220 to implement aspects of the system 100 outlined herein.
[0053] Processor 220 comprises hardware elements or logic adapted to execute the software programs (e.g., instructions) and manipulate data structures 245. An operating system 242, portions of which are typically resident in memory 240 and executed by the processor 220, functionally organizes device 102 by, inter alia, invoking operations in support of software processes and / or services executing on the device 102. In some embodiments, these software processes and / or services may include retinal vasculitis detection algorithm 290 that implement aspects of the system 100 described herein wherein the CNN classifier 120 has been trained to identify ultrawide-field fundus photographs 110 / 112 / 114 that display retinal vascular leakage among a plurality of ultrawide-field color fundus photographs 110 / 112 / 114 evaluated by the trained neural network 120. Note that while retinal vasculitis detection algorithm 290 is illustrated in centralized memory 240, alternative embodiments provide for the process to be operated within the network interfaces 210, such as a component of a MAC layer, and / or as part of a distributed computing network environment.
[0054] It will be apparent to those skilled in the art that other processor and memory types, including various computer-readable media, may be used to store and execute program instructions pertaining to the techniques described herein. Also, while the description illustrates various processes, it is expressly contemplated that various processes may be embodied as modules or engines configured to operate in accordance with the techniques herein (e.g., according to the functionality of a similar process). In this context, the term module and engine may be interchangeable. In general, the term module or engine refers to model or an organization of interrelated software components / functions. Further, while the retinal vasculitis detection algorithm 290 is shown as a standalone process, thoseskilled in the art will appreciate that this algorithm may be executed as a routine or module within other processes.Discussion
[0055] In summary, the preliminary results demonstrate that the retinal vasculitis detection algorithm 290 employed in the CNN classifier 120 of FIG. 6 and the Ensemble Learning Model 118 of FIG. 7 can differentiate between patients with and without retinal vasculitis based on evaluation of ultrawide-field color fundus photographs 110 / 112 / 114 of those patients which would be extremely difficult for a doctor to perform manually. The present disclosure contemplates an Ensemble Learning Model 118 having a plurality of CNN classifiers 120 to further improve the classification accuracy for detecting RVL in ultrawide-field color fundus photographs 110 / 112 / 114 as well as perform analysis using heatmap visualization to better understand the areas that the retinal vasculitis detection algorithm 290 is detecting. Finally, it further contemplated to classify RVL into a three-step classification and investigate whether the retinal vasculitis detection algorithm 290 can predict the extent and severity of RVL exhibited in color fundus photographs and red-free color fundus photographs 110. These findings have implications for the potential of non- invasively monitoring posterior segment uveitis.
[0056] It should be understood from the foregoing that, while particular embodiments have been illustrated and described, various modifications can be made thereto without departing from the spirit and scope of the invention as will be apparent to those skilled in the art. Such changes and modifications are within the scope and teachings of this invention as defined in the claims appended hereto.
Claims
CLAIMSWhat is claimed is:
1. A system comprising: a processor in communication with memory, the memory including instructions executable by the processor to: receive, at a CNN classifier formulated at the processor, a plurality of color fundus photographs, the CNN classifier having been trained using a machine learning algorithm to identify color fundus photographs that exhibit retinal vascular leakage; and determine, at the CNN classifier, whether any of the plurality of color fundus photographs exhibits retinal vascular leakage.
2. The system of claim 1 , the memory further including instructions executable by the processor to: train the CNN classifier using original color fundus photographs, cropped color fundus photographs, or color fundus photographs having a region of interest highlighted to detect retinal vascular leakage in any of the plurality of color fundus photographs.
3. The system of claim 1 , wherein the CNN classifier comprises: a feature extraction block operable for extracting one or more features of the plurality of color fundus photographs; and a classification block operable for evaluating the extracted features for determining whether retinal vascular leakage is exhibited in one or more of the plurality of color fundus photographs.
4. The system of claim 3, wherein the classification block of the CNN classifier comprises an output layer that assigns a binary class 1 if retinal vascular leakage is detected in the extracted features by the classification block and abinary class 0 if retinal vascular leakage is not detected in the extracted features by the classification block.
5. The system of claim 1 , wherein the classification block comprises: a global average pooling layer; a first fully connected layer in communication with the global average pooling layer; a second fully connected layer in communication with the first fully connected layer; and an output layer that assigns a binary class 1 if retinal vascular leakage is detected in the extracted features by the classification block and a binary class 0 if retinal vascular leakage is not detected in the extracted features by the classification block.
6. The system of claim 5, wherein the first fully connected layer includes a first dense layer and a first dropout layer, which communicates with a second fully connected layer having a second dense layer and a second dropout layer.
7. The system of claim 5, wherein the global average pooling layer comprises the extracted features received from the feature extraction block.
8. The system of claim 3, wherein the extracted features comprise physical features of the eye image captured in a respective one of the plurality of color fundus photographs.
9. A system comprising: a processor in communication with memory, the memory including instructions executable by the processor to: receive, at a plurality of CNN classifiers formulated at the processor, a plurality of color fundus photographs, each of the CNN classifiers having been trained using a respective machine learning algorithm to identify color fundus photographs that exhibit retinal vascular leakage;extract, at a feature extraction block of each of the plurality of CNN classifiers, features extracted from each of the plurality of color fundus photographs; and classify, at a classification block of each of the plurality of CNN classifiers, whether the extracted features from each of the plurality of color fundus photographs exhibit retinal vascular leakage.
10. The system of claim 9, further comprising: aggregating, at a plurality of CNN classifiers formulated at the processor, the classification of each of the plurality of CNN classifiers to determine whether the aggregated classification exhibits retinal vascular leakage.11 . The system of claim 10, wherein the aggregated classification is determined to exhibit retinal vascular leakage or determined not to exhibit retinal vascular leakage using a hard voting method.
12. The system of claim 9, wherein the classification block includes an output layer operable for assigning a binary class 1 if retinal vascular leakage is detected in the extracted features by the classification block and assigning a binary class 0 if retinal vascular leakage is not detected in the extracted features by the classification block.
13. The system of claim 9, wherein the classification block comprises: a global average pooling layer; a first fully connected layer in communication with the global average pooling layer; a second fully connected layer in communication with the first fully connected layer; and an output layer operable for assigning a binary class 1 if retinal vascular leakage is detected in the extracted features by the classification block and a binary class 0 if retinal vascularleakage is not detected in the extracted features by the classification block.
14. The system of claim 9, wherein output layer assigns a binary class 1 if retinal vascular leakage is detected in the extracted features by the classification block and assigns a binary class 0 if retinal vascular leakage is not detected in the extracted features by the classification block.
15. The system of claim 13, wherein the first fully connected layer includes a first dense layer and a first dropout layer, which communicates with a second fully connected layer having a second dense layer and a second dropout layer.