Retinal scan image classification
The computer-implemented retinal scan image classification method uses an integrated network to classify retinal scans with multiple CNNs, solving the problem of inefficient IRD diagnosis in the prior art and achieving more efficient and accurate diagnosis.
Patent Information
- Application Number
- CN202380074123.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-09-26
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively diagnose genetic retinal disease (IRD), especially in the absence of resources and expertise, which leads to inefficient genetic diagnosis.
The computer-implemented retinal scanning image classification method is used, and the retinal scanning is classified using an integrated network (ensemble network) combined with multiple convolutional neural networks (CNNs) to generate the probability of the existence of each ocular lesion.
The process of genetic diagnosis is significantly accelerated through integrated networks, improving the diagnostic accuracy and efficiency of rare eye diseases such as IRD, especially when resources are limited.
Smart Images

Figure CN120077415A_ABST
Abstract
Description
Field of the Invention
[0001] The present invention relates to a method for classifying retinal scan images, a neural network for classifying retinal scan images, and a method for training a network for classifying retinal scan images. Furthermore, the present invention relates to a data processing device, a computer program, and a computer-readable medium for classifying retinal scan images. Specifically, the present invention relates to a method and a network for estimating the probability of the presence of multiple eye diseases within a retinal scan. Background of the Invention
[0003] Inherited retinal diseases (IRDs) are a group of rare genetic disorders that affect 1 in 3,000 people, and over 300 different IRD-related genes have been identified to date. IRDs cause degeneration of the retina, the light-sensitive tissue at the back of the eye responsible for vision. Some patients with IRDs may have severely impaired vision from birth, while others may find their peripheral and / or central vision deteriorating over time. Collectively, IRDs are a leading cause of childhood blindness and the most common cause of blindness in the working-age population, with significant psychological and socioeconomic implications.
[0004] Revealing the genetic cause of IRD is a prerequisite for optimally determining prognosis, providing genetic counseling, eligibility for gene-directed therapy, and inclusion in gene-directed clinical trials. However, in over 40% of cases, such genetic diagnosis remains elusive, mainly due to lack of access to services or inefficient diagnostic services, such as insufficient evidence, lack of clinical experience, and inability to successfully link or communicate data and research findings between research groups. In many parts of the world, almost 100% of cases remain genetically unresolved due to lack of access to and resources for funding molecular testing.
[0005] IRDs can have unique phenotypic features that clinicians have learned to identify using modern retinal imaging techniques, which rapidly and non-invasively acquire images of the retina through a dilated pupil with minimal inconvenience to the patient. These scans can be performed using various imaging modalities, such as fundus autofluorescence (FAF), infrared (IR) imaging, and spectral domain optical coherence tomography (SD-OCT), each of which conveys different information about the retinal structure.
[0006] Figure 1 Example retinal scans (also referred to as eye or ophthalmic scans or images) acquired using these three example imaging modalities are provided. Panel A is an IR fundus scan acquired at 30 degrees. Panel B is a FAF scan acquired at 55 degrees. Panel C is a single B-scan from an SD-OCT volume.
[0007] For example, FAF images can help identify patterns of photoreceptor dysfunction and apoptosis by specifically detecting lipofuscin and related compounds based on their autofluorescence properties. Lipofuscin and related compounds mainly accumulate in the retinal pigment epithelium (RPE) promoted by oxidative stress. Infrared imaging helps identify vascular abnormalities and also helps identify melanin (a natural pigment that protects the eyes) and melanolipofuscin (a granule containing melanin and lipofuscin). SD-OCT helps identify and characterize multiple cell layers of the retina, such as the RPE, photoreceptors, outer nuclear and plexiform, inner nuclear and plexiform, ganglion cell, and nerve fiber layers. The main cells affected by IRD are photoreceptors, and SD-OCT can depict the highly reflective ellipsoid zone (EZ), which is used to identify the inner and outer segment layers of photoreceptors and is usually disrupted early in IRD.
[0008] This high-resolution, in-depth multimodal information enables ophthalmologists to identify some gene-specific patterns of the disease, thus allowing prediction of disease-related genes in certain cases. However, given the rarity of these diseases, the expertise and experience required to make an accurate clinical diagnosis are not widely available and are limited to a few experts and hospitals that have developed such expertise over the decades.
[0009] This limitation is particularly evident when considering the diagnosis of IRD; an increasing number of people are now targeted for clinical trials, and approved treatments are now available. However, to obtain treatment, a genetic diagnosis needs to be established early enough. It is crucial that timely identification of the genetic cause remains challenging.
[0010] The inventors have recognized that timely identification of eye lesions, especially those with a genetic determinant, is desirable. Summary of the Invention
[0012] According to one aspect of the present invention, there is provided a computer-implemented method for classifying retinal scan images. The method includes the step of receiving an input data set including at least one retinal scan obtained using one of a plurality of imaging modalities. That is, after obtaining the retinal scan (or possibly multiple retinal scans), the classification method involves accepting the retinal scan for processing. This can be done, for example, directly via the acquisition machine or by a user uploading the retinal scan for processing.
[0013] The method further includes passing an input data set through an ensemble network that includes a plurality of convolutional neural networks (CNNs). Each CNN herein is configured to classify a retinal scan obtained using one of a plurality of imaging modalities based on a plurality of eye diseases. That is, the first CNN is configured to classify a retinal scan obtained using the first imaging modality based on an eye disease (e.g., a disease with a genetic basis), and the second CNN is configured to classify a retinal scan obtained using the second imaging modality based on an eye disease, and so on. Each CNN produces a probability of the presence of each eye disease within the retinal scan. The (multiple) retinal scans (which are passed into the ensemble network) are processed using the CNN corresponding to the imaging modality used to obtain the (multiple) retinal scans. For example, if the (multiple) retinal scans are obtained using the first imaging modality, the (multiple) retinal scans are processed using the CNN corresponding to the first imaging modality. The CNNs are trained. The CNN (or group thereof) for a single imaging modality can be referred to as a predictor block or a modality-specific ensemble model.
[0014] The CNN can be configured to output probabilities for a predefined number of classes, the number of classes corresponding to the number of diseases being studied. For example, if the ensemble network is to classify retinal scans for 10 classes, 10 probabilities will be output from the ensemble network.
[0015] The ensembled probabilities are not intended to replace genetic testing, as phenotyping can never fully replace molecular testing, especially when treatments such as gene therapy are to be based on a genetic diagnosis. Instead, the retinal scan image classification technique helps to acquire diagnostic expertise (currently only available at a limited number of expert centers globally), thus significantly accelerating the process of genetic diagnosis. Using CNNs to classify eye diseases is a promising approach to accelerating medical diagnosis, such as the genetic diagnosis of patients with IRD, especially considering the increasing number of treatable IRDs, and rapid diagnosis can lead to improved patient treatment outcomes.
[0016] The method further includes generating an integrated (or aggregated) probability of the presence of each eye disease within the retinal scan. In the case where only a single retinal scan is passed through the integration network and there is only a single CNN for the imaging modality used to acquire that single retinal scan, the integrated probability is the probability generated from the CNN. In the case where only a single retinal scan is passed through the integration network and there are multiple CNNs within the modality-specific predictor block for the imaging modality used to acquire that single retinal scan, the integrated probability can be the average of the probabilities generated from the multiple CNNs. The integrated probability can be output in the form of a list of probabilities, one probability corresponding to each classification category.
[0017] Optionally, the input data set can include multiple retinal scans, each retinal scan acquired using one of multiple imaging modalities. In this case, each retinal scan can be processed using the CNN (or multiple CNNs within the modality-specific predictor block) corresponding to the imaging modality used to acquire the retinal scan. For example, the input data set can include at least one retinal scan from each imaging modality. The input data set can include at least one retinal scan from each of two imaging modalities. That is, in the case where a first retinal scan acquired using a first imaging modality and a second retinal scan acquired using a second imaging modality are passed through the integration network, the integrated probability is some function (e.g., average) of the probabilities generated from the first CNN (for the first imaging modality) and the second CNN (for the second imaging modality). Multiple retinal scans using the same imaging modality increase the confidence in the eye disease classification of the CNN for that imaging modality. Multiple retinal scans using different imaging modalities similarly increase the confidence in the eye disease classification.
[0018] Optionally, the multiple imaging modalities can include any one or all of the following: fundus autofluorescence (FAF); infrared (IR); and spectral domain optical coherence tomography (SD-OCT). These are the imaging modalities commonly used to acquire retinal scans. It has been found that an integration network built on CNNs corresponding to these imaging modalities is successful in identifying eye diseases. For example, due to the comprehensive genetic testing framework for IRDs and the transformative improvements in imaging technology embedded in specialized healthcare services over the past decade, there are now a sufficient number of molecularly characterized patients with detailed retinal phenotyping (using, for example, the above imaging modalities to construct a representative data set for deep learning). Incidentally, FAF scans and IR scans can be acquired with different angular extents, such as 30 degrees and 55 degrees; these can be considered the same or different imaging modalities.
[0019] Optionally, multiple eye pathologies are genetic abnormalities (or mutations). In this way, the integrated network is well-suited to help detect rare eye diseases, such as IRD, which are otherwise difficult to diagnose genetically. IRD is a typical monogenic disease and a leading cause of blindness in children and working-age adults worldwide. Then, wider access to expertise, which was previously limited to human experts in the field, can be provided by an AI system trained to detect gene-specific patterns from retinal scans. Even when limited to a single imaging modality, the integrated network performs at least as well as human experts in classifying genetic abnormalities in this task and often significantly better.
[0020] Optionally, the method may further include passing the integrated probability through a trained linear classifier. This supplements or refines the integrated probability output from the CNN based on further information encoded in the classifier. For example, the linear classifier can be configured to refine the integrated probability based on the subject age and / or mode of inheritance (MOI) of the genetic abnormality. The use of the linear classifier has been shown to improve the classification accuracy. Other factors, such as the subject's ethnicity, can be additionally or alternatively encoded within the trained linear classifier.
[0021] Optionally, the integrated network can include multiple (k) CNNs for each imaging modality. The retinal scan is processed using k CNNs, each corresponding to the imaging modality used to acquire the retinal scan. That is, multiple trained CNNs for a single imaging modality can be applied simultaneously to a single retinal scan. Then, the method can include the step of averaging the probabilities of the presence of each eye pathology within the retinal scan from each of the k CNNs. The averaged probabilities (there can be multiple averaged probabilities, one set for each group of CNNs) can then be used to generate the integrated probability. Using multiple CNNs for each imaging modality has been shown to improve the classification accuracy relative to using a single CNN.
[0022] Optionally, the method may further include: when the estimation of the presence of each eye disease in the retinal scan is below a predefined confidence threshold, outputting a prompt to re-acquire the retinal scan. For example, in cases where the training dataset for training the CNN is sparsely populated (e.g., having retinal scans for rare diseases), when there is uncertainty, the trained CNN may tend to over-predict the most common classes and under-predict the rarer classes; thus introducing a threshold introduces further certainty for retinal scan image classification. In addition, lower-quality retinal scans can also confound the results, especially if age-related. For example, retinal scans of young children tend to be noisier, partly due to poorer compliance or movement during image acquisition; the use of the confidence threshold quickly flags such scans to the user.
[0023] Optionally, the method may further include processing the output dataset to provide a classification of the genetic causes of eye diseases. For example, in cases where the eye disease is a genetically-based disease, the method speeds up the diagnostic time as it can point to specific genes to be studied. The method will help the vast majority of ophthalmologists and other vision healthcare workers who are not proficient in rare IRDs to indicate when it is worthwhile to consider molecular testing. Genetic prediction can also be used to score and thus prioritize genetic variants that match the gene and phenotype. In diseases with very different phenotypes, the "match phenotype" PP4 criterion in the ACMG guidelines provides a higher level of evidence in variant classification, which may be crucial for the transition from variants of unknown significance to likely pathogenic variants.
[0024] According to one aspect of the present invention, there is provided an integrated network for classifying retinal scan images, particularly for use in the above-mentioned retinal scan image classification method, the integrated network including a plurality of CNNs. Each CNN is configured to classify a retinal scan acquired using one of a plurality of imaging modalities based on a plurality of eye diseases, and each CNN generates a probability of the presence of each eye disease in the retinal scan.
[0025] According to one aspect of the present invention, there is provided a computer-implemented method for training an integrated network for classifying retinal scan images, particularly the above-mentioned integrated network. The training method includes receiving an input training dataset, the input training dataset including training retinal scans acquired using an acquisition imaging modality and marked with known eye diseases, each training retinal scan being acquired using one of a plurality of imaging modalities. In one example, the training dataset can be a combination of retinal scans and a lookup table (e.g., a CSV document) that relates to each of the retinal scans and various file metadata (e.g., patient.id, laterality, date, file.name, scan.id, scan.number, gene, modality, etc.).
[0026] The training method includes dividing the training retinal scans into multiple sets of training retinal scans, where each set of training retinal scans includes training retinal scans acquired using the same imaging modality. The division can further produce a held-out validation set; this set is not used for actual training but rather for quantifying the efficacy of the trained CNN.
[0027] The training method includes, for each set of training retinal scans, training at least one CNN to minimize the error between the predicted classification of multiple eye diseases and the known eye diseases. The training can start from scratch (i.e., for example, starting from a CNN architecture with random weights and biases), or it can involve fine-tuning the existing weights of the CNN (e.g., starting from ImageNet weights). Fine-tuning avoids excessive computational costs.
[0028] The training method includes, for each CNN, outputting the trained CNN weights. That is, after training (e.g., when a predefined number of epochs have passed, or when the training loss meets a predefined criterion), the CNN is "frozen" and the checkpoint or weights are saved. The output weights can be stored locally or transmitted elsewhere so that the retinal scan image classification method can be implemented locally or elsewhere without having to perform the training each time.
[0029] Optionally, the training method can include training the CNN using k-fold cross-validation, such that the method includes, for each set of training retinal scans, training k CNNs. It has been shown that training multiple networks for each imaging modality results in improved classification accuracy compared to training only a single network for each imaging modality. The value of k in this case can be determined based on the training dataset at hand; suitable values include 2, 5, 10, 20, or n, where n is the size of the dataset for each test sample (leave-one-out cross-validation). Of course, the number of CNNs used for one imaging modality does not need to be the same as the number used for another imaging modality.
[0030] Optionally, the training method can further include augmenting the input training dataset. For example, the training dataset can be augmented in any or all of the following ways: rotation; flipping; scaling; cropping; adjusting brightness; blurring; and adding noise. Of course, those skilled in the art will understand that other augmentation techniques are available and suitable. By augmenting the training dataset, the size and diversity of the dataset can be increased, thereby offsetting certain biases in the training data. In this way, the training method can avoid overfitting.
[0031] Optionally, the training method may further include preprocessing the input training data set. For example, the training data set can be preprocessed in any or all of the following ways: filtering by median pixel intensity; filtering by noise level; filtering by identification of image artifacts; and filtering by BRISQUE score. Of course, those skilled in the art will understand that other preprocessing techniques are available and suitable.
[0032] Another aspect of the embodiments includes a data processing apparatus including means adapted to perform the method of classifying retinal scan images according to an embodiment or the method of training an integrated network for classifying retinal scan images according to an embodiment. The preprocessing removes low-quality and defective training images, thus resulting in a more accurate trained CNN.
[0033] Another aspect of the embodiments includes a computer program including instructions that, when executed by a computer, cause the computer to perform the method of classifying retinal scan images according to an embodiment or the method of training an integrated network for classifying retinal scan images according to an embodiment. The computer program can be stored on a computer-readable medium. The computer-readable medium can be non-transitory.
[0034] Accordingly, another aspect of the embodiments includes a non-transitory computer-readable medium including instructions that, when executed by a computer, cause the computer to perform the method of classifying retinal scan images according to an embodiment or the method of training an integrated network for classifying retinal scan images according to an embodiment.
[0035] The present invention can be implemented in digital electronic circuits, or in computer hardware, firmware, software, or in combinations thereof. The present invention can be implemented as a computer program or a computer program product, i.e., a computer program tangibly embodied in a non-transitory information carrier (e.g., in a machine-readable storage device) or a propagated signal for execution or control of the operation of one or more hardware modules by one or more hardware modules.
[0036] The computer program can be in the form of a stand-alone program, a part of a computer program, or more than one computer program, and can be written in any form of programming language (including compiled language or interpreted language), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a data processing environment. The computer program can be deployed to execute on one or more modules at one site or distributed across multiple sites and interconnected via a communication network.
[0037] The method steps of the present invention can be performed by one or more programmable processors executing a computer program to perform the functions of the present invention by operating on input data and generating output. The apparatus of the present invention can be implemented as programmed hardware or dedicated logic circuitry, including, for example, FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit).
[0038] Processors suitable for executing a computer program include, for example, any one or more processors of a general purpose microprocessor and a dedicated microprocessor as well as any type of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions, which is coupled to one or more memory devices for storing the instructions and data.
[0039] The present invention is described in terms of specific embodiments. Other embodiments are within the scope of the appended claims. For example, the steps of the present invention can be performed in a different order and still achieve the desired result. Multiple test script versions can be edited and invoked as a unit without using object-oriented programming techniques; for example, the elements of a script object can be organized in a structured database or file system, and the operations described as being performed by the script object can be performed by a test control program.
[0040] The elements of the present invention have been described using the terms "processor", "input device", etc. Those skilled in the art will understand that such functional terms and their equivalents can refer to parts of a system that are spatially separated but combined to serve the defined function. Similarly, the same physical part of a system can provide two or more of the defined functions. For example, the same memory and / or processor can be appropriately used to implement separately defined devices. Brief Description of the Drawings
[0042] By way of example only, reference is made to the drawings, in which:
[0043] Figure 1 is a composite image of a retinal scan obtained using three imaging modalities commonly used to examine the retinas of patients with hereditary retinal diseases;
[0044] Figure 2 is a flowchart providing a method for classifying retinal scan images according to an embodiment;
[0045] Figure 3 is a schematic diagram showing the contribution of each of fifteen CNNs in an example image classification method using three imaging modalities;
[0046] Figure 4 is a schematic diagram showing an example of 5-fold cross-validation for training an ensemble network;
[0047] Figure 5 is a graph of the training loss within 100 training epochs during the example training process;
[0048] Figure 6 is an overview of the preprocessing filtering process for filtering the example training dataset;
[0049] Figure 7 is an overview of the data augmentation for augmenting the example training dataset;
[0050] Figure 8 is a set of receiver operating characteristic (ROC) curves obtained from cross - validation data for six out of 36 predicted genes in the example for three different imaging modalities;
[0051] Figure 9 is a schematic diagram showing four different methods for classifying retinal scan images considered during performance evaluation;
[0052] Figure 10 is a set of ROC curves obtained from cross - validation data for combined predictions using an ensemble network, covering predictions for six out of 36 predicted genes in the example for three different imaging modalities;
[0053] Figure 11 is a calibration plot for a single retinal scan processed using an ensemble network;
[0054] Figure 12 is a schematic diagram showing the distribution of the age at the earliest presentation of the subjects for gene diagnosis in the example training dataset;
[0055] Figure 13 is a schematic diagram showing the known genetic patterns of each gene in the example training dataset;
[0056] Figure 14 is a set of ROC curves obtained from cross - validation data for linear classifiers trained on different feature combinations for six out of 36 predicted genes in the example;
[0057] Figure 15 is a representative screen of the graphical user interface, which shows six retinal scans as inputs to the ensemble network and the ensemble probabilities of the resulting gene abnormalities; and
[0058] Figure 16 is a schematic diagram of the suitable hardware for implementing the embodiments of the present invention.
[0059] Detailed Description
[0060] Figure 2It is a flowchart depicting a computer-implemented method for classifying retinal scan images.
[0061] At S10, the computer receives an input data set including at least one retinal scan obtained using one of a plurality of imaging modalities.
[0062] At S20, the computer passes the input data set through an ensemble network that includes a plurality of CNNs. Each CNN is configured to classify a retinal scan obtained using one of a plurality of imaging modalities based on a plurality of eye diseases, and each CNN produces a probability of the presence of each eye disease within the retinal scan. The retinal scan is processed using the CNN corresponding to the imaging modality used to obtain the retinal scan.
[0063] At S30, the computer generates an ensemble probability of the presence of each eye disease within the retinal scan.
[0064] The ensemble network used in the retinal scan image classification method includes at least one neural network for each imaging modality used to obtain the retinal scans intended to be processed by the ensemble network. For example, if two different imaging modalities are used, such as FAF and SD-OCT, there is at least one neural network intended to process FAF scans and at least one neural network intended to process SD-OCT scans. Of course, multiple neural networks can be used for each imaging modality.
[0065] As an example, consider an ensemble network including fifteen constituent neural networks. In this example, each neural network is an Inception-v3 deep CNN (CNN) that can obtain one or more retinal scans of three different imaging modalities from a given patient. Each constituent neural network in this example outputs gene-level prediction scores for 36 individual IRD genes. In this example, these 36 genes are selected because together they cover more than 80% of IRD cases in the European population.
[0066] Given a single input retinal scan of one of the three supported modalities, the example ensemble network can apply each of five networks corresponding to the modality of the scan to obtain a single integrated image-level prediction. For a set of retinal scans from a single patient, the appropriate set of CNNs can be applied to each scan in turn. Then, the resulting network-wise predictions and image-wise predictions can be combined to produce a single integrated prediction (or list of probabilities) for the patient by taking the average of the individual (post-softmax) predictions.
[0067] Figure 3It is an explanatory diagram of the contributions of each of the fifteen constituent CNNs in the above example. Predictor blocks specific to each imaging modality (FAF, IR, and SD-OCT) consist of 5 associated CNNs, which are configured to provide IRD gene predictions (i.e., predictions of gene abnormalities or mutations) given a retinal scan of the associated imaging modality. Example retinal scans are inserted within each predictor block. The CNNs in this example each give predictions for 36 individual genes. Given a set of retinal scans from a patient, the appropriate predictor block corresponding to the modality of each scan can be applied.
[0068] The predictions of all the CNNs within the predictor block specific to each imaging modality are averaged. For example, using the shown FAF predictor block, the 5 individual estimated probabilities (thin bars) for gene abnormalities for BEST1, ABCA4, CNGB3, and PRPH2 are averaged to provide the average estimated probability (thick, lower bar) for gene abnormalities for each item. For each imaging modality, only the top 4 genes (in terms of average estimated probability) are depicted. That is, for a retinal scan of a specific modality, the 5-model predictor block can be applied by averaging the output probabilities for each gene.
[0069] The average estimated probabilities over all scans in the set (from all imaging modalities) can be taken and used as the final prediction or ensemble probability for the set of retinal scans. As shown in the rightmost panel (outside the predictor block), the ensemble probability is obtained by combining the predictions of its 15 constituent networks using an ensemble approach.
[0070] Of course, if a retinal scan obtained using only a single imaging modality is processed, the ensemble probability will correspond to the average estimated probability of the single predictor block (for the same single imaging modality).
[0071] In the above example, the number of constituent CNNs (15) is presented only as an example. A reader skilled in the art will understand that other numbers can also be used. 15 was chosen in the example to enable 5-fold cross-validation (see below, regarding training the ensemble network); in k-fold cross-validation, the value of k can be chosen such that each training and test group of data samples is large enough to statistically represent a broader dataset. In this case, k = 5 was found to be appropriate.
[0072] Training
[0073] Any known training technique suitable for training an image classification model can be applied to the constituent CNNs of the embodiments, provided that each resulting trained neural network is capable of classifying retinal scan images.
[0074] In this case, all the training code was written in Python using the Keras library with the TensorFlow backend. Using the Inception-v3 preprocessing capabilities, images were loaded into the model training routine via the built-in Keras data loader. This loaded the images as RGB images, resized them to the correct input dimensions (256×256 pixels), and rescaled the pixel values to be between -1 and 1 (by dividing by 127.5 and subtracting 1). Although the input scans in this example training process were monochromatic, for convenience, the inventors loaded them with three color channels (all set to the same value).
[0075] To load the images during the training phase, the images were also automatically augmented by applying random transformations. [The transformations used]
[0076] For the network architecture, the inventors used the built-in Inception-v3 model in Keras, used the saved ImageNet weights, and stripped the final output layer, and included a drop-out layer (eventually, the drop-out layer was not used in the final experiments in this example; see below) and a final output layer with 36 outputs and softmax normalization.
[0077] The network was trained in a standard manner, using the Adam optimizer on the weighted cross-entropy loss of the predicted gene classes against the true genes. This was done using the Keras fit function, with logging and monitoring via TensorBoard. At the end of training (after 100 epochs), the network weights were saved, and the inventors also saved the network configuration file with other network settings, for recording the save and for easy future loading of the network via a wrapper class. The classes were re-weighted inversely proportional to the frequency with which images of each class appeared in the training dataset.
[0078] The training in this example was done using an Nvidia Quadro P6000 GPU implemented on a Dell desktop. Other GPUs (such as the Nvidia TITAN Xp) are also suitable. Additionally, a regular CPU is also capable of performing the training; a CPU is typically well-suited for smaller training tasks.
[0079] As described above, k-fold cross-validation is a suitable framework for training the constituent CNNs of an ensemble network. The training data (i.e., the data used to train the CNNs) includes at least retinal scans and an indication of the lesions present in the retina of each scan (e.g., an indication of genetic anomalies).
[0080] Figure 4It is an illustrative diagram of k-fold cross-validation for training an ensemble network of 5 CNNs for three imaging modalities (i.e., 5-fold cross-validation, resulting in 15 CNNs).
[0081] In this example, the above training process is repeated for each fold of each modality, and the trained network weights are saved at the end of each run (thus not using early stopping), giving a total of 15 sets of network weights.
[0082] The patients in the development set are divided into 5 folds of approximately equal size as follows:
[0083] 1. Prepare a count table for each fold-modality-gene combination and initialize it to zero.
[0084] 2. For each patient, determine the genes associated with the patient and the number of images for each modality.
[0085] 3. Randomly shuffle the patients using the set seed (to ensure reproducibility).
[0086] 4. For each patient, sequentially select the fold with the lowest count on all modalities of the patient's genes in the table and assign the patient to that fold. For each modality, the counts in the table are incremented by the number of images of that patient.
[0087] To facilitate the training and evaluation of our individual networks, the inventors divided the (CSV format) dataset into a series of training CSVs and test CSVs based on separate folds (so for each modality, there are files [...]_train_0.csv, [...]_test_0.csv, [...]_train_1.csv, [...]_test_1.csv, and so on), where [...]_train_N.csv includes folds 0 to 4, excluding fold N, and [...]_test_N.csv includes fold N and fold -1.
[0088] To ensure a reasonable amount of training data for each gene, the inventors excluded any gene that did not have at least 5 images in all folds / modalities and excluded all images / patients related to those genes from the training and test sets. This gave a list of 36 genes used in all experiments.
[0089] In the illustrated example, training is performed using retinal scans obtained from patients with some form of IRD from Moorfields Eye Hospital (MEH). These retinal scans were obtained from patients with IRD attending MEH who had undergone genetic testing and in whom a confirmed genetic cause had been identified by an accredited diagnostic laboratory.
[0090] The MEH IRD cohort was previously described by Pontikos, N. et al. (“Genetic basis of inherited retinal disease in a molecularly characterised cohort of over 3000 families from the United Kingdom”. Ophthalmology (2020).) and contains 4236 individuals with IRD caused by variants in 135 different genes, of which 452 individuals (with variants in 66 genes) were less than 18 years of age as of 2019 - 08 - 02. Patients with IRD and a genetic diagnosis confirmed by an accredited genetic diagnostic laboratory were identified, and information on genetic diagnosis, age at presentation, and inheritance pattern was exported from the MEH electronic health record (OpenEyes) using SQL queries on the Microsoft SQL Server hospital data warehouse database.
[0091] For records between 5 June 2006 and 5 April 2018, images of these patients were exported from the MEH Heidelberg Imaging (Heyex) database (Heidelberg Engineering, Heidelberg, Germany) according to the hospital numbers of all patients with IRD. This resulted in a dataset of 1,196,038 images from 2871 IRD patients who had pathogenic variants in 132 different genes.
[0092] Table 1 shows the number of patients and images for each of the 36 genes and modalities (FAF, IR, and SD - OCT) included in the example development dataset.
[0093]
[0094]
[0095] Note that the complete dataset in this example is subject to quality control (see the section below). In this way, the dataset can be limited (after quality control) to genes that have at least 5 images for each gene in each of the three modalities, leaving 36 individual genes. The distribution of the 36 selected genes is presented in Table 1 above. For all 132 genes, some genes did not have a sufficient number of patients with sufficient images to correctly assign at least 5 images to each fold and were thus excluded.
[0096] Note that the complete dataset obtained from MEH in this example includes (following data quality control; see below) 15,692 FAF scans, 23,631 IR scans, and 13,099 SD-OCT scans from 2,171 patients from 36 different genes. This complete dataset was split into a "development" set of 1,907 patients and a held-out internal test set of 264 patients. The internal test set allows testing of the final integrated network. The development set allows investigation of network-level properties and integration methods.
[0097] As Figure 4 shown, the training in this example was performed using a total of 44,817 retinal scans obtained from 1,907 patients (3,749 eyes, 6,397 appointments) with some form of IRD from MEH. These scans were divided into three different modalities: FAF (N = 13,509), IR (N = 20,098), and SD-OCT (N = 11,344). For each of the three modalities, five different neural networks were trained using 5-fold cross-validation on the development data, resulting in a total of fifteen different neural networks.
[0098] The patients in the development set were divided into 5 patient subsets of approximately equal size, taking care to ensure a roughly even distribution among each different gene in each of the three different modalities. For each of the three modalities, the images from each of the 5 patient subsets were used as the folds in a 5-fold cross-validation setting. On each combination of four patient subsets, a deep CNN was trained, so that each patient subset was left out of the training set of exactly one network (such that the images from that subset could be used as test data for that network). In addition to diversifying the constituent CNNs of the ensemble, this also allows one to view the variation in the network-level model (ensemble network) due to different training / test sets, which can provide insight into how sensitive the ensemble network is to the training dataset.
[0099] For each of the 15 datasets (three modalities, 5 folds each), the CNN for 36 classes was trained for 100 epochs (going through the entire dataset), which was found to be sufficient. Figure 5Shows the example network loss on the training set during training. As shown, in a preliminary survey of all images from this example dataset, it was found that 100 epochs were sufficient for training to converge for all hyperparameter settings.
[0100] To obtain the values of the various hyperparameters used to train each CNN, for each of the three modalities, the first dataset can be split into a further training dataset and a validation dataset according to a 75:25 split (giving a valid overall 60:20:20 training / validation / test split). The network can then be trained on these new training sets using various random hyperparameter settings; the training accuracy and average per-gene AUC score (defined below) on the validation sets for 10 random hyperparameter settings are recorded in Table 2 below.
[0101] Table 2 Hyperparameter (learning rate; batch size; dropout) settings tested, and the highest results for each column highlighted. Overall, the hyperparameters used for Run 6 produced the best results for the example dataset across the three imaging modalities.
[0102]
[0103] As seen above, dropout with various dropout probabilities can be used in the penultimate layer of the network. No dropout was found to work well across all three modalities, so a dropout probability of 0% was used throughout this embodiment, meaning that dropout effectively does not exist in the final CNN.
[0104] For cross-validation, in theory, hyperparameter tuning should ideally be applied independently to each fold to avoid biasing the final test results. However, in practice, this significantly increases the computational requirements, and the benefits may be negligible. Therefore, in this example, the best hyperparameter settings were also used for training the remaining folds. This effect only applies to the results on the in-house development set here and does not affect any results on the in-house or external test sets.
[0105] For each image, a corresponding class label is given by the genetic diagnosis of the underlying patient. For the neural network, the Inception-v3 architecture initialized with pre-trained ImageNet weights is used, where the final output layer is replaced by a linear layer with 36 outputs followed by a softmax normalization. The Inception-v3 architecture was chosen because it is the smallest in model size among the choices of architectures with similar performance applied in other ophthalmic contexts. Readers skilled in the art will understand that any other architecture suitable for image classification would also be appropriate. The loss function chosen in this example is the cross-entropy loss, and its class weights are inversely proportional to the gene frequencies in the dataset. For the Adam optimizer, the default parameters used in the Keras library are chosen (β 1 = 0.9, β 2 = 0.999).
[0106] In summary, for this example, hyperparameter tuning is performed by random sampling in 10 trials on an 80 / 20 train / validation split on the first training fold. After hyperparameter tuning, a batch size of 128, a learning rate of 0.0001, and a dropout probability of 0% are found to be optimal. These hyperparameters are used to train the network on all three modalities. To avoid overfitting on the training data, data augmentation techniques can also be applied (see the section below).
[0107] Returning to Figure 4 , 5-fold cross-validation is used to train the CNN for each imaging modality used during dataset acquisition.
[0108] Training Data Quality Control and Enhancement
[0109] Before training, the dataset at hand can be preprocessed to provide quality control. Note that this step is not always necessary and depends on the dataset in question. In the above dataset (Heyex database from MEH), quality control was implemented. Figure 6 A schematic overview of an example quality control process is provided. The table in Panel A provides the quality control steps and inclusion thresholds applied to the three imaging modalities in the above example.
[0110] In the database query, the retinal scans are divided by modality, with 51,376 FAF scans, 43,746 IR scans, and 1,095,082 SD-OCT scans (leaving 5,834 images of other modalities). Since SD-OCT produces multiple B-scans, for each B-scan, only the image corresponding to the central B-scan traversing the fovea is used, as it is likely the most informative B-scan. After this, 33,849 OCT B-scans remain.
[0111] For all three modalities, filtering is applied. Any corrupted scans are discarded, as well as any scans that are not of size 768×768 pixels (for FAF scans and IR scans) or 512×496 pixels (for OCT scans). Additionally or alternatively, the size of the scans can be adjusted. Both IR scans and FAF scans have two different imaging magnification levels, namely 30 degrees and 55 degrees. In each case, only the most common mode (55 degrees for FAF and 30 degrees for IR) is retained, and all other scans are discarded. These two modes can be automatically distinguished by examining the number of black background pixels (RGB 0,0,0), where a total of 107,577 black pixels are found to correspond to the 55-degree image.
[0112] To remove low-quality and defective images, multiple quantitative filters are applied. The type of filter can be determined by an initial examination of the data to identify any specific problems. The thresholds for these filters are set by: examining the scan samples rejected by the filter and adjusting the threshold until more than 25% of the rejected scans are judged to have a reasonable quality level.
[0113] In particular, the FAF dataset and the IR dataset contain multiple scans that are too dark. Thus, for FAF, scans with a median pixel intensity below 0.05 are rejected, and for IR, scans with a median pixel intensity below 0.1 are rejected.
[0114] Furthermore, many FAF scans have a large amount of noise. To measure this noise level, the inventors introduced a "pixel noise level" score, where the original scan is compared to a blurred version of the same image (using a normalized box filter with a 5×5 kernel), and the sum of the squared differences between the two images is taken. FAF scans with a total squared difference exceeding 2200 are rejected. Many OCT images contain artifacts consisting of large regions with a pixel value of 1.0 (i.e., the maximum value in this case). To remove these, any image with a maximum pixel value of 1.0 and a median intensity greater than 0.2 is rejected.
[0115] Finally, the BRISQUE score (Blind / Referenceless Image Spatial Quality Evaluator, a reference-free image quality score) is calculated for each scan using the PyBRISQUE library, and scans above a certain threshold are discarded. In this example, the threshold is set to 120 for FAF images, 80 for IR images, and 150 for OCT images.
[0116] After quality control and after restricting the dataset to only genes with at least 5 scans, 15,692 FAF scans, 23,631 IR scans, and 13,099 OCT scans remained in 36 of the most common genes, as shown at the bottom of Panel B.
[0117] Those skilled in the art will understand that the above specific preprocessing quality control procedures are dataset-dependent and that the exact parameters will depend on the dataset at hand.
[0118] Panel C provides examples of low-quality scans rejected by the above example filters. Sub-panel A is a FAF scan where it is too dark to discern any details. Sub-panel B is a FAF scan with excessive noise. Sub-panel C is an infrared scan where it is too dark to discern any details. Sub-panel D is an SD-OCT scan where large artifacts obscure part of the scan.
[0119] As described above, data augmentation can be used to avoid overfitting on training data. Retinal scans can be augmented before training such that the augmented training dataset includes the same number of retinal scans as the original training dataset or such that the augmented retinal scans are concatenated to the original training dataset.
[0120] Figure 7 Eight data augmentation techniques that can be applied to retinal scans are shown. The techniques depicted include horizontal flipping; rotation; brightness adjustment; random scaling; Gaussian blur; adding Gaussian noise; adding salt-and-pepper noise; and adding speckle noise. As can be seen from the images in the bottom row, techniques can be applied sequentially; for example, the original image is horizontally flipped and a brightness factor of -1 is applied, resulting in the third image from the left. Any or all of the data augmentation techniques can be applied to retinal scans obtained from any imaging modality.
[0121] Performance Evaluation - Effectiveness of Example Cross-Validation Test Set and Retained Validation Test Set
[0122] As in the above example, the constituent neural networks of the ensemble network can be trained by a 5-fold cross-validation method where each individual network is trained on a different subset of the training data.
[0123] For the above example, for the network trained on FAF images, the inventors observed an average top-1 accuracy across networks of 44.3% (CI 95%= 42.5 - 46.2), top-5 accuracy was 72.7% (71.0 - 74.4), and the average per-gene ROC AUC was 0.827 (0.812 - 0.842). For the network trained on IR images, the inventors observed an accuracy of 45.8% (43.6 - 47.9), top-5 accuracy of 73.7% (72.2 - 75.2), and AUC of 0.834 (0.824 - 0.844). Finally, for the network trained on SD-OCT images, the inventors observed an accuracy of 51.6% (48.4 - 54.9), top-5 accuracy of 77.4% (75.2 - 79.6), and AUC of 0.845 (0.831 - 0.859).
[0124] The results on the held-out patient sets were similar, with an average cross-model accuracy of 51.4% for FAF, 46.9% for IR, and 56.3% for the SD-OCT model (see Table 4 below; specifically individual model means, individual image parts (upper left)).
[0125] More specifically, to evaluate the efficacy of the ensemble network method, the inventors simulated the scenario of applying the method to scans taken during a patient visit by applying the method to scans from the internal held-out dataset (validation set) introduced above (see Figure 4 ). The Moorfields held-out dataset was used for internal testing. As described above, the MEH internal test dataset in this example consisted of 1900 scans from 264 patients diagnosed across 32 genes. The first row of Table 3 below gives the breakdown of scans for each modality.
[0126] Table 3 Number of patients, scans, and genes in the internal and external test datasets.
[0127]
[0128] The scans in the internal test dataset were grouped by patient and acquisition date. The predictions were compared to the ground truth gene diagnosis for the corresponding patients. Using only the scans from the first appointment for each patient, the ensemble network method achieved a top-1 accuracy of 66.7%, top-5 accuracy of 85.6%, and an average per-gene receiver operating characteristic AUC of 0.935 (see Table 4; 5-model ensemble, multiple image / modality parts (lower right)).
[0129] The ensemble network method was applied across all available scans for each patient's each appointment (defined as all scans from a given patient on a given date) and at each patient level (using all available scans for each patient across all appointments).
[0130] Table 4 provides an overview of the test accuracies on the held-out patient sets for different methods predicted using the combined model. Each visit refers to all the ocular scans for each patient visit. Each patient refers to all the ocular scans across multiple patient visits.
[0131]
[0132] For the modality-specific ensemble models, the model was applied to all the scans of that modality from the internal test set. The model predictions were then compared with the ground truth gene diagnosis for each patient to compute the overall accuracy of the model on the test data, the top-k accuracies (proportion of images where the correct gene is within the top k predictions of the network) for k = 2, 3, 5, 10, and the area under the receiver operating characteristic curve (AUROC) averaged per class.
[0133] The ROC (receiver operating characteristic) shows the trade-off between the specificity (false positive rate) and sensitivity (true positive rate) for a given gene. For each gene, the AUC (area under the curve) of the corresponding ROC curve of the model was obtained by comparing the model predictions for the given gene in a one-versus-all setting.
[0134] The confidence intervals for the accuracies and AUCs on the development data were obtained by taking the standard deviation of the accuracies / AUCs for each of the three modalities across 5 networks and estimating the 95% confidence interval through the expression.
[0135] Performance Evaluation - Ensemble Methods vs. Other Methods
[0136] By decomposing the ensemble network into its constituent networks, it can be seen that the overall accuracies across each modality are relatively similar. Figure 8 A collection of per-gene receiver operating characteristic (ROC) curves for cross-validation data for three different imaging modalities for 6 out of the 36 genes from the above example. The FAF classification is depicted by the dotted curve; the IR classification is depicted by the dashed curve; the OCT classification is depicted by the solid curve. Each curve takes the average ROC curve across 5 models for each modality. The percentage next to the gene name represents the percentage of the total images (across all three modalities) for the given gene. The shaded region represents the area within one standard error of the mean. For simplicity, only 6 genes (ABCA4, RPGR, CNGB3, TIMP3, CRB1, and PRPH2) out of the 36 classified genes in the example are depicted here.
[0137] It can also be seen that while the overall accuracy across each modality is relatively similar, there are some significant differences for certain genes. For example, the FAF network and the SD-OCT network are significantly better than the IR network in detecting TIMP3. This is notable because TIMP3-retinopathy is associated with increased signals in the perifovea in FAF imaging and drusen-like deposits in SD-OCT scans. Similarly, the SD-OCT network is better than the IR network and the FAF network in detecting CRB1. CRB1-retinopathy is typically characterized by a thicker retina with fewer defined layers, which cannot be distinguished in FAF or IR imaging.
[0138] Many genes follow a pattern similar to PRPH2, with little deviation between modalities. However, for some genes, there are significant differences in the ability of the networks of different modalities to distinguish the genes in question. For example, for CRB1, a gene associated with retinal thickening, SD-OCT is the most predictive modality.
[0139] Therefore, the integrated approach for retinal scan image classification is able to identify the best image modality at each level of lesion (e.g., at each gene level).
[0140] It was found that combining predictions across multiple models (ensemble) and across multiple images acquired during one or more patient visits contributed to the performance of the ensemble network.
[0141] To apply the ensemble of five models for each modality to individual images, the inventors observed accuracies of 59.2% / 55.3% / 61.5% for FAF / IR / OCT, top-5 accuracies of 81.3% / 80.8% / 82.9%, and AUCs of 0.881 / 0.885 / 0.911 (see Table 4; 5-model ensemble, single-image section (lower left)). This method produced improved results compared to the individual model average, single-image method.
[0142] For the individual model predictions that combine across multiple images at each visit level on the held-out test set in the above example (without ensembling the 5 models for each modality), the inventors saw an overall average accuracy across models of 62.8%, a top-5 accuracy of 84.6%, and an average AUC of 0.914 (see Table 4; single-model ensemble, multiple images / modalities section (upper right), each visit column).
[0143] Figure 9It is a schematic overview of four types of models benchmarked here. Panel A shows a single-image model that takes a single image from one eye as input and outputs a single classification probability from a single network. Panel B shows a single-image ensemble model (five networks for each modality) that takes a single image from one eye as input and outputs classification probabilities from each of the five networks in the ensemble, which are combined into a single classification probability. Panel C shows a multi-image model (a single network for each modality) that takes multiple images from different modalities (i.e., FAF, OCT, and IR) as input and outputs classification probabilities for each image, which are combined into a single classification probability for the patient. Panel D shows a multi-image ensemble (five networks for each modality) that takes multiple images from different modalities (i.e., FAF, OCT, and IR) as input and outputs classification probabilities from each of the five networks for each image, which are combined into a single classification probability for the patient.
[0144] Table 5 below summarizes the comparison of the accuracy when classifying multiple scans using a single convolutional neural network for each imaging modality versus when using an ensemble network of 5 models for each imaging modality. Note that the single-image accuracy results (left column) are the same as the single-image accuracy results provided in Table 4 (left column) and are included here for ease of reference.
[0145] Table 5 Overview of test accuracy on the held-out patient set using single scans versus multiple scans.
[0146]
[0147] The inventors found that, based on per-gene ROC AUC, combining all images across all modalities performs better than the best-performing modality (i.e., images limited to that modality) for most genes, demonstrating the advantage of the multimodal approach. Figure 10 is a collection of per-gene ROC curves (black solid curves; "E2G ensemble") from cross-validation data used to make combined predictions for all images from a given patient appointment, overlaid on the multi-image ROC curves when limited to images of a given modality. The percentage next to the gene name represents the percentage of the total appointments corresponding to the given gene. For brevity, only 6 of the 36 ROC curves for all gene classifications in this example are provided (the same 6 as in Figure 8 . This again demonstrates the advantage of the multimodal approach.
[0148] In cases where multiple models for each modality applied to individual images are integrated and individual model predictions across multiple images are combined, these cases provide results that are better than single-network results but not as good as the holistic integrated model approach employed in the embodiments, indicating that both integrating and combining predictions across multiple images are beneficial.
[0149] Performance Evaluation - Robustness
[0150] To ensure that the results regarding the above example cases are generalizable across hospitals, the inventors applied the classifier to data from four IRD clinics, namely Oxford Eye Hospital (UK), Liverpool University Hospital (UK), University Hospital Bonn (Germany), and Federal University of Paulo (Brazil). As shown in Table 3 above, Oxford Eye Hospital (UK) provided a sample of 346 scans from 59 patients who had different genetic diagnoses in 29 different genes. Liverpool University Eye Hospital (UK) provided a sample of 210 scans from 35 patients who had different genetic diagnoses in 15 different genes. University Hospital Bonn Eye Hospital (Germany) provided a sample of 473 scans from 127 patients who had different genetic diagnoses in 11 different genes. Federal University of São Paulo (Brazil) provided a sample of 104 scans from 15 patients who had different genetic diagnoses in 3 different genes.
[0151] Patients and scans were selected by clinicians at these different clinics who were instructed to select patients with a confirmed genetic diagnosis who had multiple scans available for each modality. The selection of scans and the genetic diagnosis for each patient were shared with the inventors, except for the data from University Hospital Bonn, where they ran the embodiments locally at University Hospital Bonn to avoid having to transmit any patient data.
[0152] The inventors ran the integrated network on the images from each patient and compared the predictions of the integrated network with the ground truth genetic diagnosis. Combining all the data from all four sites (1133 retinal scans from 236 patients), an overall accuracy of 65.3% and a top 5 accuracy of 86.4% were found (the breakdown of the results at each site level is shown in Table 6 below).
[0153] Since the gene distribution in the external test data is different from the main MEH dataset (and thus different from the internal holdout test set), and the accuracy of the ensemble network in this example varies by gene, the inventors also recorded the per-gene prevalence (proportion of the dataset) and sensitivity (accuracy on that particular gene) for each gene at each external site. To understand to what extent the differences in accuracy on the external data are due to differences in gene distribution, for each site, the inventors calculated a reweighted accuracy on the internal test data to match the gene distribution of the target external dataset. Doing so reveals that, assuming gene-for-gene performance consistency, one would expect to see an overall accuracy of 67.2% on the external data, which is still slightly higher than the actual figure of 65.3%. This pattern is consistent across all four sites, with the actual accuracy at all four sites found to be a few percentage points lower than predicted by the internal test data.
[0154] In summary, when faced with previously unseen datasets, the ensemble network trained according to the embodiments is robust and generalizable across sites.
[0155] Table 6 Overview of the results of the ensemble network on external data across different sites. Gene-For-Gene Extrapolation from the internal dataset is the sum of the per-gene sensitivities on the MEH internal test dataset multiplied by the gene frequencies of the external dataset, and represents the accuracy one would expect to see for the gene distribution of that site.
[0156]
[0157] Performance Evaluation - Against Human Experts
[0158] To consider the performance of this example embodiment relative to human experts, the inventors asked nine ophthalmologists with different levels of experience to predict disease-causing genes based on individual FAF images from the internal holdout test set in the above example across 50 different patients (which were selected to obtain a uniform representation of all genes in the dataset). In this task, the ophthalmologists achieved an average accuracy of 17%, compared to an accuracy of 38% for the trained ensemble network on the same images. The inventors also asked four additional ophthalmologists to predict disease-causing genes selected from patients in the external test set from the University of Bonn Hospital, asking them to select only from 11 genes in the dataset. In this task, they achieved an average accuracy of 42%, compared to an accuracy of 49% for the trained ensemble network (where the trained ensemble network still had to select among 36 genes).
[0159] More specifically, the results of human benchmark testing by ophthalmologists are presented in Table 7 below. In this task, ophthalmologists were asked to identify the correct diagnostic gene when presented with FAF scans of patients with IRD. As expected, human performance tended to improve with increasing levels of experience. The inventors also applied the trained example integration network to the same dataset, limited to single image predictions (so only using the 5-model FAF integration) for fair comparison. The performance of the trained integration network was generally better than any single human expert (except for Ophthalmologist 11 who achieved an accuracy of 60.3%), and the frequency with which the correct gene appeared in the top 5 predictions of the trained integration network was comparable to its frequency among all human guesses.
[0160] Thus, the classification of retinal scan images using an exemplary trained integration network is on par with experts.
[0161] Table 7 For comparison, ophthalmologists and a trained integration network according to an embodiment classified the 50 FAF retinal scans of 50 patients from the internally held test dataset at Moorfields Eye Hospital and also the results on the Bonn external test dataset, which included 73 FAF retinal scans from 37 patients.
[0162] Note: Ophthalmologists 8 and 9 took the quiz together.
[0163]
[0164] Performance Evaluation - Ensemble Network Prediction Calibration
[0165] Although the example integration network outputs predictions in terms of the probability of each gene, for any given image, each underlying network typically predicts a single gene class with confidence, making it difficult to determine model confidence solely from the output of any single network.
[0166] However, the inventors found that when combining the constituent network predictions in the integration network, the output predictions appeared to be relatively well-calibrated, except for a consistent 10% over-prediction. For example, if the integration network predicts a given gene with a probability of 70%, that prediction is correct approximately 60% of the time. This is shown in Figure 11 which shows a comparison of the actual accuracy of the integration network on individual images versus the model confidence, along with the best fit line (slope = 1.04, intercept = -11.2%). Compared to the "well-calibrated" 1:1 correspondence with the actual accuracy (shown by the dotted line), the integrated predictions are slightly overconfident. Also shown is the total proportion of the dataset above each confidence threshold (dashed line).
[0167] This is useful for providing accurate and informative feedback to end-users regarding model confidence. In particular, this aspect can be used in the implementation of embodiments in a web application (app) (see below) to identify unclear input scans that may benefit from being retaken. This can be achieved by flagging individual images that fall below a certain confidence threshold when passed through the model.
[0168] Refining Ensemble Network Output - Including Age and Genetic Patterns
[0169] In addition to retinal scans, further information such as the age of the patient when admitted to the hospital and their family history can be used to determine eye diseases, such as causative genes. Since this information is often available along with the patient scans, the method for classifying retinal scan images using an integrated network can be modified to incorporate these additional features. In particular, the modification to the embodiment enables the combination of the age at first onset (age) and the mode of inheritance (MOI) of the patient with the output of the integrated network model, where the mode of inheritance (MOI) can be inferred from the family history.
[0170] One way to incorporate this information is by using a linear classifier, such as implementing logistic regression without regularization. For the above example, by training a linear classifier (using logistic regression without regularization) to predict the gene class based on the output of the integrated network as well as age and MOI, it was found that this increased the overall accuracy from 66.7% to 68.9% and the top-5 accuracy from 85.9% to 89.4%. In this case, this did not improve the average per-gene AUROC, which remained at 0.935. A detailed breakdown of the efficacy of the method combination is shown in Table 8 below.
[0171] Table 8 Results of various classification methods on the internal test dataset.
[0172]
[0173]
[0174] More specifically, age and mode of inheritance (MOI) are important factors to consider in the diagnosis of IRD genes because these provide strong prior evidence for specific genes. For example, if the patient is young and the MOI is X-linked, this may indicate that RP2 is the associated gene. Figure 12Shows the distribution of the earliest age of onset for each genetic diagnosis for the examples described throughout this specification. The horizontal axis provides the age of the subject at which each genetic diagnosis (gene abnormality as provided in the vertical axis) first occurred. For example, the earliest age of onset for the CACNA1F gene abnormality is within childhood and is most strongly present around 1 to 12 years of age; the earliest age of onset for the MTTL1 gene abnormality is in a later stage of life and is most strongly present in the late forties and early fifties. Figure 13 Shows the known inheritance patterns for each gene for the same examples; for example, the MOI for an abnormality in the MTTL1 gene may be mitochondrially based. Note that some genes (such as BEST1 and PROML1) can be either dominant or recessive.
[0175] To combine these non-imaging features, a model stacking approach can be employed, where a linear classifier can be trained on the combined predictions from an ensemble network and additional input features, and then this linear classifier produces a single prediction based on the original model predictions and additional features.
[0176] In this example, age is treated as a normalized numerical variable (subtracted by the mean and divided by the standard deviation), and MOI is treated as a one-hot vector with five categories: recessive, dominant, X-linked, mitochondrial, and unknown. For example, a recessive gene can be encoded with a 1×5 row vector with values [1 0 0 0 0]. For each image / appointment from the five patient sets in this example, the inventors concatenated these features with the final output of the corresponding network (where each gene prediction is treated as a separate feature) and used these together with the ground truth gene labels to fit a logistic regression classifier (without regularization).
[0177] The concatenation of these feature vectors with the predictions from the ensemble network from the cross-validation experiment (i.e., using only the test results of each network) forms a 41-wide feature vector, on which logistic regression is used to fit a 36-class linear classifier (targeting patient genes as described above). As described above, using the concatenated age + MOI features and ensemble network predictions, this trained classifier is applied to an internal test set, and the output is compared with the ground truth gene for each patient.
[0178] To test the resulting combined model, the inventors applied the trained model to an internal test set, using the ensemble predictions from the full ensemble network connected to age and MOI, and using the resulting predictions as new gene-level predictions, performing the same top-k and per-gene average AUC analyses as for the other performance evaluations. The inventors also performed cross-validation analysis based on five folds in the development set, where the model outputs were used as training data on 4 folds and tested on the remaining fold, repeating this for each fold.
[0179] Figure 14 Set of ROC curves of linear classifiers trained on different feature combinations (age, inheritance pattern, and model output of the trained ensemble network ("Eye2Gene")) for the internal test data set (bottom row) using cross-validation data (top row). For the cross-validation data, the linear model was trained on four folds and tested on the remaining fold; the ROC curves shown are the average over five folds, and the shaded area is one standard error from the average. For the internal test data set, the linear model was trained on the development set and tested on the internal test data set. In both cases, the improvement brought about by refining the output of the trained ensemble network with further information can be seen.
[0180] Those skilled in the art will understand that other patient characteristics can be integrated with the classifier to further refine the ensemble network output. For example, the patient's ethnic background can be additionally or alternatively encoded for classification. In addition, those skilled in the art will understand that classifiers implementing other algorithms can be used, such as logistic regression with regularization.
[0181] Packaging Retinal Scan Image Classification Method
[0182] The retinal scan image classification method described herein can be used to assist clinicians in diagnosing IRD patients. The ensemble network can be deployed as, for example, an online web application. That is, the trained ensemble network can be accessed as an online application.
[0183] Figure 15This is an example of a suitable screen of a graphical user interface (GUI) that provides example results to a user after performing an integrated network. The screen (and any other screen) can be presented to the user on a handheld device or a computer screen. To use the example application, the user uploads a series of scans (e.g., as PNG images) and can enter basic case information such as age at onset and MOI (if known). In this case, an age of 29 and an MOI of "occult" are used. The user is then requested to specify the scan type for each image (in this example, choose between SD-OCT, 55-degree FAF, and 30-degree IR). These images are then passed to a trained integrated network (to one or all of the trained CNNs for the relevant imaging modality), which outputs a set of prediction scores for each of the 36 genes for each input scan. This information is then aggregated (or integrated) into an overall prediction score for the case and presented to the user.
[0184] In this example, a bar graph of the top 5 genes predicted by the integrated network, along with the model probability scores for each gene, is presented to the user. These predictions are broken down into the contributions of three different modalities, which are shown as different colors or shadings on the bar graph. The input images and patient information color-coded by modality (here shown as a solid line bounding the FAF scan, a long dashed line bounding the OCT scan, and a short dashed line bounding the IR scan) are included at the top of the display. A full breakdown of the predicted probabilities for all 36 genes is given in a table below the bar graph. A link in the upper right corner of the table takes the user to a breakdown of the integrated network's predictions for each uploaded scan.
[0185] Applicability to Other Ocular Diseases
[0186] The above example provides a method and network for classifying retinal scan images, which classify retinal scans based on gene abnormalities. Different imaging modalities highlight different aspects of retinal diseases by providing different "views", and thus have the potential to contribute different complementary evidence for accurate disease diagnosis. Therefore, analyzing multiple imaging modalities in parallel provides more insight into the pathology of the disease than examining one imaging modality at a time.
[0187] Therefore, the multimodal method disclosed herein is applicable to all eye diseases (genetic and non-genetic), which include retinal disorders, age-related macular degeneration (AMD), diabetic retinopathy (DR), glaucoma, retinopathy of prematurity (ROP) and vein occlusion, optic disc diseases (e.g., glaucoma and optic neuropathy), and conditions affecting the cornea (e.g., keratoconus, corneal dystrophy, and keratitis).
[0188] For example, consider AMD. The blue autofluorescence imaging modality highlights the distribution of lipofuscin in the retinal pigment epithelium over a relatively wide 50-degree field of view. Lipofuscin accumulation is a byproduct of intracellular aging and is thus an important biomarker of retinal aging such as in AMD.
[0189] The optical coherence tomography (OCT) imaging modality highlights the different layers in the cross-section of the retina and can thus detect any abnormalities such as the accumulation of retinal fluid, as is the case in AMD.
[0190] Infrared imaging highlights abnormalities in the retinal vasculature and can thus identify abnormal vascularisation such as that which occurs in AMD.
[0191] This can be supplemented by OCT angiography imaging, which can highlight abnormal vascularisation in the different layers of the retina.
[0192] Depending on the location and stage of disease occurrence, the above imaging modalities can also be combined with wider-angle imaging to capture changes in the peripheral retina.
[0193] Thus, in the case of AMD retinal scan classification, the method according to the present invention can be employed for four (or five) imaging modalities and an integrated probability of the presence of AMD within the retinal scan can be generated.
[0194] Of course, a similar multimodal approach can be used to identify cases of glaucoma, where the disease can be detected in the cornea, optic disc, peripheral retina, and can be detected in OCT by thinning of the retinal nerve fiber layer.
[0195] Hardware
[0196] Figure 16 is a block diagram of a computing device (such as a data storage server) that embodies the present invention and can be used to implement aspects of the method for retinal scan image classification as described herein. The computing device includes a processor 993 and a memory 994. Optionally, the computing device further includes a network interface 997 for communicating with other computing devices.
[0197] For example, an embodiment can consist of a network of such computing devices. Optionally, the computing device further includes one or more input mechanisms (such as a keyboard and mouse 996) and a display unit (such as one or more monitors 995). The components can be connected to each other via a bus 992.
[0198] Memory 994 may include a computer-readable medium, which term may refer to a single medium or multiple media (e.g., a centralized or distributed database and / or associated cache and servers) configured to carry computer-executable instructions or have data structures stored thereon. The computer-executable instructions may include, for example, instructions and data that are accessible by a general-purpose computer, a special-purpose computer, or a special-purpose processing device (e.g., one or more processors) and that cause the general-purpose computer, special-purpose computer, or special-purpose processing device (e.g., one or more processors) to perform one or more functions or operations. Thus, the term “computer-readable storage medium” may also include any medium that is capable of storing, encoding, or carrying a set of instructions for execution by a machine and that causes the machine to perform any one or more of the methods of the present disclosure. Thus, the term “computer-readable storage medium” may be understood to include, without limitation, solid-state memory, optical media, and magnetic media. By way of example and not limitation, such computer-readable media may include non-transitory computer-readable storage media including random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage or other magnetic storage devices, flash memory devices (e.g., solid-state memory devices).
[0199] Processor 993 is configured to control the computing device and perform processing operations, such as executing code stored in memory 994 to implement various different functions of the retina scan image classification method or the retina scan image classification training process as described herein and in the claims.
[0200] Memory 994 may store data that is read and written by processor 993, such as data from training or classification tasks executed on processor 993. For example, memory 994 may store information about the selected (s) architecture of each CNN and may store weights for implementation using the (s) architecture.
[0201] As mentioned herein, processor 993 may include one or more general-purpose processing devices, such as a microprocessor, a central processing unit, a GPU, etc. Processor 993 may include a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or a combination of instruction sets. Processor 993 may also include one or more special-purpose processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. In one or more embodiments, processor 993 is configured to execute instructions for performing the operations and steps discussed herein.
[0202] The network interface (network I / F) 997 can be connected to a network (such as the Internet), and can be connected to other computing devices via the network. The network I / F 997 can control data input / output from / to other devices via the network.
[0203] The methods embodying aspects of the present invention can be executed on a computing device (such as Figure 16 the computing device shown in Figure 16 ). Such a computing device does not need to have each component shown in and can be composed of a subset of those components. The methods embodying aspects of the present invention can be executed by a single computing device communicating with one or more data storage servers via a network or by multiple computing devices operating in cooperation with each other. Cloud services implementing the computing devices can be deployed.
Claims
1. A computer-implemented method for classifying retinal scan images, the method comprising: receiving an input data set comprising at least one retinal scan obtained using one of a plurality of imaging modalities; passing the input data set through an ensemble network comprising a plurality of convolutional neural networks, wherein: each convolutional neural network is configured to classify a retinal scan obtained using one of the plurality of imaging modalities based on a plurality of eye diseases, each convolutional neural network generating a probability of the presence of each of the eye diseases within the retinal scan, and processing the retinal scan using the convolutional neural network corresponding to the imaging modality used to obtain the retinal scan; and generating an integrated probability of the presence of each of the eye diseases within the retinal scan.
2. The method for classifying retinal scan images according to claim 1, wherein: the input data set comprises a plurality of retinal scans, each retinal scan obtained using one of the plurality of imaging modalities, and each retinal scan is processed using the convolutional neural network corresponding to the imaging modality used to obtain the retinal scan.
3. The method for classifying retinal scan images according to any one of the preceding claims, wherein, the plurality of imaging modalities comprises any one or all of the following: fundus autofluorescence FAF; infrared IR; and spectral domain optical coherence tomography SD-OCT.
4. The method for classifying retinal scan images according to any one of the preceding claims, wherein, the plurality of eye diseases are genetic abnormalities.
5. The method for classifying retinal scan images according to claim 4, further comprising passing the integrated probability through a linear classifier, wherein, the linear classifier is configured to refine the integrated probability based on the age of the subject with genetic abnormalities and / or the mode of inheritance MOI.
6. The method for classifying retinal scan images according to any one of the preceding claims, wherein: the ensemble network comprises k convolutional neural networks for each imaging modality, and the retinal scan is processed using k convolutional neural networks, each convolutional neural network corresponding to the imaging modality used to obtain the retinal scan, the method further comprising: averaging the probabilities of the presence of each of the eye diseases within the retinal scan from each of the k convolutional neural networks.
7. The method for classifying retinal scan images according to any one of the preceding claims, further comprising: when the estimate of the presence of each of the eye diseases within the retinal scan is below a predefined confidence threshold, outputting a prompt to re-acquire the retinal scan.
8. The method for classifying retinal scan images according to any one of the preceding claims, further comprising processing the output data set to provide a classification of the genetic cause of the eye disease.
9. An ensemble network for classifying retinal scan images, particularly for use in the method according to any one of the preceding claims, the ensemble network comprising a plurality of convolutional neural networks, wherein: Each convolutional neural network is configured to classify a retinal scan obtained using one of a plurality of imaging modalities based on a plurality of eye diseases, and each convolutional neural network produces a probability of the presence of each of the eye diseases within the retinal scan.
10. A computer-implemented method for training an ensemble network for classifying retinal scan images, particularly the ensemble network according to claim 9, the method comprising: receiving an input training data set comprising training retinal scans obtained using an acquisition imaging modality and known eye disease labels, each training retinal scan being obtained using one of a plurality of imaging modalities; dividing the training retinal scans into a plurality of training retinal scan sets, each training retinal scan set comprising training retinal scans obtained using the same imaging modality; for each training retinal scan set, training a convolutional neural network to minimize the error between the predicted classification of a plurality of eye diseases and the known eye diseases; and for each convolutional neural network, outputting the trained convolutional neural network weights.
11. The method for training an ensemble network for classifying retinal scan images according to claim 10, wherein, training the convolutional neural network uses k-fold cross-validation such that the method comprises: for each training retinal scan set, training k convolutional neural networks.
12. The method for training an ensemble network for classifying retinal scan images according to claim 10 or claim 11, further comprising augmenting the input training data set by any or all of the following: rotation; flipping; scaling; cropping; adjusting brightness; blurring; and adding noise.
13. The method for training an ensemble network for classifying retinal scan images according to any one of claims 10 to 12, further comprising preprocessing the input training data set by any or all of the following: filtering by median pixel intensity; filtering by noise level; filtering by identification of image artifacts; and filtering by BRISQUE score.
14. A data processing apparatus comprising a memory and a processor, the data processing apparatus being configured to execute the method according to any one of claims 1 to 8 and claims 10 to 13.
15. A computer program comprising instructions which, when executed by a computer, cause the computer to execute the method according to any one of claims 1 to 8 and claims 10 to 13.
16. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to execute the method according to any one of claims 1 to 8 and claims 10 to 13.