Retinal scan image classification

An ensemble network of CNNs classifies retinal scans to accelerate genetic diagnosis of IRDs, addressing the challenge of limited expertise and timely diagnosis, enhancing diagnostic accuracy and accessibility.

JP2025536899APending Publication Date: 2025-11-12UCL BUSINESS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025520714
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2023-09-26
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

The diagnosis of inherited retinal diseases (IRDs) is challenging due to the limited distribution of expertise and the need for timely genetic diagnosis, which is often elusive in a significant number of cases, hindering access to effective diagnostic services and gene-directed therapies.

Method used

A computer-implemented method using an ensemble network of convolutional neural networks (CNNs) to classify retinal scans from multiple imaging modalities, generating probabilities for ocular pathologies, which can be supplemented by a linear classifier for improved accuracy, facilitating rapid genetic diagnosis.

Benefits of technology

The method accelerates the genetic diagnosis of IRDs by providing broad access to diagnostic expertise, improving classification accuracy, and reducing the time to diagnosis, especially for rare diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536899000001_ABST
    Figure 2025536899000001_ABST
Patent Text Reader

Abstract

A computer-implemented method for retinal scan image classification, the method including: receiving an input dataset including at least one retinal scan, the retinal scan acquired using one of a plurality of imaging modalities; passing the input dataset through an ensemble network, the ensemble network including a plurality of trained convolutional neural networks, each convolutional neural network configured to classify retinal scans acquired using one of a plurality of imaging modalities based on a plurality of ocular lesions, each convolutional neural network generating a presence probability for each ocular lesion in the retinal scan, the retinal scan being processed using a convolutional neural network corresponding to the imaging modality used to acquire the retinal scan; and generating an ensembled presence probability for each ocular lesion in the retinal scan.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for classifying retinal scan images, a neural network for use in classifying retinal scan images, and a method for training a network for use in classifying retinal scan images. Further, the present invention relates to a data processing apparatus, a computer program, and a computer-readable medium for classifying retinal scan images. In particular, the present invention relates to a method and network for estimating the probability of the presence of multiple ocular pathologies in a retinal scan. [Background technology]

[0002] Inherited retinal diseases (IRDs) are a group of rare genetic disorders that affect 1 in 3,000 people, and more than 300 IRD-associated genes have been identified to date. IRDs cause a decline in the function of the retina, the light-sensitive tissue at the back of the eye that is responsible for vision. Some patients with IRDs are born with severely impaired vision, while others experience a gradual decline in peripheral and central vision over time. Collectively, IRDs are a leading cause of childhood blindness and the most common cause in the working-age population, with significant psychological and socioeconomic impacts.

[0003] Identifying the genetic cause of IRD is a prerequisite for optimal prognosis determination, provision of genetic counseling, eligibility for gene-directed therapy, and inclusion in gene-directed clinical trials. However, genetic diagnosis remains elusive in more than 40% of cases, primarily due to insufficient evidence, clinical experience, and inefficiencies in diagnostic services, including lack of access to and inefficiencies in diagnostic services, such as the ability to effectively link and communicate data and findings between research groups. In many parts of the world, nearly 100% of cases remain genetically unresolved due to a lack of access to and funding for molecular testing.

[0004] IRDs can have distinct phenotypic characteristics that clinicians can learn to recognize using modern retinal imaging techniques that rapidly and noninvasively acquire images of the retina with dilated pupils, with minimal patient inconvenience. These scans can be performed with a variety of imaging modalities, including fundus autofluorescence (FAF), infrared (IR) imaging, and spectral-domain optical coherence tomography (SD-OCT), each of which conveys different information about retinal structure.

[0005] Figure 1 shows examples of retinal scans (also called ocular scans or images, and ophthalmic scans or images) acquired using these three example imaging modalities. Panel A is an IR fundus scan acquired at 30 degrees. Panel B is an FAF scan acquired at 55 degrees. Panel C is a single B-scan from an SD-OCT volume.

[0006] For example, FAF imaging can help identify both photoreceptor dysfunction and apoptosis by specifically detecting lipofuscin and related compounds, which accumulate in the retinal pigment epithelium (RPE) primarily due to oxidative stress, through their autofluorescence properties. IR imaging can also help identify vascular abnormalities, as well as melanin, a naturally occurring pigment that protects the eye, and melanin-lipofuscin complexes, which are granules containing both melanin and lipofuscin. SD-OCT can help identify and characterize multiple retinal cell layers, including the RPE, photoreceptors, outer nucleus and plexus, inner nucleus and plexus, ganglion cells, and nerve fiber layer. The primary cells affected by IRD are photoreceptors, and SD-OCT can be used to distinguish between the inner and outer layers of photoreceptors and visualize the highly reflective ellipsoid zone (EZ), which is often destroyed in the early stages of IRD. DISCLOSURE OF THE INVENTION [Problem to be solved by the invention]

[0007] This high-resolution, detailed, multimodal information allows ophthalmologists to identify some gene-specific disease patterns and, in some cases, predict disease-associated genes. However, due to the small number of these diseases, the expertise and experience required for accurate clinical diagnosis is not widely distributed and is limited to a handful of specialists and hospitals that have developed this expertise over decades.

[0008] This limitation is particularly evident when considering the diagnosis of IRD. Clinical trials are now available to an increasing number of patients, and approved treatments are becoming available. However, this requires that a genetic diagnosis be established early enough. Importantly, identifying the genetic cause in a timely manner remains challenging.

[0009] The present inventors have realized that it would be desirable to timely identify ocular pathologies, particularly ocular pathologies that have a genetic determinant.

[0010] According to one aspect of the present invention, a computer-implemented method for retinal scan image classification is provided. The method includes receiving an input dataset including at least one retinal scan, the retinal scan being acquired using one of a plurality of imaging modalities. That is, following acquisition of the retinal scan (or possibly multiple retinal scans), the classification method includes accepting the retinal scan(s) for processing. This can be done, for example, directly via an acquisition device or by a user uploading the retinal scan for processing.

[0011] The method further includes passing the input dataset through an ensemble network, the ensemble network comprising a plurality of convolutional neural networks (CNNs), each configured to classify retinal scans acquired using one of a plurality of imaging modalities based on a plurality of ocular pathologies. That is, a first CNN is configured to classify retinal scans acquired using the first imaging modality based on ocular pathologies, e.g., pathologies having a genetic basis, a second CNN is configured to classify retinal scans acquired using the second imaging modality based on ocular pathologies, and so on. Each CNN generates a probability of the presence of each ocular pathology in the retinal scan. The retinal scans (passed to the ensemble network) are processed using a CNN corresponding to the imaging modality used to acquire the retinal scans. For example, if the retinal scan(s) were acquired using the first imaging modality, the retinal scan(s) are processed using the CNN corresponding to the first imaging modality. The CNNs are trained. A CNN (or a group thereof) for a single imaging modality may be referred to as a prediction block or a modality-specific ensemble model.

[0012] The CNN can be configured to output probabilities for a predefined number of classes, where the number of classes corresponds to the number of pathologies under investigation. For example, if the ensemble network classifies retinal scans with respect to 10 classes, then the ensemble network will output 10 probabilities.

[0013] Ensembled probabilities are not intended to replace genetic testing, as phenotyping cannot fully replace molecular testing, especially when treatments such as gene therapy are implemented based on a genetic diagnosis. Rather, retinal scan image classification techniques facilitate access to diagnostic expertise currently available only at a limited number of specialized centers worldwide, thereby dramatically accelerating the journey to a genetic diagnosis. Especially given that the number of treatable IRDs is increasing and rapid diagnosis leads to improved patient outcomes, using CNNs to classify ocular pathologies is a promising approach to accelerating medical diagnostics, such as genetic diagnosis for IRD patients.

[0014] The method further includes generating an ensemble (or aggregate) probability for the presence of each of the ocular pathologies in the retinal scan. If only a single retinal scan is passed through the ensemble network and there is only a single CNN for the imaging modality used to acquire the single retinal scan, the ensemble probability is the probability generated from that CNN. If only a single retinal scan is passed through the ensemble network and there are multiple CNNs in the modality-specific prediction block for the imaging modality used to acquire the single retinal scan, the ensemble probability may be the average of the probabilities generated from the multiple CNNs. The ensemble probability may be output in the form of a list of probabilities, one for each class of the classification.

[0015] Optionally, the input dataset may include multiple retinal scans, each acquired using one of multiple imaging modalities. In this case, each retinal scan can be processed using a CNN (or multiple CNNs in a modality-specific prediction block) corresponding to the imaging modality used to acquire the retinal scan. For example, the input dataset may include at least one retinal scan from each of the imaging modalities. The input dataset may include at least one retinal scan from each of two imaging modalities. That is, if a first retinal scan acquired using a first imaging modality and a second retinal scan acquired using a second imaging modality are passed through an ensemble network, the ensembled probability is some function (e.g., the average) of the probabilities generated from the first CNN (for the first imaging modality) and the second CNN (for the second imaging modality). Using multiple retinal scans from the same imaging modality increases the reliability of the CNN(s) for classifying eye pathologies for that imaging modality. The use of multiple retinal scans from different imaging modalities similarly increases the confidence in ocular pathology classification.

[0016] Optionally, the multiple imaging modalities may include any or all of the following: fundus autofluorescence (FAF); infrared (IR); and spectral-domain optical coherence tomography (SD-OCT). These are commonly used imaging modalities for acquiring retinal scans. Ensemble networks constructed with multiple CNNs corresponding to these imaging modalities have been shown to be successful in identifying ocular pathologies. For example, with the revolutionary improvements in imaging technology over the past decade and the incorporation of comprehensive genetic testing for IRD into specialized medical services, there are now a sufficient number of molecularly characterized patients with detailed retinal phenotypes (e.g., using the above imaging modalities to construct a representative dataset for deep learning). Incidentally, FAF and IR scans may be acquired at different angular spreads, e.g., 30 degrees and 55 degrees, and these may be considered the same or different imaging modalities.

[0017] Optionally, multiple ocular lesions are genetic abnormalities (or mutations). Ensemble networks are thus well-suited to assist in the detection of rare eye diseases, such as IRD, which are difficult to genetically diagnose. IRD is typically a monophenotype and is the leading cause of blindness in children and adults worldwide. AI systems trained to detect gene-specific patterns from retinal scans may provide broader access to expertise previously limited to human experts in this field. Even when limited to a single imaging modality, the accuracy of ensemble networks is at least as good, and often significantly better, than human experts in this task when classifying genetic abnormalities.

[0018] Optionally, the method further includes passing the ensemble probabilities through a trained linear classifier, which supplements or refines the ensemble probabilities output from the CNN(s) based on additional information encoded in the classifier. For example, the linear classifier can be configured to refine the ensemble probabilities based on the subject's age and / or the mode of inheritance (MOI) of the genetic abnormality. The use of a linear classifier has been shown to improve classification accuracy. Other factors, such as the subject's ethnicity, can additionally or alternatively be encoded into the trained linear classifier.

[0019] Optionally, the ensemble network can be composed of k CNNs per imaging modality. A retinal scan is processed using k CNNs, each corresponding to an imaging modality used to acquire the retinal scan. That is, multiple trained CNNs for a single imaging modality can be applied simultaneously to a single retinal scan. The method can then include averaging the respective probabilities of presence of ocular pathologies in the retinal scan from each of the k CNNs. The averaged probabilities (there can be multiple sets, one set for each group of CNNs) are then used to generate an ensemble probability. Using multiple CNNs per imaging modality has been shown to improve classification accuracy over using a single CNN.

[0020] Optionally, the method can further include outputting a prompt to reacquire the retinal scan if the estimate of the presence of each ocular pathology in the retinal scan falls below a predefined confidence threshold. For example, when the training dataset used to train the CNN is sparse (e.g., retinal scans of rare pathologies), the trained CNN tends to overpredict the most common classes and underpredict rare classes in the presence of uncertainty, and introducing a threshold can provide additional confidence in retinal scan image classification. Furthermore, poor quality retinal scans, especially when associated with age, can confound results. For example, retinal scans of young children tend to be noisy due to poor performance or movement during image acquisition.

[0021] Optionally, the method may further process the output dataset to provide a classification of the genetic cause of the ocular pathology. For example, if the ocular pathology is genetically based, the method can indicate specific genes to investigate, thereby shortening the time to diagnosis. This method can help the majority of ophthalmologists and other vision care professionals who do not specialize in rare IRDs to indicate when molecular testing should be considered. Gene predictions can also be used to score and prioritize genetic variants whose genes are consistent with the phenotype. For diseases with very distinct phenotypes, the ACMG guideline's "concordant with phenotype" PP4 criterion provides a higher level of evidence in variant classification, which can be crucial in moving from variants of unknown significance to variants with a high likelihood of pathogenicity.

[0022] According to one aspect of the present invention, there is provided an ensemble network for retinal scan image classification, particularly for use in the above-described method of retinal scan image classification, the ensemble network comprising a plurality of CNNs, each configured to classify a retinal scan obtained using one of a plurality of imaging modalities based on a plurality of ocular pathologies, and each CNN generating a probability of the presence of each of the ocular pathologies in the retinal scan.

[0023] According to one aspect of the present invention, there is provided an ensemble network for retinal scan image classification, and in particular a computer-implemented method for training such an ensemble network. The training method includes receiving an input training dataset including training retinal scans labeled with acquisition imaging modalities and known ocular pathologies, where each training retinal scan is acquired using one of a plurality of imaging modalities. In one example, the training dataset can be a combination of retinal scans and a lookup table (e.g., a CSV document) that references each of the retinal scans along with various file metadata (e.g., patient ID, laterality, date, file name, scan ID, scan number, genes, modality, etc.).

[0024] The training method includes partitioning the training retinal scans into multiple sets of training retinal scans, each set including training retinal scans acquired using the same imaging modality. The partitioning may further generate a held-out validation set, i.e., this set is not used in actual training, but instead is used to quantify the effectiveness of the trained CNN.

[0025] The training method involves training at least one CNN to minimize the error between predicted classifications of multiple eye pathologies and known eye pathologies for each set of training retinal scans. Training can be done from scratch (i.e., from a CNN architecture with randomized weights and biases) or by fine-tuning the CNN's existing weights (e.g., from ImageNet weights). Fine-tuning avoids excessive computational costs.

[0026] The training method includes outputting the trained CNN weights for each CNN. After training (e.g., after a predefined number of epochs or when the training loss meets a predefined criterion), the CNN is "frozen" and the checkpoints or weights are saved. The output weights can be saved locally or transmitted elsewhere, allowing the retinal scan image classification method to be run locally or elsewhere without having to run the training each time.

[0027] Optionally, the training method may include training k CNNs for each set of training retinal scans using k-fold cross-validation. Training multiple networks for each imaging modality has been shown to improve classification accuracy over training only a single network for each. The value k in this example can be determined depending on the training dataset at hand. Suitable values ​​include 2, 5, 10, 20, or n, where n is the dataset size for each test sample (leave-one-out cross-validation). Of course, the number of CNNs for one imaging modality does not have to be the same as the number for a second imaging modality.

[0028] Optionally, the training method may further include augmenting the input training dataset. For example, the training dataset may be augmented by any or all of rotating, flipping, zooming, cropping, adjusting brightness, blurring, and adding noise. Of course, those skilled in the art will recognize that other augmentation techniques are available and suitable. Augmenting the training dataset can increase the size and diversity of the dataset, thereby countering certain biases in the training data. In this way, the training method can avoid overfitting.

[0029] Optionally, the training method may further include preprocessing the input training data set. For example, the training data set may be preprocessed by filtering by median pixel intensity, by noise level, by identifying image artifacts, and / or by BRISQUE score. Of course, those skilled in the art will recognize that other preprocessing techniques are available and suitable.

[0030] Other aspects of the present invention include a data processing apparatus comprising means suitable for performing a method for classifying retinal scan images according to an embodiment, or a method for training an ensemble network for classifying retinal scan images according to an embodiment, wherein pre-processing removes poor quality and defective training images, thereby resulting in a more accurate trained CNN.

[0031] Another aspect of the present invention includes a computer program comprising instructions that, when executed by a computer, cause the computer to perform a method of retinal scan classification according to an embodiment or a method of training an ensemble network for retinal scan classification according to an embodiment. The computer program can be stored on a computer-readable medium. The computer-readable medium can be non-transitory.

[0032] Accordingly, another aspect of the embodiment includes a non-transitory computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform a method of retinal scan image classification according to an embodiment, or a method of training an ensemble network for retinal scan image classification according to an embodiment.

[0033] The invention can be implemented in digital electronic circuitry, in computer hardware, firmware, software, or in combinations of them. The invention may also be implemented as a computer program or computer program product, i.e., a computer program embodied in a non-transitory information carrier, for example a machine-readable storage device or a propagated signal, for execution by, or to control the operation of, one or more hardware modules.

[0034] A computer program may be in the form of a stand-alone program, a computer program portion, or multiple computer programs, may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a data processing environment. A computer program may be deployed to be executed on one module or on multiple modules at one site, or may be distributed across multiple sites and interconnected by a communications network.

[0035] The method steps of the present invention may be performed by one or more programmable processors executing computer programs that perform the functions of the present invention by operating on input data and generating output. Apparatus of the present invention may be implemented as programmed hardware or as special purpose logic circuitry including, for example, an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0036] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random access memory, or both. An essential element of a computer is a processor for executing instructions coupled to one or more memory devices for storing instructions and data.

[0037] The present invention is described with reference to specific embodiments. Other embodiments are within the scope of the following claims. For example, the steps of the present invention can be performed in a different order and still achieve desirable results. Multiple test script versions can be edited and launched as a unit without using object-oriented programming techniques. For example, elements of script objects can be organized in a structured database or file system, and operations described as being performed by the script objects can be performed by a test control program.

[0038] Elements of the present invention have been described using terms such as "processor," "input device," and the like. Those skilled in the art will understand that such functional terms, and their equivalents, may refer to parts of a system that are spatially separated but combined to perform a defined function. Similarly, the same physical portion of a system may provide two or more of the defined functions. For example, separately defined means may be implemented using the same memory and / or processor, as appropriate.

[0039] By way of example only, reference is made to the accompanying drawings in which: [Brief explanation of the drawings]

[0040] [Figure 1] Figure 1 is a composite image of retinal scans acquired with three imaging modalities commonly used to examine the retina of patients with inherited retinal diseases. [Figure 2] FIG. 2 is a flow diagram providing a method for classifying retinal scans according to an embodiment. [Figure 3] FIG. 3 shows the contribution of each of the 15 CNNs in an example image classification method using three imaging modalities. [Figure 4] FIG. 4 is a diagram illustrating an example of 5-fold cross-validation for training an ensemble network. [Figure 5] FIG. 5 shows a plot of the training loss over 100 training epochs in an example training process. [Figure 6] FIG. 6 provides an overview of the pre-processing filtering process used to filter the exemplary training dataset. [Figure 7] FIG. 7 shows an overview of the data augmentation used to expand the exemplary training dataset. [Figure 8] FIG. 8 shows a series of receiver operating characteristic (ROC) curves obtained from cross-validation data of three different imaging modalities across six of the 36 predictive genes in one example. [Figure 9] FIG. 9 illustrates four different approaches to retinal scan image classification that were considered during performance evaluation. [Figure 10] Figure 10 shows a set of ROC curves obtained from cross-validation data for composite prediction using ensemble networks, overlaying the prediction values ​​of three different imaging modalities across six of the 36 predictor genes in one example. [Figure 11] Figure 11 is a calibration chart of a single retinal scan processed using the ensemble network. [Figure 12] FIG. 12 shows the distribution of subject ages at earliest presentation of a genetic diagnosis in an exemplary training dataset. [Figure 12C] FIG. 12C shows the distribution of subject ages at earliest presentation of a genetic diagnosis in an exemplary training dataset. [Figure 13]FIG. 13 shows the known inheritance patterns for each gene in an exemplary training dataset. [Figure 13C] FIG. 13C shows the known inheritance pattern for each gene in an exemplary training dataset. [Figure 14] FIG. 14 shows a set of ROC curves obtained from cross-validation data for linear classifiers trained on different feature combinations across 6 of the 36 predictive genes in one example. [Figure 15] Figure 15 is a representative screen shot of the graphical user interface showing six retinal scans as input to the ensemble network and the resulting ensembled probability of genetic abnormality. [Figure 15C] FIG. 15C is a representative screen shot of the graphical user interface showing six retinal scans as input to the ensemble network and the resulting ensembled probability of genetic abnormality. [Figure 16] FIG. 16 is a diagram of hardware suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0041] FIG. 2 is a flow chart illustrating a computer-implemented method for classifying retinal scan images.

[0042] At S10, a computer receives an input data set including at least one retinal scan, the retinal scan being acquired using one of a plurality of imaging modalities.

[0043] At S20, the computer passes the input data set through an ensemble network, the ensemble network comprising a plurality of CNNs, each configured to classify a retinal scan acquired using one of a plurality of imaging modalities based on a plurality of ocular pathologies, each CNN generating a probability of the presence of each of the ocular pathologies in the retinal scan. The retinal scan is processed using the CNN corresponding to the imaging modality used to acquire the retinal scan.

[0044] At S30, the computer generates an ensemble probability of the presence of each of the ocular pathologies in the retinal scan.

[0045] The ensemble network used in the retinal scan image classification method includes at least one neural network for each imaging modality used to acquire the retinal scans intended to be processed through the ensemble network. For example, if there are two different imaging modalities used, such as FAF and SD-OCT, there will be at least one neural network intended to process FAF scans and at least one neural network intended to process SD-OCT scans. Of course, multiple neural networks may be used for each image modality.

[0046] As an example, consider an ensemble network composed of 15 neural networks. In this example, each neural network is an Inception-v3 deep CNN (CNN) capable of ingesting one or more retinal scans from three different imaging modalities from a given patient. Each constituent neural network in this example outputs a gene-level prediction score for 36 individual IRD genes. These 36 genes were selected in this example because, collectively, they cover over 80% of IRD cases in European populations.

[0047] Given a single input retinal scan in one of the three supported modalities, the exemplary ensemble network can apply each of the five networks corresponding to that scan's modality to obtain a single ensemble image-level prediction. For a series of retinal scans from a single patient, the appropriate group CNN can be applied to each scan in turn. The resulting network-level and image-level predictions can be combined to generate a single ensemble prediction (or list of probabilities) for that patient by averaging the individual (post-softmax) predictions.

[0048] Figure 3 is an illustration showing the contribution of each of the 15 constituent CNNs in the above example. Each imaging modality-specific prediction block (FAF, IR, SD-OCT) is composed of five associated CNNs configured to provide IRD gene predictions (i.e., predictions of genetic abnormalities or mutations) given retinal scans of the associated imaging modality. An example retinal scan is inset within each prediction block. In this example, the CNNs each provide predictions for 36 individual genes. Given a collection of retinal scans from a patient, the appropriate prediction block corresponding to the modality of each scan is applied.

[0049] The predictions of all CNNs within each modality-specific prediction block are averaged. For example, in the illustrated FAF prediction block, the estimated probabilities of five separate gene abnormalities—BEST1, ABCA4, CNGB3, and PRPH2 (thin bars)—are averaged to provide the average estimated probability of each gene abnormality (thick bottom bar). For each imaging modality, only the top four genes (by average estimated probability) are depicted. That is, for retinal scans of a particular modality, a five-model prediction block can be applied by averaging the output probabilities per gene.

[0050] The average estimated probability across all scans (from all imaging modalities) in the collection can be taken and used as the final prediction, or ensemble probability, for the collection of retinal scans. As seen in the right-most panel (outside the prediction block), the ensemble probability is obtained by combining the predictions of the 15 constituent networks using an ensemble approach.

[0051] Naturally, when retinal scans acquired using only a single imaging modality are processed, the ensembled probability corresponds to the average estimated probability of a single prediction block (for the same single imaging modality).

[0052] In the above example, the number of constituent CNNs (15) is shown merely as an example. Experienced readers will recognize that other numbers may be used. 15 was chosen in this example to allow for 5-fold cross-validation (discussed later in ensemble network training). In K-fold cross-validation, the value of k can be chosen so that each training and testing data sample is large enough to be statistically representative of the broader dataset. In this case, k=5 proved appropriate.

[0053] study Any known training technique suitable for training image classification models can be applied to the constituent CNNs of the embodiments, provided that each resulting trained neural network is capable of performing retinal scan image classification.

[0054] In this case, all training code was written in Python using the Keras library with the TensorFlow backend. Images were loaded into the model training routine via the built-in Keras data loader using Inception-v3's preprocessing functions. This loads the images as RGB images, resizes them to the correct input dimensions (256x256 pixels), and rescales pixel values ​​to be between -1 and 1 (divide by 127.5 minus 1). For convenience, we load three color channels (all set to the same value) even though the input scan in this example training process is monochromatic.

[0055] To load the images during the training phase, the images were automatically augmented by applying random transformations. [Using Transformations]

[0056] For the network architecture, we used the built-in Inception-v3 model in Keras using the saved ImageNet weights, removed the final output layer, and included a dropout layer (which was ultimately not used in the final experiments in this example). The final output layer had 36 outputs and softmax normalization.

[0057] The network was trained in the standard manner, using the Adam optimiser for weighted cross-entropy loss of predicted gene classes against true genes. This was done using the Keras fit function, with logging and monitoring via TensorBoard. The network weights were saved at the end of training (after 100 epochs), and we also saved a network configuration file containing other network settings for record-keeping and easy loading of future networks via a wrapper class. Classes were reweighted inversely proportional to the frequency with which images of each class appeared in the training dataset.

[0058] In this example, training was performed using an Nvidia Quadro P6000 GPU in a Dell desktop. Other GPUs, such as the Nvidia TITAN Xp, are also suitable. Additionally, training can be performed on a traditional CPU. In general, CPUs are better suited for smaller training tasks.

[0059] As mentioned above, k-fold cross-validation is a suitable framework for training the constituent CNNs of ensemble networks. The training data (i.e., the data on which the CNN is trained) includes at least retinal scans and a representation of the pathology present in the retina in each scan (e.g., a representation of genetic abnormalities).

[0060] FIG. 4 is an exemplary illustration of k-fold cross-validation for training an ensemble network with five CNNs for three image modalities (i.e., five-fold cross-validation yields 15 CNNs).

[0061] In this example, the training process described above was repeated for each fold (split group) of each modality, and the trained network weights were saved at the end of each run (so early stopping was not used), resulting in 15 sets of network weights.

[0062] Patients in the development set were divided into five folds of approximately equal size as follows: 1. A table of counts for each fold-modality-gene combination was created and initialized to zero. 2. For each patient, we identified the genes associated with that patient and the number of images for each modality. 3. Patients were randomly shuffled (to ensure reproducibility). 4. For each patient in turn, we selected the fold with the lowest count across all modalities for that patient's genes in the table and assigned that patient to that fold. The count in the table was incremented by the number of images for that patient for each modality.

[0063] To facilitate training and evaluation of individual networks, we split the dataset (in CSV format) into a set of training and testing CSVs based on individual folds (thus, for each modality, there are files [....]_train_0.csv, [...]_test_0.csv, [...]_train_1.csv, [...]_test_1.csv, etc.), where [...]_train_N.csv contains folds 0-4 excluding N, and [...]_test_N.csv contains folds N and -1.

[0064] To ensure a reasonable amount of training data for each gene, we removed genes that did not have at least five images across all folds / modalities and removed all images / patients associated with those genes from the training and test sets, resulting in a list of 36 genes that were used throughout all experiments.

[0065] In the example shown, the system is trained using retinal scans taken from patients with IRD at Moorfields Eye Hospital (MEH) who were seen at MEH and had undergone genetic testing and had a genetic cause confirmed by an accredited diagnostic laboratory.

[0066] The MEH IRD cohort was previously reported by Pontikos, N. et al. ("Genetic basis of inherited retinal disease in a molecularly characterized cohort of over 3,000 families from the United Kingdom," Ophthalmology (2020)) and included 4,236 individuals with IRD caused by mutations in 135 different genes, of which 452 (with mutations in 66 genes) were under 18 years of age as of August 2, 2019. Patients with IRD and confirmed genetic diagnoses by a certified genetic testing laboratory were identified, and information on genetic diagnosis, age at onset, and mode of inheritance was exported from the MEH electronic medical record (OpenEyes) using SQL queries of the hospital data warehouse database in Microsoft SQL Server.

[0067] Images were exported from the MEH Heidelberg Imaging (Heyex) database (Heidelberg Engineering, Heidelberg, Germany) for all patients with IRD based on hospital number for records between 2006-06-05 and 2018-04-05, resulting in a dataset of 1,196,038 images from 2,871 IRD patients with disease-causing mutations in 132 different genes.

[0068] Table 1: Number of patients and images per gene and modality (FAF, IR, SD-OCT) for the 36 genes included in the example development dataset. [Table 1]

[0069] It should be noted that the complete dataset for this example was subjected to quality control (see section below). In this way, the dataset was restricted to only genes with at least five images per gene in each of the three modalities (after quality control), leaving 36 individual genes. The distribution of the 36 selected genes is shown in Table 1 above. For all 132 genes, some were excluded due to an insufficient number of patients with sufficient images to properly assign at least five images to each fold.

[0070] Note that the full dataset obtained from MEH in this example (after data quality control; see below) contains 15,692 FAF scans, 23,631 IR scans, and 13,099 SD-OCT scans from 2,171 patients across 36 different genes. This full dataset is split into a "development" set of 1,907 patients and a held-out internal test set of 264 patients. The internal test set allows for testing of the final ensemble network. The development set allows for the investigation of network-level properties and ensemble approaches.

[0071] As can be seen in Figure 4, training in this example is performed using a total of 44,817 retinal scans acquired from 1,907 patients with any IRD (3,749 eyes, 6,397 appointments) from MEH. These scans are split across three different modalities: FAF (N=13,509), IR (N=20,098), and SD-OCT (N=11,344). For each of the three modalities, five different neural networks were trained using 5-fold cross-validation on the development data, resulting in a total of 15 different neural networks.

[0072] The patients in the development set were divided into five patient subsets of approximately equal size, taking care to ensure an approximately equal split across each of the different genes in each of the three different modalities. For each of the three modalities, images from each of the five patient subsets were used as folds in a five-fold cross-validation setup. For each combination of the four patient subsets, a deep CNN was trained, with each patient subset excluded from the training set of exactly one network (images from that subset could be used as test data for that network). This not only diversified the constituent CNNs of the ensemble, but also allowed us to observe the variability of the network-level model (ensemble network) due to different training / test sets, providing insight into how sensitive the ensemble network is to the training dataset.

[0073] For each of the 15 datasets (three modalities, five folds per modality), a 36-class CNN was trained for 100 epochs (through the entire dataset), which proved sufficient. Figure 5 shows an example of the network's loss on the training set during training. As shown, preliminary studies on all images from this example dataset found that 100 epochs was sufficient for convergence for all hyperparameter settings.

[0074] To obtain a variety of hyperparameter values ​​for training each CNN, for each of the three modalities, the initial dataset can be further split into training and validation datasets according to a 75:25 split (resulting in an overall 60:20:20 training / validation / test split). Networks can then be trained on these new training sets using a variety of random hyperparameter settings. The training accuracy and average AUC scores (defined below) per gene on the validation set for the 10 random hyperparameter settings are recorded in Table 2 below.

[0075] Table 2: Hyperparameter (learning rate; batch size; dropout) settings tested, with the best results in each highlighted column. The hyperparameters used in Run 6 yielded the best results on the example dataset, overall and across the three image modalities. [Table 2]

[0076] As seen above, dropout in the penultimate layer of the network can be used with a variety of dropout probabilities. No dropout has been found that performs well across all three modalities. Therefore, in this example, a 0% dropout probability is used, meaning that there is essentially no dropout in the final CNN.

[0077] In cross-validation, in theory, hyperparameter tuning should ideally be applied independently to each fold to avoid biasing the final test results. In practice, however, this significantly increases computational effort, with negligible benefits. Therefore, in this example, the optimal hyperparameter settings are also used to train the remaining folds. This effect here only applies to the results on the internal development set, and does not affect the results on either the internal or external test sets.

[0078] For each image, a corresponding class label is given by the underlying patient's genetic diagnosis. The neural network uses the Inception-v3 architecture, initialized with pre-trained ImageNet weights, with the final output layer replaced by a linear layer with 36 outputs and softmax normalization. The Inception-v3 architecture was chosen because it offers the smallest model size among architectures with similar performance applied in other ophthalmology domains. Those skilled in the art will recognize that other architectures suitable for image classification are also suitable. The loss function selected in this example is the cross-entropy loss with class weights inversely proportional to the gene frequency in the dataset. The default parameters used in the Keras library (β1 = 0.9, β2 = 0.999) were selected for the Adam optimizer.

[0079] In summary, in this example, hyperparameter tuning is performed by random sampling over 10 trials with an 80 / 20 training / validation split in the first training fold. After tuning the hyperparameters, a batch size of 128, a learning rate of 0.0001, and a dropout probability of 0% were found to be optimal. These hyperparameters are used to train the network across all three modalities. Data augmentation techniques can also be applied to avoid overfitting the training data (see the following section).

[0080] Returning to Figure 4, we train a CNN for each imaging modality used during the acquisition of the dataset using 5-fold cross-validation.

[0081] Quality control and expansion of training data Prior to training, the dataset at hand may undergo preprocessing for quality control. It should be noted that this step is not always necessary and depends on the dataset in question. For the dataset mentioned above (from MEH's Heyex database), quality control has been performed. Figure 6 outlines an example of the quality control process. The table in Panel A shows the quality control steps and inclusion thresholds applied for the three imaging modalities in the example mentioned above.

[0082] Of the database queries, retinal scans were separated by modality: 51,376 FAF scans, 43,746 IR scans, and 1,095,082 SD-OCT scans (5,834 images from other modalities). Because SD-OCT generates multiple B-scans, only the images corresponding to the central B-scan across the fovea were used for each B-scan, as these are likely the most informative B-scans. This left 33,849 OCT B-scans.

[0083] Filtering is applied for all three modalities. Corrupted scans and scans with a size other than 768 × 768 pixels (for FAF and IR scans) or 512 × 496 pixels (for OCT scans) are discarded. Additionally or alternatively, scans may be resized. Both IR and FAF scans have two imaging magnifications: 30° and 55°. In each case, only the most common mode (55° for FAF and 30° for IR) is saved; all other scans are discarded. These two modes can be automatically distinguished by checking the number of black background pixels (RGB0,0,0), and we found that the total number of black pixels corresponding to the 55° image was 107,577.

[0084] A number of quantitative filters are applied to remove low-quality or defective images. The type of filter is determined by an initial inspection of the data to identify specific problems. The thresholds for these filters are set by examining a sample of the scans that have been filtered out and adjusting the threshold until at least 25% of the filtered scans are judged to be of an acceptable quality level.

[0085] In particular, the FAF and IR datasets contain a large number of scans that are excessively dark, so scans with median pixel intensities below 0.05 for FAF and 0.1 for IR are excluded.

[0086] Furthermore, many FAF scans are characterized by a large amount of noise. To measure this noise level, we introduced a "pixel noise level" score. This score compares the original scan with a blurred version of the same image (using a normalized box filter with a 5 × 5 kernel) and takes the sum of the squared differences between the two images. FAF scans with a sum of squared differences greater than 2200 are rejected. Many OCT images contained artifacts consisting of large regions with a pixel value of 1.0 (i.e., the maximum value in this case). To remove these, all images with a maximum pixel value of 1.0 and a median intensity greater than 0.2 are excluded.

[0087] Finally, a BRISQUE score (Blind / Reference-less Image Spatial Quality Evaluator) is calculated for each scan using the PyBRISQUE library, and scans above a certain threshold are discarded, which in this example are set to 120 for FAF images, 80 for IR images, and 150 for OCT images.

[0088] After quality control, we restricted the dataset to only genes with at least five scans, leaving 15,692 FAFs, 23,631 IRs, and 13,099 OCT scans for the 36 most common genes, as seen in the bottom of panel B.

[0089] Those skilled in the art will appreciate that the specific pre-processing quality control procedures described above are data set dependent, with the exact parameters depending on the data set at hand.

[0090] Panel C is an example of a poor quality scan that was filtered out by the example filter above. Subpanel A is an FAF scan that is too dark to see any detail. Subpanel B is an FAF scan that is too noisy. Subpanel C is an infrared scan that is too dark to see any detail. Subpanel D is an SD-OCT scan where a large artifact obscures part of the scan.

[0091] As mentioned above, data augmentation may be used to avoid overfitting the training data: retinal scans may be augmented before training so that the augmented training data set contains the same number of retinal scans as the original training data set, or the augmented retinal scans are concatenated to the original training data set.

[0092] Figure 7 illustrates eight data augmentation techniques that can be applied to retinal scans. The techniques depicted include horizontal flipping, rotation, brightness adjustment, random zoom, Gaussian blur, adding Gaussian noise, adding salt and pepper noise, and adding speckle noise. As can be seen in the bottom row of images, techniques can be applied sequentially. For example, horizontally flipping the original image and applying a brightness factor of -1 results in the third image from the left. Any or all of the data augmentation techniques can be applied to retinal scans acquired from any imaging modality.

[0093] Performance evaluation - effectiveness on the cross-validation test set and left-out validation test set of examples Similar to the example above, the constituent neural nets of the ensemble network may be trained via 5-fold cross-validation, where each individual network is trained on a different subset of the training data.

[0094] In the above example, for a network trained on FAF images, we achieved an average Top-1 accuracy of 44.3% (CI 95% We observed a mean ROC AUC of 0.827 (0.812-0.842), a top-5 accuracy of 72.7% (71.0-74.4), and a mean ROC AUC per gene of 0.827 (0.812-0.842). For networks trained on IR images, we observed an accuracy of 45.8% (43.6-47.9), a top-5 accuracy of 73.7% (72.2-75.2), and an AUC of 0.834 (0.824-0.844). Finally, for networks trained on SD-OCT images, we observed an accuracy of 51.6% (48.4-54.9), a top-5 accuracy of 77.4% (75.2-79.6), and an AUC of 0.845 (0.831-0.859).

[0095] Results for the excluded patient set were similar, with average accuracy between models of 51.4% for FAF, 46.9% for IR, and 56.3% for the SD-OCT model (see Table 4 below, especially the Single Model Average, Single Image section (top left)).

[0096] More specifically, to evaluate the effectiveness of the ensemble network approach, we simulate the scenario of applying the approach to scans taken during a patient visit by applying it to scans from the internal leave-out dataset (validation set) introduced above (see Figure 4). The Moorfields leave-out dataset is used for internal testing. As mentioned above, the MEH internal testing dataset in this example contains 1,900 scans from 264 patients across 32 genetic diagnoses. The breakdown of scans by modality is shown in the first row of Table 3 below.

[0097] Table 3: Number of patients, scans, and genes in the internal and external test datasets [Table 3]

[0098] Scans in the internal test dataset are grouped by patient and acquisition date. Predictions are compared to each patient's underlying genetic diagnosis. An ensemble network approach using only scans from each patient's first visit achieved a Top-1 accuracy of 66.7%, a Top-5 accuracy of 85.6%, and an average receiver operator characteristic AUC per gene of 0.935 (Table 4; see the 5-model ensemble, multi-image / modality section (bottom right)).

[0099] The ensemble network approach is applied across all available scans per patient per appointment (defined as all scans from a given patient on a given day), and at the per-patient level (using all available scans per patient across all appointments).

[0100] Table 4: Summary of test accuracy for the excluded patient set using different combination methods of model prediction. Per visit means all eye scans per patient visit. Per patient means all eye scans across multiple visits for a patient. [Table 4]

[0101] For each modality-specific ensemble model, we apply the model to all scans for that modality from the internal test set. Model predictions are then compared to each patient's underlying genetic diagnosis to calculate the model's overall accuracy on the test data, top-k accuracy (the percentage of images for which the correct gene is within the network's top k predictions) for k=2, 3, 5, and 10, and the average area under the receiver operator characteristic curve (AUROC) per class.

[0102] The ROC (Receiver Operator Characteristic) indicates the trade-off between specificity (false positive rate) and sensitivity (true positive rate) for a gene. For each gene, the AUC (Area Under the Curve) of the model's corresponding ROC curve is obtained by comparing the model predictions for a given gene in one-versus-all setups.

[0103] The confidence intervals for accuracy and AUC in the development data are obtained by taking the standard deviation of the accuracy / AUC for each of the three modalities across the five networks and estimating the 95% confidence interval by the formula:

[0104] Performance Evaluation - Ensemble Approach vs. Others By decomposing the ensemble network into its constituent networks, we can see that the overall accuracy across each modality is relatively similar. Figure 8 shows a collection of per-gene receiver operating characteristic (ROC) curves from cross-validation data for three different imaging modalities across six of the 36 predictor genes in the example above. The FAF classification is depicted as a dashed-dotted line, the IR classification as a dashed line, and the OCT classification as a solid line. Each curve is the average of the ROC curves for five models per modality. The percentage next to the gene name indicates the percentage of all images (across all three modalities) for the given gene. The shaded area indicates the area within one standard error of the mean. For simplicity in this example, only six genes (ABCA4, RPGR, CNGB3, TIMP3, CRB1, and PRPH2) from the 36 classified genes are depicted here.

[0105] We also find that while the overall accuracy of each modality is relatively similar, there are notable differences for specific genes. For example, the FAF and SD-OCT networks are significantly better at detecting TIMP3 than the IR network. This is notable because TIMP3 retinopathy is associated with increased peripheral macular signal in FAF images and drusen-like deposits in SD-OCT scans. Similarly, the SD-OCT network is better at detecting CRB1 than the IR and FAF networks. CRB1 retinopathy is characterized by a thickened retina with poorly defined layers, which cannot be distinguished in FAF or IR images.

[0106] Many genes follow a similar pattern to PRPH2, with little deviation between modalities. However, for some genes, there are significant differences in the ability of the networks for different modalities to discriminate between them. For example, for CRB1, a gene associated with retinal thickening, SD-OCT is the most predictive modality.

[0107] In this way, an ensemble approach to retinal scan classification can identify the optimal imaging modality at a per-lesion (pathology) level (e.g., per-gene level).

[0108] We found that combining predictions across multiple models (ensembles) and across multiple images acquired during one or more patient visits contributes to the performance of the ensemble network.

[0109] When applying an ensemble of five models per modality to individual images, we observed accuracies of 59.2%, 55.3%, and 61.5% for FAF, IR, and OCT, top-5 accuracies of 81.3%, 80.8%, and 82.9%, and AUCs of 0.881, 0.885, and 0.911, respectively (see Table 4; 5-model ensemble, single-image section (bottom left)). This approach yields improved results compared to the single-model average, single-image approach.

[0110] When combining predictions from individual models across multiple images at the per-visit level (without ensembling five models per modality), we observe an average overall model accuracy of 62.8%, a top-5 accuracy of 84.6%, and an average AUC of 0.914 on the leave-out test set in the example above (Table 4; see single-model ensemble, multiple images / modalities section (top right), per-visit column).

[0111] Figure 9 shows a schematic of the four models benchmarked here. Panel A shows a single-image model, which takes a single image from one eye as input and outputs a single classification probability from a single network. Panel B shows a single-image ensemble model (five networks for each modality), which takes a single image from one eye as input and outputs classification probabilities from each of the five networks in the ensemble, which combine these into a single classification probability. Panel C shows a multi-image model (one network for each modality), which takes multiple images from different modalities (i.e., FAF, OCT, and IR) as input and outputs classification probabilities for each image, which combine these into a single classification probability for the patient. Panel D shows a multi-image ensemble (five networks for each modality), which takes multiple images from different modalities (i.e., FAF, OCT, and IR) as input and outputs classification probabilities for each image from each of the five networks, which combine these into a single classification probability for the patient.

[0112] Table 5 below summarizes the accuracy of classifying multiple scans using a single convolutional neural network for each imaging modality compared to the accuracy of classifying multiple scans using a five-model ensemble network for each imaging modality. Note that the single-image accuracy results (left column) are the same as the single-image accuracy results shown in Table 4 (left column) and are included here for ease of reference.

[0113] Table 5: Summary of test accuracy for the excluded patient set using single scan v. multiple scans [Table 5]

[0114] Based on per-gene ROC AUC, we found that combining all images across all modalities outperformed the best-performing modality (i.e., limited to images from that modality) for the majority of genes, demonstrating the superiority of the multi-modality approach. Figure 10 shows a collection of per-gene ROC curves from cross-validated data for the combined prediction (solid black curve; "E2G ensemble") across all images from a given patient appointment, superimposed on the multi-image ROC curve limited to images from a given modality. The percentage next to the gene name indicates the percentage across appointments corresponding to that gene. For simplicity, this example shows only six of the 36 ROC curves classifying all genes (the same six as in Figure 8). This again demonstrates the benefits of the multi-modality approach.

[0115] In both the ensemble of multiple models per modality applied to individual images, and the combination of predictions of individual models across multiple images, these cases provide results that are better than those of single networks, but worse than the overall ensemble model approach taken in the embodiments, demonstrating the advantages of both employing ensembles and combining predictions across multiple images.

[0116] Performance evaluation - robustness To confirm the generalizability of the results across hospitals, we applied the classifier to data from four IRD clinics, including Oxford Eye Hospital (UK), Liverpool University Hospital (UK), Bonn University Hospital (Germany), and Federal University of São Paulo (Brazil). As shown in Table 3, Oxford Eye Hospital (UK) provided 346 scan samples from 59 patients with different genetic diagnoses in 29 different genes. Liverpool University Eye Hospital (UK) provided 210 scan samples from 35 patients with different genetic diagnoses in 15 different genes. Bonn University Eye Hospital (Germany) provided 473 scan samples from 127 patients with different genetic diagnoses in 11 different genes. Federal University of São Paulo (Brazil) provided 104 scan samples from 15 patients with different genetic diagnoses in three different genes.

[0117] Patients and scans were selected by clinicians at these different clinics, who were instructed to select patients with a confirmed genetic diagnosis who had the number of scans available for each modality. The selection of scan images was shared with us along with each patient's genetic diagnosis, except for data from the University of Bonn, where the implementation was performed locally to avoid transferring patient data.

[0118] We ran the ensemble network on each patient's image and compared the ensemble network's predictions with the underlying genetic diagnosis. Combining all data from all four sites (1133 retinal scans from 236 patients), we achieved an overall accuracy of 65.3% and a top-5 accuracy of 86.4% (a breakdown of results by site is shown in Table 6 below).

[0119] Because the distribution of genes in the external test data differs from the main MEH dataset (and therefore the internal leave-out test set), and the accuracy of the ensemble network in this example varies by gene, we also recorded the per-gene prevalence (percentage of the dataset) and sensitivity (accuracy for that particular gene) for each gene at each external site. To understand whether the differences in accuracy in the external data are due to differences in gene distribution, we calculated a reweighted accuracy on the internal test data for each site to match the gene distribution of the external dataset of interest. As a result, we found that, assuming constant per-gene performance, the overall accuracy on the external data would be expected to be 67.2%, but still slightly higher than the actual figure of 65.3%. This pattern was consistent across all four sites, with actual accuracy being several percentage points lower than predicted from the internal test data.

[0120] In summary, ensemble networks trained according to embodiments are robust when faced with previously unseen datasets and generalizable across sites.

[0121] Table 6: Summary of ensemble network results on external data across different sites. Gene-by-gene extrapolation from the internal dataset is the sum of gene-by-gene sensitivity on the MEH internal test dataset multiplied by the gene frequency in the external dataset, representing the expected accuracy for the gene distribution at that site. [Table 6]

[0122] Performance Evaluation - Related to Human Resource Specialists To illustrate the performance of this example embodiment compared to human experts, we asked nine ophthalmologists with varying levels of experience to predict causative genes based on single FAF images across 50 different patients (selected to ensure uniform representation of all genes in the dataset) from the leave-one-out test set of the example described above. The ophthalmologists achieved an average accuracy of 17% on this task, compared to 38% for an ensemble network trained on the same images. We also asked four ophthalmologists to select patients from an external test set at Bonn University Hospital and predict causative genes from the 11 genes in the dataset. They achieved an average accuracy of 42% on this task, compared to 49% for the trained ensemble network (which still had to select from 36 genes).

[0123] More specifically, the results of the human benchmark performed by ophthalmologists are shown in Table 7 below. In this task, ophthalmologists were shown FAF scans of IRD patients and asked to identify the correct diagnostic genes. As expected, human performance tended to improve with experience level. For a fair comparison, we also applied examples of ensemble networks trained on the same dataset, but limited to single-image predictions (i.e., using only the five-model FAF ensemble). The performance of the trained ensemble networks was generally superior to that of a single human expert (except for ophthalmologist 11, who achieved 60.3% accuracy), and the frequency with which correct genes appeared in the top five predictions of the trained ensemble networks was comparable to their frequency in all human guesses.

[0124] Thus, classification of retinal scans using examples from the trained ensemble network is comparable to that of experts.

[0125] Table 7: Trained ensemble networks by ophthalmologist and embodiment for comparison classified the internal leave-out test dataset containing 50 FAF retinal scans of 50 patients from Moorfields Eye Hospital and the Bonn external test dataset containing 73 FAF retinal scans of 37 patients. Note: Ophthalmologists 8 and 9 performed the test together. [Table 7]

[0126] Performance Evaluation - Ensemble Network Prediction Calibration In the example of an ensemble network, predictions are output with a probability for each gene, but for any given image, each underlying network typically predicts a single gene class with confidence, making it difficult to ascertain the reliability of the model from the output of only a single network.

[0127] However, we found that when the predictions of the constituent networks were combined in an ensemble network, the output predictions appeared relatively well-calibrated, except for a consistent 10% overprediction. For example, if the ensemble network predicted a gene with a 70% probability, the prediction was correct approximately 60% of the time. This is illustrated in Figure 11, which shows the actual accuracy compared to the model confidence of the ensemble network for each individual image, along with a best-fit line (slope = 1.04, intercept = -11.2%). The ensemble predictions are slightly overconfident compared to a "well-calibrated" 1:1 correspondence with the actual accuracy (shown by the dotted line). The total percentage of datasets exceeding each confidence threshold is also shown (dashed lines).

[0128] This is useful for providing accurate and informative feedback to the end user regarding the reliability of the model. In particular, this aspect can be used in the implementation of embodiments in a web app (see below) to identify unclear input scans where it may be beneficial to re-image. This can be implemented by flagging individual images that fall below a certain confidence threshold when run through the model.

[0129] Refining ensemble network output to include age and genetic pattern In addition to retinal scans, additional information, such as the age at which the patient presented to the clinic and family history, can help identify eye diseases, e.g., causative genes. Because such information is frequently available along with patient scans, methods for classifying retinal scans using ensemble networks can be modified to incorporate these additional features. In particular, variations on the embodiment allow for combining the patient's age at first visit and mode of inheritance (MOI), which can be inferred from family history, with the output of the ensemble network.

[0130] One way to incorporate this information is to use a linear classifier, for example, implementing logistic regression without regularization. For the example above, we found that training a linear classifier (using logistic regression without regularization) that predicts gene class based on the output of the ensemble network, age, and MOI improved overall accuracy from 66.7% to 68.9% and Top-5 accuracy from 85.9% to 89.4%. In this example, the average AUROC per gene did not improve and remained unchanged at 0.935. See Table 8 below for a detailed breakdown of the effectiveness of combining approaches.

[0131] Table 8: Results of different classification approaches on the internal test dataset [Table 8]

[0132] More specifically, age and mode of inheritance (MOI) are important factors to consider in the genetic diagnosis of IRD, providing strong prior evidence for a specific gene. For example, if the patient is young and the MOI is X-linked, RP2 is suggested as the associated gene. Figure 12 shows the distribution of age at onset (first visit) for each genetic diagnosis for the examples described throughout this specification. The horizontal axis provides the age of the subject at which each genetic diagnosis (genetic abnormality as provided on the vertical axis) was first presented. For example, abnormalities in the CACNA1F gene appear earliest in childhood, most strongly from around 1 to 12 years of age, while abnormalities in the MTTL1 gene appear earliest in the second half of life, most strongly in the late 40s to early 50s. Figure 13 shows the known mode of inheritance for each gene for the same example. For example, it is known that the MOI of abnormalities in the MTTL1 gene is likely mitochondrial. Some genes, such as BEST1 and PROML1, are inherited dominantly, while others are inherited recessively.

[0133] To incorporate these non-image features, a model stacking approach may be adopted, where a linear classifier may be trained with combined predictions from the ensemble network and additional input features.

[0134] In this example, age is treated as a normalized numeric variable (subtract the mean and divide by the standard deviation), and MOI is treated as a one-hot vector with five classes (recessive, dominant, X-linked, mitochondrial, and unknown). For example, a recessive gene can be encoded by a 1x5 row vector with the value 10000. For each image / appointment from the set of five patients in this example, we concatenate these features with the final output of the corresponding network (where each gene prediction is treated as an individual feature) and use these, along with the true gene labels, to fit a logistic regression classifier (without regularization).

[0135] These feature vectors are concatenated with the ensemble network predictions obtained from cross-validation experiments (i.e., using only the test results of each network) to form a 41-wide feature vector, on which a 36-class linear classifier is fitted using logistic regression (targeting patient genes as before). This trained classifier is then applied to an internal test set, using the concatenated age+MOI features and ensemble network predictions as before, and the output is compared to the underlying genes for each patient.

[0136] To test the resulting combined model, we applied this trained model to an internal test set using ensemble predictions from the full ensemble network concatenated with age and MOI, and used the resulting predictions as new gene-level predictions, conducting the same top-k and gene-wise average AUC analyses as in the other performance evaluations. We also performed cross-validation analysis based on five folds of the development set, using the model output as training data in four folds and testing on the remaining fold, repeating this process for each fold.

[0137] Figure 14 shows a collection of ROC curves for linear classifiers trained on different feature combinations (age, genetic pattern, and model output from a trained ensemble network ("Eye2Gene")) using cross-validation data (top) and an internal test dataset (bottom). For the cross-validation data, a linear model is trained on four folds and tested on the remaining folds. The ROC curves shown are the average of the five folds, with the shaded region representing one standard error away from the mean. For the internal test dataset, a linear model is trained on the development set and tested on the internal test dataset. In both cases, we see that adding more information to the output of the trained ensemble network improves performance.

[0138] Those skilled in the art will understand that other patient characteristics can be integrated with the classifier to further refine the ensemble network output. For example, the patient's ethnic background can be additionally or alternatively encoded and classified. Furthermore, those skilled in the art will understand that classifiers implementing other algorithms, such as logistic regression with regularization, can be used.

[0139] Packaged retinal scan image classification method The retinal scan image classification methods described herein can be used to assist clinicians in diagnosing patients with IRD. The ensemble network can be deployed, for example, as an online web app, i.e., the trained ensemble network can be accessed as an online app.

[0140] Figure 15 shows an example of a suitable screen for a graphical user interface (GUI) that provides example results to the user after running the ensemble network. This screen (and any other screen) can be presented to the user on a handheld device or computer screen. To use the example app, the user can upload a series of scans (e.g., as PNG images) and enter basic case information, such as age at presentation and MOI, if known. In this example, the age is 29 years old, and the MOI is "recessive." The user is then asked to specify the type of scan (in this example, SD-OCT, 55° FAF, or 30° IR) for each image. These images are then passed to a trained ensemble network (trained CNNs for one or all of the relevant imaging modalities), which outputs a set of prediction scores for each of the 36 genes for each of the input scans. This information is then aggregated (or ensembled) into an overall prediction score for the case and presented to the user.

[0141] In this example, a bar graph of the top five genes predicted by the ensemble network is displayed, along with each gene's model probability score. These predictions are broken down into the contributions of three different modalities, displayed as different colors or shades on the bar graph. Input images color-coded by modality (here, FAF scans are shown as solid lines, OCT scans as long-dashed lines, and IR scans as short-dashed lines) and patient information are shown at the top of the display. A complete breakdown, including predicted probabilities for all 36 genes, is shown in the table below the bar graph. Links in the top right of the table take you to a breakdown of the ensemble network's predictions for each uploaded scan.

[0142] Applicability to other eye diseases The above example provides a method and network for retinal scan image classification, classifying retinal scans based on genetic abnormalities. Different image modalities may highlight different aspects of retinal disease by providing different "views," and thus provide different complementary evidence toward accurate disease diagnosis. Therefore, analyzing multiple image modalities in parallel can provide a deeper understanding of disease pathology than examining them one by one.

[0143] Thus, the multimodal approach disclosed herein is applicable to all ocular diseases (genetic and non-genetic), including retinal diseases, age-related macular degeneration (AMD), diabetic retinopathy (DR), glaucoma, retinopathy of prematurity (ROP) and vein occlusion, diseases of the optic nerve such as glaucoma and optic neuropathies, and diseases affecting the cornea such as keratoconus, corneal dystrophies and keratitis.

[0144] Consider AMD as an example. Blue autofluorescence imaging modality can highlight lipofuscin distribution in the retinal pigment epithelium over a fairly wide field of view (50 degrees). Lipofuscin accumulation is a by-product of intracellular aging, making it an important biomarker of retinal aging, such as AMD.

[0145] Optical coherence tomography (OCT) imaging modality highlights different layers in cross-sections of the retina and can therefore detect abnormalities such as the accumulation of retinal fluid as seen in AMD.

[0146] Infrared imaging highlights abnormalities in the retinal vasculature and may therefore be able to identify abnormal blood vessel formation such as that which occurs in AMD.

[0147] This can be complemented by OCT angiographic imaging, which may highlight abnormal vascularization in various layers of the retina.

[0148] Depending on the location and stage of the lesion, the above imaging modalities can be combined with wider-angle imaging to capture changes in the peripheral retina.

[0149] Thus, for AMD retinal scan classification, the method according to the invention can be employed with respect to four (or five) imaging modalities to generate an ensemble probability for the presence of AMD in the retinal scan.

[0150] Naturally, a similar multimodal approach is also used to identify cases of glaucoma, which can be detected in both the cornea, thalamus, and peripheral retina, and can be detected by thinning of the retinal nerve fiber layer with OCT.

[0151] Hardware 16 is a block diagram of a computing device, such as a data storage server, that embodies the present invention and can be used to perform aspects of the methods for retinal scan image classification described herein. The computing device comprises a processor 993 and a memory 994. Optionally, the computing device also includes a network interface 997 for communicating with other computing devices.

[0152] For example, an embodiment may consist of a network of such computing devices. Optionally, the computing devices also include one or more input mechanisms, such as a keyboard and a mouse 996, and one or more display devices, such as one or more monitors 995. These components may be connected to one another via a bus 992.

[0153] The memory 994 may include a computer-readable medium, which term may refer to a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) configured to carry computer-executable instructions or have data structures stored thereon. Computer-executable instructions may include, for example, instructions and data that can be accessed by a general-purpose computer, a special-purpose computer, or a special-purpose processing device (e.g., one or more processors) to cause it to perform one or more functions or operations. Accordingly, the term “computer-readable storage medium” may also include any medium capable of storing, encoding, or carrying a set of instructions for execution by a machine, causing the machine to perform any one or more of the methods disclosed herein. Accordingly, the term “computer-readable storage medium” includes, but is not limited to, solid-state memory, optical media, and magnetic media. By way of example and not limitation, such computer-readable media may include non-transitory computer-readable storage media including random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory devices (e.g., solid-state memory devices).

[0154] The processor 993 is configured to control the computing device and perform processing operations, for example, executing code stored in the memory 994 to perform various different functions of the retinal scan image classification method or retinal scan image classification learning process(es) as described and claimed herein.

[0155] The memory 994 may store data read from and written to the processor 993, such as data from a learning or classification task performed on the processor 993. For example, the memory 994 may store information regarding the selected architecture(s) of each CNN and store weights for implementation on the architecture(s).

[0156] As referred to herein, processor 993 may include one or more general-purpose processing devices, such as a microprocessor, a central processing unit, a GPU, etc. Processor 993 may also include a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or a combination of instruction sets. Processor 993 may also include one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. In one or more embodiments, processor 993 is configured to execute instructions to perform the operations and steps described herein.

[0157] The network interface (network I / F) 997 may be connected to a network such as the Internet, and may be connected to other computing devices via the network. The network I / F 997 may control the input and output of data to and from other devices via the network.

[0158] Methods embodying aspects of the present invention may be performed on a computing device such as that illustrated in Figure 16. Such a computing device need not have all of the components illustrated in Figure 16, but may be configured with a subset of those components. Methods embodying aspects of the present invention may be performed by a single computing device in communication with one or more data storage servers over a network, or by multiple computing devices operating in cooperation with each other. Cloud services implementing the computing devices may also be deployed.

Claims

1. 1. A computer-implemented method for classifying retinal scan images, the method including: receiving an input dataset including at least one retinal scan, wherein the retinal scan is acquired using one of a plurality of imaging modalities; passing the input dataset through an ensemble network, wherein the ensemble network comprises a plurality of convolutional neural networks, Each convolutional neural network is configured to classify retinal scans acquired using one of the plurality of imaging modalities based on a plurality of ocular pathologies, each convolutional neural network generating a probability of presence of each of the ocular pathologies in the retinal scan; and The retinal scan is processed using a convolutional neural network corresponding to the imaging modality used to acquire the retinal scan; and Generating an ensemble probability of the presence of each of the ocular pathologies in the retinal scan.

2. 2. The method of claim 1, further comprising: the input dataset includes a plurality of retinal scans, each retinal scan acquired using one of the plurality of imaging modalities; wherein each retinal scan is processed using a convolutional neural network corresponding to the imaging modality used to acquire the retinal scan.

3. 3. A method for classifying retinal scans according to claim 1 or 2, comprising: The multiple imaging modalities include any or all of the following methods: fundus autofluorescence (FAF); infrared (IR); and spectral domain optical coherence tomography (SD-OCT).

4. 4. A method for classifying retinal scans according to any one of claims 1 to 3, comprising: The method, wherein the plurality of ocular pathologies are genetic abnormalities.

5. 5. The method of claim 4, further comprising: Passing the ensembled probabilities through a linear classifier, wherein the linear classifier is configured to refine the ensembled probabilities based on the subject's age and / or the mode of inheritance (MOI) of the genetic abnormality.

6. 6. A method for classifying retinal scans according to any one of claims 1 to 5, comprising: the ensemble network comprises k convolutional neural networks for each imaging modality; the retinal scan is processed using k convolutional neural networks, each convolutional neural network corresponding to the imaging modality used to acquire the retinal scan; The method further comprises: Averaging the probability of presence of each of the ocular pathologies in the retinal scan from each of the k convolutional neural networks.

7. 7. The method of claim 1, further comprising: Outputting a prompt to reacquire the retinal scan if the estimate of the presence of each of the ocular pathologies in the retinal scan falls below a predefined confidence threshold.

8. 8. The method of claim 1, further comprising: Processing the output dataset to provide a classification of the genetic cause of the ocular pathology.

9. An ensemble network for retinal scan image classification, in particular for use in a method according to any one of claims 1 to 8, said ensemble network comprising a plurality of convolutional neural networks, wherein: Each convolutional neural network is configured to classify retinal scans acquired using one of a plurality of imaging modalities based on a plurality of ocular pathologies, and each convolutional neural network generates a probability of the presence of each of the ocular pathologies in the retinal scan.

10. A computer-implemented method for training an ensemble network for classification of retinal scan images, in particular as claimed in claim 9, comprising: receiving an input training dataset including training retinal scans labeled with acquisition imaging modalities and known ocular pathologies, where each training retinal scan is acquired using one of a plurality of imaging modalities; dividing the training retinal scans into a plurality of sets of training retinal scans, wherein each set of training retinal scans includes training retinal scans acquired using the same imaging modality; training a convolutional neural network to minimize an error between a predicted classification of a plurality of ocular pathologies and the known ocular pathologies for each set of training retinal scans; and For each convolutional neural network, output the trained convolutional neural network weights.

11. 11. A method for training an ensemble network for retinal scan classification according to claim 10, comprising: Training the convolutional neural networks uses k-fold cross-validation, and the method includes training k convolutional neural networks for each set of training retinal scans.

12. 12. A method of training an ensemble network for retinal scan classification according to claim 10 or 11, comprising: The method further includes augmenting the input training dataset by any or all of rotating, flipping, zooming, cropping, adjusting brightness, blurring, and adding noise.

13. 13. A method of training an ensemble network for retinal scan classification according to any one of claims 10 to 12, comprising: The method further includes preprocessing the input training data set by any or all of filtering by median pixel intensity, filtering by noise level, filtering by identifying image artifacts, and filtering by BRISQUE score.

14. A data processing apparatus comprising a memory and a processor configured to perform the method of any one of claims 1 to 8 and 10 to 13.

15. A computer program comprising instructions that, when said program is executed by a computer, cause said computer to carry out the method of any one of claims 1 to 8 and 10 to 13.

16. A computer readable medium comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8 and 10 to 13.