System for imaging and diagnosing retinal diseases

The integration of MSI and OCT through machine learning models addresses the limitations of existing imaging technologies by combining spectral and structural information for enhanced ocular disease diagnosis and biomarker identification.

JP2025532035APending Publication Date: 2025-09-29ALCON INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025515569
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-27
Filing Date
2023-09-27
Publication Date
2025-09-29

AI Technical Summary

Technical Problem

Existing imaging technologies struggle to effectively combine the rich spectral information from multispectral imaging (MSI) with the structural detail of optical coherence tomography (OCT) for early diagnosis of ocular diseases, requiring expertise to interpret OCT images and lacking comprehensive biomarker identification.

Method used

A system integrating MSI and OCT using machine learning models for feature extraction and biomarker identification, combining MSI's spectral depth with OCT's structural information to enhance disease diagnosis, utilizing neural networks for image processing and segmentation.

Benefits of technology

Enables early disease detection and severity assessment by leveraging the strengths of both MSI and OCT, providing comprehensive biomarker segmentation and diagnosis with improved accuracy and reduced reliance on human expertise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532035000001_ABST
    Figure 2025532035000001_ABST
Patent Text Reader

Abstract

In certain embodiments, a system, computer-implemented method, and computer-readable medium are disclosed for performing integrated analysis of MSI and OCT images to diagnose ocular disease. The MSI and OCT images are processed using separate input machine learning models to create input feature maps that are input to intermediate machine learning models. The intermediate machine learning models process the input feature maps and output final feature maps that are processed by one or more output machine learning models that output one or more estimated representations of pathologies in the patient's eye. A single device captures the OCT and non-OCT images using (a) a sensor in a first imaging device that is shared with a second imaging device, and / or (b) optical components for directing light that forms a retina to the first imaging device or the second imaging device.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Multispectral imaging (MSI) is a technique that involves measuring (or capturing) light from a specimen (e.g., ocular tissues / structures) at different wavelengths or spectral bands across the electromagnetic spectrum. MSI has the potential to capture more information from a specimen that may be invisible with conventional imaging, which generally uses broadband illumination and broadband imaging sensors. MSI information acquired by an MSI imaging system may be used to diagnose ocular diseases and enable real-time adjustments in the use of instruments (e.g., forceps, lasers, probes, etc.) used to manipulate ocular tissues / structures during surgery. [Background technology]

[0002] Optical coherence tomography (OCT) is a technology that uses light waves to generate two-dimensional (2D) and three-dimensional (3D) images of the eye. 2D OCT may involve the use of time-domain OCT and / or Fourier-domain OCT, the latter of which involves the use of spectral-domain OCT and swept-source OCT techniques. 3D OCT may similarly utilize time-domain OCT and Fourier-domain OCT imaging techniques. OCT imaging may also be used preoperatively or intraoperatively to diagnose ocular diseases.

[0003] It would be an advancement in the art to better utilize the capabilities of MSI and OCT to diagnose ocular diseases. Summary of the Invention [Means for solving the problem]

[0004] In certain embodiments, a system is provided, including a first imaging device configured to capture a first image of a patient's retina, where the first image is an optical coherence tomography (OCT) image. The system further includes a second imaging device configured to capture a second image of the patient's retina according to an imaging modality other than OCT. In the system, at least one of (a) a sensor of the first imaging device is shared with the second imaging device, and (b) an optical component is configured to select whether the first imaging device or the second imaging device receives light reflected from the retina to the first imaging device or the second imaging device.

[0005] So that the above-mentioned features of the present disclosure can be understood in detail, a more particular description of the present disclosure briefly summarized above can be had by reference to embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only exemplary embodiments and therefore should not be considered as limiting the scope of the invention, as other equally effective embodiments are possible. [Brief explanation of the drawings]

[0006] [Figure 1] 1 illustrates an exemplary system for performing integrated analysis of MSI and OCT images to diagnose ocular disease, according to certain embodiments. [Figure 2A] FIG. 1 illustrates a first approach for training a machine learning model to perform joint analysis of MSI and OCT images to diagnose ocular diseases, according to certain embodiments. [Figure 2B] FIG. 1 illustrates a second approach for training a machine learning model to perform joint analysis of MSI and OCT images to diagnose ocular diseases, according to certain embodiments. [Figure 2C] FIG. 1 illustrates a third approach for training a machine learning model to perform joint analysis of MSI and OCT images to diagnose ocular diseases, according to certain embodiments. [Figure 3]FIG. 1 is a flow diagram of a method for training a machine learning model to perform joint analysis of MSI and OCT images to diagnose ocular diseases, according to certain embodiments. [Figure 4A] 1 illustrates a system for capturing both OCT and MSI images and spectral information, according to certain embodiments. [Figure 4B] 1 illustrates an alternative system for capturing both OCT and MSI images and spectral information, according to certain embodiments. [Figure 4C] FIG. 1 is a more detailed diagram of a system for capturing both OCT and MSI images and spectral information, according to certain embodiments. [Figure 5A-5B] FIG. 1 illustrates a system for diagnosing eye disease using machine learning models, according to certain embodiments. [Figure 6] 1 illustrates an exemplary computing device that at least partially implements one or more functions for performing integrated analysis of images from multiple imaging modalities, according to certain embodiments.

[0007] For ease of understanding, where possible, identical elements common to the figures are designated with the same reference numerals. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further reference. DETAILED DESCRIPTION OF THE INVENTION

[0008] Various embodiments described herein provide a framework for processing information obtained from MSI and OCT images using artificial intelligence. The advantage of MSI is that MSI images contain rich information about the retina in a wide spectral range, features that cannot be seen using human vision or a fundus camera. The wide spectral range of MSI also provides a high degree of depth penetration into the retina. However, MSI images do not provide structural information. In contrast, OCT images provide structural information about the retina. However, interpreting OCT images requires a high degree of expertise. Using the approach described herein, the rich detail and deep penetration of MSI can be combined with the structural information of OCT to identify biomarkers for various pathologies and perform early disease diagnosis.

[0009] 1 shows a system 100 for performing integrated analysis of MSI images 102 and OCT images 104. System 100 may include three main stages: a feature extraction stage using machine learning models 106a and 106b, a feature boosting stage using machine learning model 110, and a biomarker and prediction stage using machine learning models 114 and 116. Through these three stages, system 100 processes MSI images 102 and OCT images 104 separately for feature extraction and then combines the extracted features to obtain a meaningful interpretation.

[0010] The MSI image 102 may be captured using any approach for implementing MSI known in the art, including so-called hyperspectral imaging (HSI). Similarly, the OCT image 104 may be acquired using any approach for performing OCT known in the art.

[0011] The MSI images 102 are acquired by illuminating a patient's eye with a multispectral band illumination source (e.g., a narrow-band illumination source, a narrow-band filter, etc.) and / or measuring reflected light with a multispectral band camera (e.g., an imaging sensor capable of detecting multiple spectral bands beyond the RGB spectral bands). Thus, each MSI image 102 represents reflected light within a specific spectral band. Differences between MSI images 102 arise from the different reflectances of different structures within the eye for different spectral bands. Thus, when considered collectively, the MSI images 102 provide additional information about the ocular structures than a single broadband image. In some implementations, the MSI images 102 are en face images of the retina used to detect retinal pathologies. However, MSI images 102 of other parts of the eye, such as the vitreous or anterior chamber, may also be used.

[0012] Optical coherence tomography (OCT) is a technology that generates two-dimensional (2D) and three-dimensional (3D) images of the eye using light waves from a coherent light source, i.e., a laser. OCT images are typically cross-sectional images of the eye for planes parallel and collinear to the eye's optical axis. However, OCT images for multiple cross-sectional planes may be used to construct a 3D image, from which 2D images for cross-sectional planes not parallel to the optical axis may be generated. For example, an en face image of the retina may be derived from the 3D image. In some embodiments, the OCT image 104 is such an en face image of the retina. OCT is capable of imaging the retina to a certain depth, such that the OCT image 104 is, in some embodiments, a collection of en face images for image planes at or above the surface of the retina, to depths within or below the retina.

[0013] Although the embodiments described herein relate to the use of MSI images 102 and OCT images 104, images from any pair of imaging modalities, or images from three or more different imaging modalities, may be used in a similar manner. For example, the additional imaging modalities may include scanning laser ophthalmology (SLO), a fundus camera, and / or a broadband visible light camera.

[0014] In the system 100, the MSI image 102 is processed by a machine learning model 106a, and the OCT image 104 is processed by a machine learning model 106b. The machine learning models 106a, 106b may be implemented as a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of neural networks.

[0015] The results of processing the images 102, 104 by the machine learning models 106a, 106b are feature maps 108a, 108b, respectively. For example, the feature maps 108a, 108b may be the output of one or more hidden layers of the machine learning models 106a, 106b. The feature maps 108a, 108b may be two-dimensional or three-dimensional arrays of values. If the feature maps 108a, 108b are two-dimensional arrays, the feature maps 108a, 108b may include the same dimensions in both dimensions, or may differ. If one or both of the feature maps 108a, 108b are three-dimensional arrays, the feature maps 108a, 108b may include the same dimensions in at least two dimensions, or may differ in any of the three dimensions.

[0016] The feature maps 108a, 108b, and possibly the images 102, 104, are processed by a machine learning model 110. The machine learning model 110 may be implemented as a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other type of neural network. The result of processing the feature maps 108a, 108b, and possibly the images 102, 104, by the machine learning model 110 is the feature map 112. For example, the feature map 112 may be the output of one or more hidden layers of the machine learning model 110, as discussed in more detail below.

[0017] The feature map 112, and possibly the images 102, 104, are then processed by machine learning models 114 and 116, which then output one or more biometric segmentation maps 118 that label ocular features represented in the images 102, 104 that correspond to one or more lesions. Each biometric segmentation map 118 may be in the form of an image having the same dimensions as the images 102, 104, with the non-zero pixels corresponding to pixels in the images 102, 104 that are identified as corresponding to the particular lesion represented by the biometric segmentation map. The biometric segmentation maps 118 may include a separate map for each of the multiple lesions, or a single map in which all pixels representing any of the multiple lesions are non-zero.

[0018] The machine learning model 114 may be implemented as a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of neural networks. For example, the machine learning model 114 may be implemented as a U-net.

[0019] The machine learning model 116 outputs a disease diagnosis 120 and, optionally, a severity score 122 corresponding to the disease diagnosis. The machine learning model 116 may be implemented as a long short-term memory (LSTM) machine learning model, a generative adversarial network (GAN) machine learning model, or other types of machine learning models. The disease diagnosis 120 may be output in the form of text specifying the name of the lesion, a numeric code corresponding to the lesion, or other representation. The severity score 122 may be a numeric value, such as a value between 1 and 10 or a value within another range. The severity score 122 may be restricted to a discrete set of values ​​(e.g., integers between 1 and 10) or may be any value within the limits of precision for the number of bits used to represent the severity score 122.

[0020] The lesions for which a biometric segmentation map 118 may be generated and a diagnosis 120 and severity score 122 may be generated include at least the following: ·Retinal tear Retinal detachment ·Diabetic retinopathy Hypertensive retinopathy Sickle cell retinopathy Central retinal vein occlusion Epiretinal membrane ·Macular hole Macular degeneration (including age-related macular degeneration) Retinitis pigmentosa Glaucoma Alzheimer's disease Parkinson's disease, which at least causes noticeable changes to the retina.

[0021] The biological segmentation map 118 may, for example, mark vascular features that correspond to lesions. Examples of vascular features that can be used to diagnose lesions are described in the following references, both of which are incorporated herein by reference in their entirety: Segmenting Retinal Vessels Using a Shallow Segmentation Network to Aid Ophthalmic Analysis,M.Arsalan et al.,Mathematics 2022,Volume 10,p.1536. PVBM:A Python(R) Vasculature Biomarker Toolbox Based on Retinal Blood Vessel Segmentation, J. Fhima et al., Cornell University (31 July, 2022).

[0022] 2A illustrates an exemplary approach for training the machine learning models 106a, 106b, 110, 114, 116. In particular, FIG. 2A illustrates a supervised machine learning approach that uses a plurality of training data entries 200, such as hundreds, thousands, tens of thousands, hundreds of thousands, or more. Each training data entry 200 may include as input an MSI image 102 and an OCT image 104. Each of the MSI images 102 represents an image obtained by detecting light in a different spectral band relative to the other MSI images 102.

[0023] The MSI image 102 and OCT image 104 of the training data entry 200 may be of the same patient eye and may be captured approximately simultaneously so that the anatomical structures represented in the images 102, 104 are approximately the same. For example, "approximately simultaneously" may mean within one second and one hour of each other. However, "approximately simultaneously" may depend on the pathology being detected, and those with very slow progression may use images 102, 104 captured with a longer time difference, such as less than one day, less than one week, or some other time difference. The MSI image 102 and OCT image 104 are preferably aligned and scaled with respect to each other so that a given pixel coordinate in the MSI image 102 represents approximately the same location in the eye (e.g., within 0.1 mm, 1 μm, or 0.01 μm) as the same pixel coordinate in the OCT image 104. This alignment and scaling may be accomplished for the entire images 102, 104, or for at least a portion of one or both of the images 102, 104 that shows the anatomical structure of interest (eg, the macula of the retina).

[0024] Alignment and scaling of the images 102, 104 relative to one another may be achieved by aligning the optical axes of the instruments used to capture the images 102, 104 and calibrating the magnification of the instruments to achieve approximately identical magnifications (e.g., within ±0.1%, 0.01%, or 0.001%). Alternatively, alignment and scaling of the images 102, 104 may be achieved by analyzing the anatomical structures represented in the images 102, 104. For example, if the MSI image 102 and the OCT image 104 represent the retina of the eye, the pattern of blood vessels represented in each image 102, 104 may be used to align and scale one or both of the images 102, 104. If the images 102, 104 are different sizes or do not completely overlap, such as after registration and scaling, the non-overlapping portions of one or both of the images 102, 104 may be cropped and / or one or both of the images 102, 104 may be padded so that the images 102, 104 are the same size and completely overlap each other.

[0025] Each training data entry 200 may include as desired output some or all of one or more biomarker segmentation maps 118, disease diagnoses 120, and severity scores 122. The same patient may have multiple lesions such that a segmentation map 118, disease diagnosis 120, and severity score 122 may be included for each lesion present or a subset of the most predominant lesions. The desired output is generated by a human expert based on evaluation of the images 102, 104, and possibly other health information about the patient obtained before and after the capture of the images 102, 104. In particular, the biomarker segmentation map 118 for a lesion may include pixels in one or both of the images 102, 104 that are marked by the human expert as corresponding to the lesion.

[0026] For each training data entry 200, the machine learning model 106a receives the MSI image 102 and generates one or more estimated biomarker segmentation maps. For example, the output of the machine learning model 106a may be a three-dimensional array of estimated biomarker segmentation maps, where each two-dimensional array along the third dimension corresponds to a lesion.

[0027] The training algorithm 202 compares the one or more estimated biomarker segmentation maps to one or more biomarker segmentation maps 118 for the training data entries 200. The training algorithm 202 then updates one or more parameters of the machine learning model 106a according to the difference between each estimated biomarker segmentation map for the lesion and the corresponding biomarker segmentation map 118 for that lesion in the training data entries 200.

[0028] The machine learning model 106b may be trained in a similar manner to the machine learning model 106a. For each training data entry 200, the machine learning model 106b receives the OCT image 104 and generates one or more estimated biomarker segmentation maps. For example, the output of the machine learning model 106a may be a three-dimensional array of estimated biomarker segmentation maps, where each two-dimensional array along the third dimension corresponds to a lesion.

[0029] A training algorithm 202, which may be the same or different from the one used to train the machine learning model 106a, compares the one or more estimated biomarker segmentation maps to one or more biomarker segmentation maps 118 of the training data entry 200. The training algorithm 202 then updates one or more parameters of the machine learning model 106b according to the difference between each estimated biomarker segmentation map for the lesion and the corresponding biomarker segmentation map 118 for that lesion in the training data entry 200.

[0030] If more than two imaging modalities are used, additional machine learning models may be present and trained in a similar manner. Each machine learning model corresponds to an imaging modality and processes the corresponding images for that imaging modality in the training data entries. The machine learning models generate one or more estimated biomarker segmentation maps that are compared by the training algorithm to one or more biomarker segmentation maps 118 of the training data entries, and the training algorithm then updates the machine learning models according to the comparison. The hidden layer of each machine learning model may generate an output that is used as a feature map for the imaging modality to which the machine learning model corresponds.

[0031] 1, the machine learning model 110 takes as input the feature maps 108a, 108b of the machine learning models 106a, 106b. The machine learning model 110 may be trained after the machine learning models 106a, 106b have been trained using some or all of the training data entries 200.

[0032] For each training data entry 200, the machine learning model 110 receives feature maps 108a, 108b resulting from processing the MSI image 102 and OCT image 104 of the training data entry 200 by the machine learning models 106a, 106b. As noted above, the feature maps 108a, 108b may be the output of a hidden layer of the machine learning models 106a, 106b, respectively, i.e., a layer other than the final layer that outputs one or more estimated biomarker segmentation maps. The machine learning model 110 may also receive the MSI image 102 and the OCT image 104 as inputs, although in other embodiments, only the feature maps 108a, 108b are used.

[0033] The machine learning model 110 processes the feature maps 108a, 108b, and optionally the MSI image 102 and the OCT image 104, to generate one or more estimated biomarker segmentation maps. For example, the output of the machine learning model 110 may be a three-dimensional array of estimated biomarker segmentation maps, with each two-dimensional array along the third dimension corresponding to a lesion. While two feature maps 108a, 108b are shown, the machine learning model 110 may process any number of feature maps, and optionally any number of images used to generate the feature maps, in a similar manner for any number of imaging modalities.

[0034] The training algorithm 204 compares the one or more estimated biomarker segmentation maps to one or more biomarker segmentation maps 118 of the training data entries 200. The training algorithm 204 then updates one or more parameters of the machine learning model 110 according to the difference between each estimated biomarker segmentation map for the lesion and the corresponding biomarker segmentation map 118 for that lesion.

[0035] 1, the machine learning models 114, 116 take as input the feature maps 112 of the machine learning model 110. The machine learning models 114, 116 may be trained after the machine learning model 110 has been trained using some or all of the training data entries 200.

[0036] For each training data entry 200, the machine learning models 114, 116 receive a feature map 112 resulting from processing the MSI image 102 and OCT image 104 of the training data entry 200 by the machine learning models 106a, 106b, 110. As noted above, the feature map 112 may be the output of a hidden layer of the machine learning model 110, i.e., a layer other than the final layer that outputs one or more putative biomarker segmentation maps. The machine learning models 114, 116 also take the MSI image 102 and the OCT image 104 as inputs, although in other embodiments, only the feature map 112 is used.

[0037] The machine learning model 114 processes the feature maps 112, and possibly the images 102, 104 from the training data entry 200, to generate one or more estimated biomarker segmentation maps. If more than two imaging modalities are used, the images according to the more than two imaging modalities from the training data entry 200 may be processed by the machine learning model 114 along with the feature maps 112 obtained from the images. The output of the machine learning model 114 may be a three-dimensional array of estimated biomarker segmentation maps, where each two-dimensional array along the third dimension corresponds to a lesion.

[0038] The training algorithm 206a compares the one or more estimated biomarker segmentation maps to one or more biomarker segmentation maps 118 of the training data entries 200. The training algorithm 206a then updates one or more parameters of the machine learning model 114 according to the difference between each estimated biomarker segmentation map for the lesion and the corresponding biomarker segmentation map 118 for that lesion.

[0039] The machine learning model 116 processes the feature maps 112, and optionally the MSI images 102 and OCT images 104, to generate one or more probable diagnoses and probable severity scores for each probable diagnosis. If more than two imaging modalities are used, the images according to the more than two imaging modalities from the training data entry 200 may be processed by the machine learning model 116 along with the feature maps 112 obtained for the images.

[0040] The output of the machine learning model 116 may be a vector, where each element of the vector, if non-zero, indicates that a lesion is predicted to be present. The output of the machine learning model 116 may also be text listing one or more prevalent lesions predicted to be present. The output of the machine learning model 116 may further include a severity score for each lesion predicted to be present, such as a vector where each element corresponds to a lesion and the value of the element indicates the severity of the corresponding lesion.

[0041] The training algorithm 206b compares the putative diagnosis and corresponding severity score to the disease diagnosis 120 and severity score 122 of the training data entry 200. The training algorithm 206b then updates one or more parameters of the machine learning model 116 according to the difference between the putative diagnosis and corresponding severity score and the disease diagnosis 120 and severity score 122 of the training data entry 200.

[0042] 2B, in some embodiments, training of one or both of the machine learning models 106a, 106b may be performed by unsupervised training algorithms 210a, 210b, respectively. FIGURE 2B illustrates training with images 102, 104, with the understanding that one or more machine learning models for additional or alternative imaging modalities may be trained in the same manner.

[0043] For the embodiment of Figure 2B, the machine learning models 110, 114, 116 may be as described above with respect to Figure 2A. In some embodiments, only one of the machine learning models 106a, 106b is trained using an unsupervised training algorithm 210a, 210b, while the other is trained using a supervised training algorithm 202 as described above with respect to Figure 2A. No labeled training data entries are used for the unsupervised machine learning algorithms 210a, 210b.

[0044] The machine learning model 106a may be trained using a corpus of sets of MSI images 102. The corpus may be curated to include a large set of MSI images of healthy eyes, e.g., retinal images, where no pathology is present, and a small portion of the corpus, e.g., less than 5% or less than 1%, that corresponds to one or more pathologies. The set of MSI images 102 may or may not be labeled with respect to whether the set of images 102 represents a pathology and / or the specific pathology represented.

[0045] The unsupervised training algorithm 210a processes the corpus with the machine learning model 106a and trains the machine learning model 106a to identify and classify anomalies detected in the set of MSI images 102 of the corpus. The unsupervised training algorithm 210a may be implemented using any approach for performing anomaly detection or other unsupervised machine learning known in the art. The output of the machine learning model 106a may be an image having the same dimensions as the individual MSI images 102 with pixels representing anomalies labeled.

[0046] The machine learning model 106b may be trained using a corpus of OCT images 104. The corpus may be curated to include a large number of OCT images of healthy eyes, e.g., retinal images, where no pathology is present, and a small portion of the corpus, e.g., less than 5% or less than 1%, that corresponds to one or more pathologies. The set of OCT images 104 may or may not be labeled with respect to whether the set of images 102 represents a pathology and / or the specific pathology represented.

[0047] The unsupervised training algorithm 210b processes the corpus with the machine learning model 106b and trains the machine learning model 106a to identify and classify anomalies detected in the OCT images 104 of the corpus. The unsupervised training algorithm 210b may be implemented using any approach for performing anomaly detection or other unsupervised machine learning known in the art. The output of the machine learning model 106b may be an image having the same dimensions as each OCT image 104 with pixels representing anomalies labeled.

[0048] The set of MSI images 102 and OCT images 104 used to train the machine learning models by the unsupervised training algorithms 210a, 210b may include images 102, 104 from the training data entry 200 used to train other machine learning models 110, 114, 116. The set of MSI images 102 and OCT images 104 may be further augmented with images of healthy eyes to facilitate identification of abnormalities corresponding to pathologies. The set of MSI images 102 and OCT images 104 may be constrained to be the same size and may be aligned with each other. For example, the images 102, 104 may be from different eyes, but may be aligned to place a representation of the center of the retinal fovea in the approximate center of each image 102, 104, e.g., within 1, 2, or 3 pixels. Other features, such as the fundus, may also be used for alignment. Although the images 102, 104 are of different eyes, the images 102, 104 may also be scaled so that the anatomical structures depicted in the images are approximately the same size. For example, the images 102, 104 may be scaled so that the fovea, fundus, or one or more other anatomical features are the same size.

[0049] Once trained, the machine learning models 106a, 106b may provide output to the machine learning model 110 in the form of one or both of the feature maps 108a, 108b, which are the output of one or more hidden layers of the machine learning models 106a, 106b, respectively (see FIGS. 1 and 2A). Alternatively, the final output of the machine learning models 106a, 106b, e.g., images with anomaly labels, may be used as input to the machine learning model 110.

[0050] 2C , in an improvement over the unsupervised machine learning approach of FIG. 2B , a supervised training algorithm 212b may compare the output of machine learning model 106a with the output of machine learning model 106b for a given set of MSI images 102 and OCT images 104 of the same patient's eye captured substantially simultaneously, as defined above. The supervised training algorithm 212b may then adjust the parameters of machine learning model 106b according to the comparison in order to train machine learning model 106b to identify the same abnormalities detected by machine learning model 106a. Note that the opposite approach may alternatively or additionally be used. The output of machine learning model 106b may be used by supervised training algorithm 212b or a different supervised training algorithm 212a to train machine learning model 106a to identify the abnormalities identified by machine learning model 106b.

[0051] In some embodiments, training may proceed in various phases, each using one of the training approaches described above with respect to FIGS. 2A, 2B, and 2C. In a first example, machine learning models 106a and 106b may first be trained using the supervised machine learning approach of FIG. 2A, and then machine learning models 106a and 106b may be trained using the unsupervised approach of FIG. 2B, with machine learning model 106b then further trained based on the output of machine learning model 106a (and / or vice versa) according to the approach of FIG. 2C. In a second example, only unsupervised learning is used. Machine learning models 106a and 106b are trained individually using the unsupervised approach of FIG. 2B, and then machine learning model 106b is trained based on the output of machine learning model 106a and / or trained based on the output of machine learning model 106b according to the approach of FIG. 2C.

[0052] 2C illustrates training with images 102, 104, with the understanding that one or more machine learning models for additional or alternative imaging modalities can be trained in the same manner. In particular, the output of a machine learning model according to one imaging modality may be used to similarly train one or more other machine learning models according to one or more other imaging modalities. The outputs of two or more first machine learning models for one or more first imaging modalities may be concatenated or otherwise combined and used to train one or more second machine learning models for one or more second imaging modalities using the approach of FIG. 2C.

[0053] 3, the illustrated method 300 may be performed by a computer system, such as computing system 600 of FIG. 6. Method 300 includes training a first input machine learning model with training images of a first imaging modality, at step 302. For example, step 302 may include training machine learning model 106a with MCI images 102 according to any of the approaches described above with respect to FIGS. 2A-2C.

[0054] Method 300 includes training a second input machine learning model with training images of a second imaging modality at step 304. For example, step 304 may include training machine learning model 106b with OCT images 104 according to any of the approaches described above with respect to FIGS.

[0055] At step 306, the method 300 includes processing an image according to a first imaging modality with a first input machine learning model to obtain an input feature map F1 and processing the image according to a second imaging modality with a second input machine learning model to obtain an input feature map F2. The feature maps F1 and F2 may be outputs of hidden layers of the first and second input machine learning models, respectively. Step 304 may include processing the MSI image 102 with the machine learning model 106a and processing the OCT image 104 with the machine learning model 106b to obtain feature maps 108a and 108b, as described above with respect to FIGS. 1 and 2A . As mentioned above, the MSI image 102 and the OCT image 104 may be part of a common training data entry 200 such that the MSI image 102 and the OCT image 104 are of the same patient's eye and are captured at approximately the same time.

[0056] The method 300 includes, at step 308, training an intermediate machine learning model using the feature maps F1 and F2. Specifically, multiple pairs of feature maps F1 and F2 may each be processed by the intermediate machine learning model, and the output of the intermediate machine learning model may be used to train the intermediate machine learning model. Each pair of feature maps F1 and F2 may be obtained for first and second modality images of the same patient's eye that are captured substantially simultaneously. Step 308 may include processing the images used to obtain each pair of feature maps F1 and F2 using the intermediate machine learning model. Step 308 may include training the machine learning model 110 using the feature maps 108a, 108b and the training data entries 200, as described above with respect to FIG. 2A .

[0057] Method 300 includes, at step 310, processing the feature pairs of feature maps F1 and F2, and optionally the training images used to obtain each pair of feature maps F1 and F2, with an intermediate machine learning model to obtain a final feature map F. The final feature map F may be obtained from the output of a hidden layer of the intermediate machine learning model. Step 310 may include processing feature maps 108a, 108b, and optionally corresponding images 102, 104, with machine learning model 110 to obtain feature map 112, as described above with respect to FIGS. 1 and 2A .

[0058] The method 300 includes, at step 312, training one or more output machine learning models with the feature map F. The one or more output machine learning models may be trained to, for a given feature map F, output estimated representations of lesions represented in the training images used to generate the feature map F using the first and second input machine learning models and intermediate machine learning models. The one or more output machine learning models may take as input images according to the first and second imaging modalities used to generate the feature map F. Step 312 may include training one or both of the machine learning models 114, 116 using the feature map 112 and optionally the corresponding images 102, 104 to output some or all of the biomarker segmentation map 118, the disease diagnosis 120, and the severity score 122.

[0059] Method 300 may include, at step 314, processing utilized images according to the first and second imaging modalities according to a pipeline of first and second input machine learning models, intermediate machine learning models, and one or more output machine learning models. Specifically, one or more utilized images according to the first imaging modality are processed using the first input machine learning model to obtain a feature map F1, one or more utilized images according to the second imaging modality are processed using the second input machine learning model to obtain a feature map F2, feature maps F1 and F2 and optionally utilized images are processed using the intermediate machine learning model to obtain a feature map F, and feature map F and optionally utilized images are processed by one or more output machine learning models to obtain an estimated representation of a lesion represented in the utilized image. The estimated representation may be output to a display device or stored in a storage device for later use or subsequent processing. The feature maps (F1, F2, F) may additionally be displayed or stored.

[0060] For example, step 314 may include processing the utilized images 102, 104 that are not part of the training data entry 200, i.e., the images 102, 104, with the machine learning models 106a, 106b, respectively, to obtain feature maps 108a, 108b, respectively, as described above with respect to Figure 1. The feature maps 108a, 108b, and possibly the utilized images 102, 104, may be processed with the machine learning model 110 to obtain a feature map 112. The feature map 112, and possibly the utilized images 102, 104, may be processed by one or both of the machine learning models 114, 116 to obtain a biomarker segmentation map 118, a disease diagnosis 120, and a severity score 122.

[0061] Steps 302-314 may be performed sequentially, i.e., the first and second input machine learning models are trained, followed by the intermediate machine learning models, followed by the training of one or more output machine learning models, followed by utilization. Steps 302-314 may additionally or alternatively be interleaved, i.e., the first and second input machine learning models, the intermediate machine learning models, and the one or more output machine learning models are trained as a group. For example, in a first phase, the first and second input machine learning models, the intermediate machine learning models, and the one or more output machine learning models are trained separately in the order listed, and in a second phase, training continues as a group, i.e., after an iteration involving processing a set of images according to the pipeline, some or all of the first and second input machine learning models, the intermediate machine learning models, and the one or more output machine learning models may be updated as part of the iteration by a training algorithm according to the outputs of the first and second input machine learning models, the intermediate machine learning models, and the one or more output machine learning models, respectively. During utilization step 314, training may continue individually or as a group, particularly unsupervised learning as described with respect to FIGS. 2B and / or 2C .

[0062] Step 314 may be performed by a different computer system than that used to perform steps 302-312. For example, the pipeline including the first, second, third, and one or more output machine learning models may be installed on one or more other computer systems for use by a surgeon or other medical professional.

[0063] Although the method 300 is described with respect to two imaging modalities, three or more imaging modalities may be used as well. For example, the imaging modalities IM i , i=1 to N, where N is assumed to be 2 or greater. For a given training data entry, the requirement of approximately identical magnification, alignment, and simultaneous imaging of the same eye of a patient is met by imaging modality IM. i The input may be filled by an image of the machine learning model ML i, i=1 to N, and each machine learning model ML i is the imaging modality IM i , and each corresponds to the corresponding imaging modality IM i By processing one or more images of i For each input, a machine learning model ML i according to any of the approaches described above to train the machine learning models 106a, 106b, and i may be trained with images of

[0064] The intermediate machine learning model in such an embodiment is therefore composed of N feature maps F i ,i=1~N as input, and optionally a feature map F i The output machine learning model takes the training images used to generate the final feature map F, and possibly the feature map F i The intermediate and output machine learning models are trained as described above for machine learning model 110 and machine learning models 114, 116.

[0065] 4A, 4B, and 4C illustrate exemplary systems 400a, 400b, and 400c that may be used to substantially simultaneously capture images from multiple imaging modalities and detect some or substantially all known retinal pathologies. For example, diabetic retinopathy is becoming increasingly common. Age-related macular degeneration (AMD) and glaucoma are also relatively common among older adults. Older adults will typically experience at least some vision loss due to one or more retinal pathologies. Because many retinal diseases are progressive, early detection is important to improve quality of life and reduce blindness.

[0066] Generally, ophthalmology settings have a variety of tools, including separate instruments such as OCT, fundus cameras, and scanning laser ophthalmoscopy (SLO). In the early stages of retinal diseases, it is difficult to distinguish between diseases. Images from multiple imaging modalities can help distinguish between diseases.

[0067] OCT images provide high-resolution structural information on all layers of the retina at the micron scale. However, structural changes are unlikely to occur in the early stages of disease. The metabolic state is usually abnormal before any structural changes become detectable by OCT.

[0068] Color fundus cameras provide color fundus images, but do not cover the 500-600 nm fluorescence range, which is important for detecting fluorophores deposited in the retinal pigment epithelium (RPE), which are present in the early stages of AMD and may develop into drusen or atrophy.

[0069] Fundus autofluorescence (FAF) allows imaging of retinal degeneration and focal pigmentation and hyperpigmentation indicators at the RPE level. FAF is the most reliable imaging modality for detecting, describing, quantifying, and monitoring the progression of outer retinal atrophy. BlinD (basal linear deposits) and BLamD (basal lamellar deposits) in the RPE are precursors of AMD and can be visualized using FAF but not OCT.

[0070] Currently, no single tool can reliably provide early detection, risk assessment, and progression monitoring of retinal disease. Figures 4A, 4B, and 4C show exemplary systems 400a, 400b, and 400c that can provide this functionality. Images acquired using systems 400a, 400b, and 400c may be processed using machine learning models to further facilitate early detection, risk assessment, and progression monitoring of retinal disease.

[0071] With particular reference to FIG. 4A , the system 400 a includes an OCT 402. The OCT 402 may be implemented as any OCT known in the art, which may include a light source 404, output optics 406, and a detector 408. The light source 404 may be a coherent light source, such as a laser or laser diode. The light source 404 may also be a low-coherence broadband light source. The output optics 406 may include a scanning mirror, focusing optics, and a mechanism for converting the depth of focus of the OCT 402, such as time-domain OCT. The detector 408 may include a spectrometer, such as a diffraction grating and charge-coupled device (CCD), a complementary metal-oxide semiconductor (CMOS) sensor, or other detector. The output of the detector 408 is either an image or a stream of analytes, which may be organized into an image based on the state of the OCT 402 when each analyte is detected.

[0072] Light from the light source 404 is directed by output optics 406 onto a retina 412 of a patient's eye 414. The light from the light source 404 may pass through one or more lenses 416 to focus the light onto the retina 412. Light reflected from the retina 412 returns to the output optics 406, at least a portion of which is directed to a detector 408.

[0073] An actuated mirror 418, such as a toggle switch mirror, can be positioned as shown or can be actuated to position 420. When OCT 402 is in use, actuated mirror 418 is placed in position 420, allowing light reflected from retina 412 to reach OCT 402 without interacting with mirror 418. Mirror 418 may also be positioned as shown in FIG. 4A to capture images according to one or more other imaging modalities. For example, the reflective surface of mirror 418 may be oriented at an angle of approximately 45 degrees (e.g., ±2 degrees) relative to the optical axis of lens 416 and / or eye 414. Mirror 418 may be manually switched between the positions shown or may be coupled to an electrical actuator 418a.

[0074] The light source 422 may be used to illuminate the retina 412 when imaging according to one or more other imaging modalities. The one or more other imaging modalities may include some or all of MSI, HSI, fundus autofluorescence (FAF), FAF spectral, infrared, ultraviolet, or other imaging modalities. The light source 422 may include (a) a single light source suitable for one or more other imaging modalities, (b) a single light source that may operate in different ways (intensity and / or spectrum) for different imaging modalities, or (c) multiple light sources, each light source being used for a different imaging modality. The light source 422 may be implemented as a broadband light source embodied as one or more LEDs.

[0075] Light from light source 422 may be incident on mirror 424, such as an annular aperture mirror, a beam splitter, or other mirror capable of partial reflection and transmission. Mirror 424 directs at least a portion of the light from light source 422 onto mirror 418, which directs the light onto retina 412, such as via lens 416.

[0076] Light reflected from retina 412 is directed by mirror 418 back to mirror 424, which allows at least a portion of the reflected light to pass through it. At least a portion of the reflected light may be directed to spectrometer 428. Spectrometer 428 captures spectral information of the reflected light.

[0077] A portion of the light reflected from the retina 412 may additionally or alternatively be reflected through a filter 430 onto a camera 432. The camera 432 may be a monochrome or color camera. The filter 430 may be a filter wheel including multiple filters, each corresponding to a different wavelength band. Multiple images of the reflected light may be captured, each image captured through a different one of the multiple filters interposed between the camera 432 and the retina 412. The multiple images may thus constitute an MSI or HSI image. The filter 430 may include an electronically controlled actuator for selecting among the multiple filters, or may be manually adjustable.

[0078] When both spectrometer 428 and camera 432 are used, beam splitter 426 may be positioned so that light reflected from retina 412 is both (a) transmitted through beam splitter 426 to one of camera 432 and spectrometer 428, and (b) reflected from beam splitter 426 to the other of spectrometer 428 and camera 432.

[0079] The OCT images output by detector 408, the FAF images and / or FAF spectral information acquired using spectrometer 428, and the MSI or HSI images acquired using camera 432 may be input to machine learning model 438. Machine learning model 438 is trained to do some or all of the following: (a) identify anatomical structures and features corresponding to retinal pathologies, (b) diagnose one or more retinal pathologies, and (c) estimate the severity of one or more retinal pathologies.

[0080] The machine learning model 438 may be embodied as the system 100 as described above, trained according to any of the embodiments described above. If more than two imaging modalities are implemented using the system 400a, the system 100 may be configured as described above to include three or more machine learning models 106a, 106b that generate three or more feature maps 108a, 108b that are input to the intermediate machine learning model 110. If more than two imaging modalities are implemented using the system 400a, the training data entries 200 used to train the machine learning models 106a, 106b, 110, 114, 116 may include images of the three or more imaging modalities in addition to or instead of the MSI images 102 and OCT images 104 described above. The machine learning models 106a, 106b, 110, 114, 116 may be trained in the same manner as described above, with each machine learning model 106a, 106b being trained with images of the imaging modality corresponding to that machine learning model 106a, 106b.

[0081] System 400a has many advantages when used in combination with system 100. System 400a can easily capture images of multiple imaging modalities substantially simultaneously without requiring the patient to move to a separate device. System 400a may also be configured so that multiple images according to multiple imaging modalities are substantially identically scaled (e.g., within 0.01% along each dimension) and substantially aligned, e.g., within 0.1 mm, 0.01 mm, or 1 μm, as measured with respect to retinal features represented in the multiple images. Substantially identical scaling may be obtained by calibrating the magnification of each imaging modality. Substantial alignment may be obtained by substantially aligning the optical axis of each imaging modality with the center of the image acquired using each imaging modality.

[0082] Images from multiple imaging modalities acquired using system 400a may be processed by machine learning models 438 upon capture. With readily available computing power, the results of the processing can be made available within minutes. An ophthalmologist, surgeon, or other medical professional can therefore provide an immediate diagnosis to a patient for virtually any retinal disease.

[0083] Referring to FIG. 4B, system 400b may be modified relative to system 400a by the use of an additional camera 436. Light reflected from retina 412 may be directed to both cameras 432, 436 by a beam splitter 434. Cameras 432, 436 may capture different types of images. For example, cameras 432, 436 may capture two or more of the following types of images: color (RGB), infrared, ultraviolet, MSI, HSI, and FAF. Images from camera 436 may be processed using machine learning model 438 along with other images according to other imaging modalities provided by system 400b, as described above with respect to FIG. 4A.

[0084] Referring to FIG. 4C, spectral domain OCT (SD-OCT) may be modified to implement the illustrated system 400c to capture images according to multiple other imaging modalities (MSI, HSI, FAF, color, infrared, ultraviolet).

[0085] The SD-OCT includes a light source 440. The light source 440 may be a low-coherence light source, such as a broadband light source that produces light across the entire visible spectrum (e.g., 380-700 nm), and possibly also in the infrared and / or ultraviolet spectrum. The light source 440 may be embodied as one or more light-emitting diodes (LEDs).

[0086] Light from a light source 440 is input to an optical fiber coupler 442. A portion of the light is transmitted to a dispersion compensator 444. The light passes through the dispersion compensator 444, strikes a mirror 446, and is reflected back through the dispersion compensator 444 to the optical fiber coupler 442.

[0087] Light from light source 440 is also coupled by fiber optic coupler 442 into optical fiber 448. Optical fiber 448 directs the light to lens 450, which directs the light received from optical fiber 448 onto scanning mirror 452. Scanning mirror 452 may be a single mirror actuated along two orthogonal axes of rotation to scan the light over a two-dimensional area of ​​retina 412. Alternatively, scanning mirror 452 may be embodied as two mirrors, each rotating about one of two orthogonal axes of rotation. For example, scanning mirror 452 may be embodied as a galvo mirror and a resonant scanner.

[0088] Light reflected from scan mirror 452 passes through scan lens 454. Scan lens 454 is actuated along its optical axis to change the focal depth of light from light source 440 that reaches retina 412. Scan lens 454 is therefore translated to different positions to image different layers of retina 412. Scan lens 454 may be actuated mechanically, or the focal depth may be changed electronically, such as by implementing scan lens 454 as an optofluidic lens.

[0089] The light emitted by the scan lens 454 may be directed onto the retina 412 by one or more other components. For example, one or more mirrors 456 may change the direction of the light emitted from the lens 454, and one or more lenses 458 may focus the light emitted from the lens 454 onto the retina 412. The position and / or orientation of the one or more mirrors 456 and the one or more lenses 458 may be adjustable.

[0090] In some embodiments, adaptive optics (AO) 460 may be positioned in the optical path between lens 450 and fiber optic coupler 442. Adaptive optics 460 improves the quality of images acquired using SD-OCT, and the characteristics of AO 460 may be selected according to any approach known in the art of SD-OCT design.

[0091] Light reflected from retina 412 follows the reverse path traversed by light traveling from fiber optic coupler 442 to retina 412. Upon reaching fiber optic coupler 442, at least a portion of the light reflected from retina 412, along with at least a portion of the light returning from dispersion compensator 444, is coupled into optical fiber 462. Optical fiber 462 receives the light output by optical fiber 462 and directs the light to spectrometer 464, such as by output lens 466, which focuses or collimates the light input to spectrometer 464.

[0092] In some embodiments, the optical fiber coupler 442 is coupled to the light source 440 and the dispersion compensator by a multimode optical fiber, and the optical fibers 462, 448 are single-mode optical fibers.

[0093] Spectrometer 464 may additionally be used for one or more other imaging modalities. Accordingly, a mirror 468, beam splitter, or other optical element may be used to direct light from multiple light sources to spectrometer 464. In the illustrated embodiment, mirror 468 is an actuated mirror. Mirror 468 may be mounted in the orientation shown to reflect light from output lens 466 to spectrometer 464. Mirror 468 may be moved in orientation 470 to allow light from one or more other light sources to enter spectrometer 464. Mirror 468 may be manually switched between the positions shown or may be coupled to an electrical actuator 468a.

[0094] Spectrometer 464 may be implemented using any type of spectrometer known in the art. In the illustrated embodiment, spectrometer 464 includes a diffraction grating 472, a lens 474, and a detector 476, such as a CCD or CMOS sensor. Light incident on diffraction grating 472 forms a wavelength-dependent fringe pattern that is focused by lens 474 onto detector 476. Thus, light incident on each point on detector 476 can be mapped to a specific wavelength. The output of detector 476 may then be processed to measure the spectrum of the light entering spectrometer 464. As light from light source 440 is scanned across retina 412, the reflectance spectrum of a spot on retina 412 may be obtained from each spectral measurement of spectrometer 464.

[0095] A mirror, beam splitter, or other optical element may be used to direct light onto the retina 412 for one or more other imaging modalities. For example, an actuated mirror 480, such as a toggle switch mirror, may be positionable as shown or actuated to position 482. When SD-OCT imaging is performed, the actuated mirror 480 is placed in position 482, allowing light to travel between the light source 440 and the retina 412 without interacting with the mirror 480. The mirror 480 may be positioned as shown in FIG. 4C to capture images according to one or more other imaging modalities. For example, the reflective surface of the mirror 480 may be oriented at an angle of approximately 45 degrees (e.g., ±2 degrees) relative to the optical axis of the lens 458 and / or the eye 414. The mirror 480 may be manually switched between the positions shown or may be coupled to an electrical actuator 480a.

[0096] Light source 484 may be used to illuminate retina 412 when imaging according to one or more other imaging modalities. The one or more other imaging modalities may include some or all of the following: MSI, HSI, FAF, color, infrared, ultraviolet, or other imaging modalities. Light source 484 may include (a) a single light source suitable for one or more other imaging modalities, (b) a single light source that may operate in different ways (intensity and / or spectrum) for different imaging modalities, or (c) multiple light sources, each light source being used for a different imaging modality. For example, light source 484 may be a broadband light source implemented using one or more LEDs.

[0097] Light from light source 484 may be incident on mirror 486, such as an annular aperture mirror, a beam splitter, or other mirror capable of partial reflection and transmission. One or more lenses 488 may be interposed between light source 484 and mirror 486 to focus or collimate the light from light source 484 onto retina 412. Mirror 486 directs at least a portion of the light from light source 484 onto mirror 480, which directs the light onto retina 412.

[0098] Light reflected from retina 412 is directed by mirror 480 back to mirror 486, which allows at least a portion of the reflected light to pass through it. A portion of the reflected light may be directed to spectrometer 464.

[0099] A portion of the light reflected from the retina 412 may additionally or alternatively be reflected through a filter 490 onto a camera 492. In some embodiments, one or more lenses 494 may be interposed between the filter 490 and the camera 492. The camera 492 may be a monochrome, color, infrared, ultraviolet, or other type of camera. The filter 490 may be a filter wheel including multiple filters, each corresponding to a different wavelength band. The filter wheel may be manually adjustable or may include an electronically controlled actuator. Multiple images of the reflected light may be captured, each captured through a different one of the multiple filters interposed between the camera 492 and the retina 412. The multiple images may thus constitute an MSI or HSI image. System 400c may be modified similarly to system 400b to use multiple cameras to capture light passing through the filter 490.

[0100] When both spectrometer 464 and camera 492 are used to image light from light source 484, beam splitter 496 may be positioned so that light reflected from retina 412 is both (a) transmitted through beam splitter 496 to one of camera 492 and spectrometer 464, and (b) reflected from beam splitter 496 to the other of spectrometer 464 and camera 492.

[0101] When system 400c is operating as an SD-OCT system, the output of spectrometer 464 is used to form an OCT image as is known in the art of SD-OCT. When mirror 480 is used to direct light from light source 484 onto retina 412, the output of spectrometer 464 may be used to form a FAF image, and the output of camera 492 may be used to form an MSI or HSI image. Other cameras used may produce FAF, color, infrared, ultraviolet, or other types of images. OCT, MSI or HSI, FAF, and other images acquired using system 400c may be processed using machine learning model 438 as described above.

[0102] 5A and 5B show additional machine learning models 500a, 500b that may be used as machine learning models 438 in addition to or instead of system 100.

[0103] 5A , the machine learning model 500a may take as input one or both of a FAF spectrum 502 and an oxygen map 504. The oxygen map measures the oxygen saturation level within the blood vessels of the retina 412. The oxygen map is obtained by analyzing the spectrum reflected from points within the retina 412. Thus, the oxygen map may be obtained using the output of the spectrometers 428, 464 using any approach known in the art for generating oxygen maps. The FAF spectrum 502 may similarly be obtained from the spectrometers 428, 464.

[0104] The FAF spectrum 502 and oxygen map 504 may each be input to a plurality of machine learning models 506a-506d, which generate corresponding outputs 508a-508d, respectively, that describe a particular metric, anatomical term, or feature corresponding to the lesion. In the illustrated embodiment, the machine learning models 506a-506d are convolutional neural networks (CNNs) 506a-506d. However, the machine learning models 506a-506d may also each be embodied as a deep neural network (DNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of machine learning models.

[0105] Outputs 508a-508d may include, for example, metabolism in the upper retinal vessels (e.g., anywhere between Bruch's membrane and the vitreous), metabolism in the choroidal vessels, metabolism in the RPE, phosphor deposits in the RPE, or other anatomical items or features corresponding to pathology that may be represented in the spectral information included in FAF spectrum 502 and / or oxygen map 504.

[0106] The outputs 508a-508d may be input to an output machine learning model 510. In the illustrated embodiment, the output machine learning model 510 is a random forest, although other machine learning models may be used, such as a long short-term memory (LSTM) machine learning model, a generative adversarial network (GAN) machine learning model, a logistic regression machine learning model, or other types of machine learning models.

[0107] The output machine learning model 510 outputs a metabolic state 512 of the retina 412 represented by the FAF spectrum and oxygen map 504. The metabolic state 512 may reflect information such as blood flow, oxygen saturation, waste removal, or other information. The metabolic state 512 may be in the form of a numerical value indicative of the overall metabolic state of the retina 412, with a low value indicating poor health and a high value indicating good health. The metabolic state 512 may be a plurality of numerical values, each corresponding to an aspect of the metabolism of the retina 412. The metabolic state may be stored and / or output to a display device. For example, the FAF spectrum 502, oxygen map 504, outputs 508a-508d, and some or all of the representations of the metabolic state may be stored and / or output to a display device for evaluation by an ophthalmologist, surgeon, or other medical professional.

[0108] Machine learning models 506a-506d and machine learning model 510 may be trained using training data entries. For example, each training data entry may include as inputs a FAF spectrum 502 and an oxygen map 504 of retina 412, and as desired outputs outputs 508a-508d as determined by human labelers and metabolic state 512 as determined by human labelers.

[0109] Thus, the training algorithm may process the FAF spectra 502 and oxygen maps 504 of the training data entries through the machine learning models 506a-506d to obtain estimated outputs 508a-508d that are compared with the outputs 508a-508d of the training data entries. The training algorithm then updates the parameters of the machine learning models 506a-506d according to the comparison.

[0110] An output machine learning model 510 may be trained by processing the outputs 508a-508d from the training data entries through the machine learning model 510 to obtain an estimated metabolic state 512. The training algorithm then compares the estimated metabolic state 512 with the metabolic state 512 from the training data entries and updates the output machine learning model 510 according to the comparison.

[0111] 5B, the machine learning model 500B takes as input one or more images 514 of the retina 412. The one or more images 514 may include an OCT image, an MSI image, an HSI image, an FAF image, a monochrome or color image, an infrared image, an ultraviolet image, or other image of the retina 412.

[0112] One or more images 514 may be input to multiple machine learning models 516, 516d, which each generate a corresponding output 518a-518d that describes a particular metric, anatomical term, or feature corresponding to the lesion. In the illustrated embodiment, the machine learning models 516a-516d are convolutional neural networks (CNNs) 516a-516d. However, the machine learning models 516a-516d may also each be embodied as a deep neural network (DNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of machine learning models.

[0113] The outputs 518a-518d may include, for example, exudates, hemorrhages, drusen, lesions, or other anatomical items or features corresponding to lesions that may be represented in some or all of the one or more images 514.

[0114] The outputs 518a-518d may be input to an output machine learning model 520. In the illustrated embodiment, the output machine learning model 520 is a random forest, although other machine learning models may be used, such as a long short-term memory (LSTM) machine learning model, a generative adversarial network (GAN) machine learning model, a logistic regression machine learning model, or other types of machine learning models.

[0115] The output machine learning model 520 outputs a diagnosis 522a-522c for each of the one or more lesions. The output machine learning model 520 may additionally output a stage 524a-524c for each of the one or more lesions indicating the progression level of the lesion. For example, the lesions may include diabetic retinopathy (DRP), age-related macular degeneration (AMD), glaucoma, or other lesions that may be represented in the image of the retina 412. The stage 524a-524c for each lesion may be one of a set of discrete values ​​for each lesion, such as a number from 1 to 10, or any other set of discrete values ​​used by medical professionals to represent the progression of the lesion.

[0116] Representations of the diagnoses 522a-522c and steps 524a-524c may be stored and / or output to a display device. Representations of some or all of the one or more images 514, outputs 518a-518d may be stored and / or output to a display device for evaluation by an ophthalmologist, surgeon, or other medical professional.

[0117] Machine learning models 516a-516d and machine learning model 520 may be trained using training data entries. For example, each training data entry may include one or more images 514 for retina 412 as input and outputs 518a-518d, diagnoses 522a-522c, and steps 524a-524c for diagnoses 522a-522c specified by a human labeler as desired outputs.

[0118] The training algorithm may process one or more images 514 of the training data entries through machine learning models 516a-516d to obtain estimated outputs 518a-518d that are compared with the outputs 518a-518d of the training data entries. The training algorithm then updates the parameters of the machine learning models 516a-516d according to the comparison.

[0119] Output machine learning model 520 may be trained by processing outputs 518a-518d from training data entries through output machine learning model 520 to obtain estimated diagnoses 522a-522c and stages 524a-524c. The training algorithm then compares estimated diagnoses 522a-522c and stages 524a-524c with diagnoses 522a-522c and stages 524a-524c from the training data entries and updates output machine learning model 520 according to the comparison.

[0120] The systems and methods disclosed herein provide at least the following advantages: A single device according to Figures 4A, 4B, or 4C may be used by an operator to acquire sufficient image and spectral information to identify virtually all retinal pathologies in a single patient position. The associated machine learning model can also output labeled images and diagnoses based on these images during the same visit in which they were captured. The associated machine learning model combines information from multiple imaging modalities to perform early detection of retinal pathologies. The convenience of being able to obtain images and diagnoses allows for earlier and more frequent testing, facilitating both earlier detection and more frequent testing to allow for observation of changes over time to assess progression of the lesion.

[0121] 6 illustrates an exemplary computing system 600 that may at least partially implement one or more of the functions described herein. An imaging device according to any of FIGS. 4A-4C may include a computing device having some or all of the attributes of computing system 600. The computing device may be coupled to and receive images from one or more cameras 432, 436, 492. The computing device may also receive the output of spectrometers 428, 464 and / or the output of OCT 402, where OCT 402 includes a detector 408 other than spectrometer 428. The computing device may control the operation of systems 400a, 400b, 400c to capture images and transition between imaging modalities as described above, including controlling some or all of the OCT 402, any light sources 422, 484, one or more mirror actuators (e.g., actuators 418a, 480a, 468a), the scanning mirror 452, one or more actuators of the scanning lens 454, and actuators of filters 430, 490, which may be embodied as a filter wheel or other device that allows selection from multiple filters.

[0122] 4A-4C may further execute machine learning model 438 and / or machine learning models 500a, 500b. Alternatively, a different computing device having some or all of the attributes of computing system 600 may receive image and spectral information from the imaging device and process the image and spectral information using machine learning model 438 and / or machine learning models 500a, 500b.

[0123] As shown, computing system 600 includes a central processing unit (CPU) 602, one or more I / O device interfaces 604 that can enable various I / O devices 614 (e.g., keyboard, display, mouse device, pen input, etc.) to be connected to computing system 600, a network interface 606 for connecting computing system 600 to a network 690, memory 608, storage 610, and an interconnect 612.

[0124] If the computing system 600 is an imaging system, the computing system 600 may further include one or more optical components for obtaining ophthalmic images of a patient's eye, and any other components known to those skilled in the art.

[0125] CPU 602 may retrieve and execute programming instructions stored in memory 608. Similarly, CPU 602 may retrieve and store application data residing in memory 608. Interconnect 612 transfers programming instructions and application data between CPU 602, I / O device interface 604, network interface 606, memory 608, and storage 610. CPU 602 is included to represent a single CPU, multiple CPUs, a CPU with multiple processing cores, etc.

[0126] The memory 608 represents volatile memory, such as random access memory, and / or nonvolatile memory, such as nonvolatile random access memory, phase change random access memory, etc. As shown, the memory 608 may store a training algorithm 616, such as any of the training algorithms 202, 204, 206a, 206b, 210a, 210b, 212a, and 212b described herein, or a training algorithm for training the machine learning models 500a and 500b as described above. The memory 608 may also store a machine learning model 618, such as any of the machine learning models 106a, 106b, 110, 114, 116, 438, 500a, and 500b described herein.

[0127] Storage 610 may be non-volatile memory such as a disk drive, solid-state drive, or a collection of storage devices distributed across multiple storage systems. Storage 610 may optionally store training data entries 620, such as training data entries 200 or training data entries for training machine learning models 500a, 500b as described above. Storage 610 may store images 622 acquired according to any of the imaging modalities described herein. Storage 610 may also store results of processing images with machine learning models 618, possibly including storing intermediate results of any of the machine learning models 618.

[0128] Note that the computing system 600 used to train the machine learning model 618 may be different from the computing system 600 that utilizes the trained machine learning model 618. Thus, the machine learning model 618 and the results 624 of the machine learning model 618 may exist without the corresponding training algorithm 616 and training data entries 620.

[0129] Additional considerations The above description is provided to enable those skilled in the art to practice the various embodiments described herein. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments. For example, changes may be made to the function and arrangement of the elements discussed without departing from the scope of the disclosure. In various embodiments, various steps or elements may be omitted, substituted, or added, as appropriate. Features described with respect to some embodiments may be combined in several other embodiments. For example, an apparatus may be implemented or a method may be practiced using any number of aspects described herein. Furthermore, the scope of the present disclosure is intended to cover similar apparatuses or methods that are implemented using structure, functionality, or structure and functionality in addition to or other than the various aspects of the disclosure described herein. It should be understood that any aspect of the disclosure disclosed herein may be implemented by one or more elements recited in the claims.

[0130] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single members. As an example, "at least one of a, b, or c" is intended to cover a, b, c, ab, ac, bc, and abc, as well as any combination with multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or a, b, and c in any other order).

[0131] As used herein, the term "determining" encompasses a variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, querying (e.g., querying a table, database, or other data structure), ascertaining, etc. "Determining" may also include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. "Determining" may also include resolving, selecting, choosing, establishing, etc.

[0132] The methods disclosed herein include one or more steps or actions for achieving the method. Method steps and / or actions may be interchangeable without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be changed without departing from the scope of the claims. Furthermore, various actions of the methods described above may be performed by any suitable means capable of performing the corresponding functions. These means may include various hardware and / or software elements and / or modules, including, but not limited to, circuits, application specific integrated circuits (ASICs), or processors. In general, wherever actions are illustrated in the figures, those actions may include corresponding means-plus-function elements that are similarly numbered.

[0133] The various illustrative logic blocks, modules, and circuits described in connection with this disclosure may be implemented or performed by a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware elements, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any commercially available processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in cooperation with a DSP core, or any other such configuration.

[0134] The processing system may be implemented with a bus architecture. The bus may include any number of interconnected buses and bridges, depending on the particular application and overall design constraints of the processing system. The bus may interconnect various circuits, including the processor, machine-readable media, and input / output devices, among others. A user interface (e.g., keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also connect various other circuits, such as timing sources, peripherals, voltage regulators, power management circuits, etc., which are known in the art and will not be described further. The processor may be implemented with one or more general-purpose and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuits capable of executing software. Those skilled in the art will recognize how to best implement the described functionality of the processing system, depending on the particular application and the overall design constraints imposed on the overall system.

[0135] If implemented in software, the functions described above may be stored on or transmitted as one or more instructions or code on a computer-readable medium. Software is construed broadly to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. Computer-readable media includes both computer storage media and communication media, such as any medium that facilitates transfer of a computer program from one place to another. A processor may be responsible for managing the bus and general processing, including the execution of software modules stored on the computer-readable storage medium. The computer-readable storage medium may be coupled to the processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integral to the processor. By way of example, computer-readable media may include a transmission line, a carrier wave modulated with data, and / or a computer-readable storage medium on which instructions are stored separately from a wireless node, all of which may be accessed by the processor via a bus interface. Alternatively or additionally, the computer-readable medium, or any portion thereof, may be integrated into the processor, such as in the case of a cache and / or general-purpose register file. Examples of machine-readable storage media include, for example, RAM (random access memory), flash memory, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), registers, magnetic disks, optical disks, hard drives, or any other suitable storage medium or any combination thereof. The machine-readable medium may be embodied in a computer program product.

[0136] A software module may include a single instruction or many instructions and may be distributed across several different code sections, across different programs, and across multiple storage media. A computer-readable medium may include many software modules. A software module contains instructions that, when executed by a device such as a processor, cause a processing system to perform various functions. A software module may include a transmitting module and a receiving module. Each software module may reside on a single storage device or be distributed across multiple storage devices. For example, a software module may be loaded from a hard drive into RAM when a trigger event occurs. During execution of a software module, a processor may load some of the instructions into a cache to speed access. One or more cache lines may then be loaded into a general-purpose register file for execution by the processor. When referring to the functionality of a software module, it is understood that such functionality is implemented by the processor when executing instructions from that software module.

[0137] The following claims are not limited to the embodiments set forth herein but are to be accorded the full scope consistent with the language of the claims. In the claims, when an element is referred to in the singular, it does not mean "only one" unless specifically stated otherwise, but rather "one or more." The term "some" refers to one or more unless specifically stated otherwise. No element of a claim is to be construed under 35 U.S.C. §112(f) unless the element is expressly recited using the phrase "means for," or, in the case of a method claim, unless the element is recited using the phrase "step of." All structural and functional equivalents of the elements of various aspects described throughout this disclosure that are known or later become known to those skilled in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Furthermore, nothing disclosed herein is intended to be made available to the public, regardless of whether such disclosure is expressly recited in the claims.

Claims

1. a first imaging device configured to capture a first image of the patient's retina, the first image being an optical coherence tomography (OCT) image; a second imaging device configured to capture a second image of the patient's retina according to an imaging modality other than OCT; and Equipped with At least one of the following: the sensor of the first imaging device is shared with the second imaging device, or the optical component is configured to select whether the first imaging device or the second imaging device receives light reflected from the retina to the first imaging device or the second imaging device; system.

2. The system of claim 1 , wherein the imaging modality other than OCT is fundus autofluorescence (FAF).

3. The system of claim 2 , wherein the sensor of the first imaging device is shared with the second imaging device, and the sensor is a spectrometer.

4. the optical component is a first actuated mirror; the first imaging device comprises a first light source, and the second imaging device comprises a second light source; the first actuating mirror is further configured to select which of the first light source and the second light source is used to illuminate the retina. The system of claim 3 .

5. 5. The system of claim 4, further comprising a second actuating mirror configured to select between (a) allowing light from a first light source to enter the spectrometer and (b) allowing light from the second light source to enter the spectrometer.

6. The system of claim 1 , wherein the imaging modality other than OCT is a multispectral imaging camera.

7. The system of claim 1 , wherein the first imaging device is a spectral domain OCT (SD-OCT) device.

8. one or more processing devices and one or more memory devices coupled to the one or more processing devices, the one or more memory devices storing executable code that, when executed by the one or more processing devices, causes the one or more processing devices to: receiving the first image and the second image; processing the first image and the second image with a machine learning model to obtain at least one of a diagnosis of a retinal pathology and an identification of features represented in at least one of the first and second images and corresponding to the retinal pathology; one or more processing devices and one or more memory devices The system of claim 1 further comprising:

9. one or more processing devices and one or more memory devices coupled to the one or more processing devices, the one or more memory devices storing executable code that, when executed by the one or more processing devices, causes the one or more processing devices to: receiving the first image and the second image; processing the first image and the second image with a plurality of first machine learning models to obtain a plurality of outputs representative of features of the retina; processing the plurality of outputs with a second machine learning model to obtain a diagnosis of a retinal pathology; one or more processing devices and one or more memory devices The system of claim 1 further comprising:

10. 10. The system of claim 9, wherein the executable code, when executed by the one or more processing devices, further causes the one or more processing devices to process the plurality of outputs with the second machine learning model to obtain a stage of the retinal pathology.

11. 10. The system of claim 9, wherein the executable code, when executed by the one or more processing devices, further causes the one or more processing devices to identify a metabolic state of the retina.

12. 12. The system of claim 11, further comprising a display, wherein the executable code, when executed by the one or more processing devices, further causes the one or more processing devices to output a representation of the metabolic state of the retina to the display.

13. The system of claim 11 , wherein the metabolic state includes at least one of blood flow, oxygen saturation, or waste removal.

14. The system of claim 1 , wherein the first imaging device and the second imaging device are calibrated to include the same magnification.

15. The system of claim 1 , wherein the optical axes of the first imaging device and the second imaging device are aligned.