System for imaging and diagnosis of retinal disease

By combining multispectral imaging and optical coherence tomography technology and applying machine learning models for image analysis, the problem of limited ophthalmic diagnostic efficiency and accuracy in the prior art is solved, and early disease diagnosis and detailed biomarker recognition are achieved.

CN119947635APending Publication Date: 2025-05-06ALCON INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380067827.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-27
Filing Date
2023-09-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize multispectral imaging and optical coherence tomography to provide comprehensive analysis in ophthalmic diagnosis, resulting in limited diagnostic efficiency and accuracy.

Method used

Early disease diagnosis is achieved by combining multispectral imaging and optical coherence tomography techniques, using machine learning models to feature extraction, enhancement and biomarker recognition of captured images.

Benefits of technology

It improves the efficiency and accuracy of ophthalmic diagnosis, can identify retinal pathology early and provide detailed biomarker information to support more accurate treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119947635A_ABST
    Figure CN119947635A_ABST
Patent Text Reader

Abstract

In certain embodiments, a system, computer-implemented method, and computer-readable medium for performing a comprehensive analysis of MSI and OCT images to diagnose ocular disorders are disclosed. The MSI and OCT are processed using separate input machine learning models to create input feature maps that are input to the intermediate machine learning model. An intermediate machine learning model processes the input feature map and outputs a final feature map processed by one or more output machine learning models that output one or more estimated representations of a pathology of the patient's eye. A single device captures OCT images and non-OCT images using (a) a sensor of a first imaging device shared with a second imaging device and / or (b) optical components for directing light from the retina to the first imaging device or the second imaging device.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Multispectral imaging (MSI) is a technique that involves measuring (or capturing) light from a sample (e.g., ocular tissue / structure) at different wavelengths or spectral bands within the electromagnetic spectrum. MSI can capture more information from a sample that may not be visible through conventional imaging, which typically uses broadband illumination and broadband imaging sensors. The MSI information obtained by an MSI imaging system can be used to diagnose ocular disorders and enable real-time adjustments to the use of instruments (e.g., forceps, lasers, probes, etc.) used to manipulate ocular tissue / structure during surgical procedures.

[0002] Optical coherence tomography (OCT) is a technique that uses light waves to generate two-dimensional (2D) and three-dimensional (3D) images of the eye. 2D OCT can involve the use of time-domain OCT and / or Fourier domain OCT, which involves the use of spectral domain OCT methods and swept source OCT methods. 3D OCT can similarly utilize time-domain OCT imaging techniques and Fourier domain OCT imaging techniques. OCT imaging can also be used to diagnose eye disorders preoperatively or during surgery.

[0003] Better utilization of the capabilities of MSI and OCT to diagnose ocular disorders would be an advancement in the field. Summary of the invention

[0004] In certain embodiments, a system is provided. The system includes a first imaging device configured to capture a first image of a patient's retina, the first image being an optical coherence tomography (OCT) image. The system further includes a second imaging device configured to capture a second image of the patient's retina according to an imaging mode other than OCT. In the system, at least one of the following: (a) a sensor of the first imaging device is shared with the second imaging device, and (b) an optical component is configured to select which of the first imaging device and the second imaging device receives light reflected from the retina to the first imaging device or the second imaging device. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] In order to be able to understand the above-mentioned features of the present disclosure in detail, the present disclosure briefly summarized above can be described in more detail with reference to the embodiments, some of which are shown in the accompanying drawings. However, it should be noted that the accompanying drawings only show exemplary embodiments and are therefore not to be considered as limiting the scope thereof, and other equally effective embodiments may be allowed.

[0006] Figure 1 An example system for performing combined analysis of MSI images and OCT images to diagnose eye disorders in accordance with certain embodiments is presented.

[0007] Figure 2A is a diagram illustrating a first method for training a machine learning model to perform comprehensive analysis of MSI images and OCT images to diagnose eye disorders according to certain embodiments.

[0008] Figure 2B is a diagram illustrating a second method for training a machine learning model to perform comprehensive analysis of MSI images and OCT images to diagnose eye disorders according to certain embodiments.

[0009] Figure 2C is a diagram illustrating a third method for training a machine learning model to perform comprehensive analysis of MSI images and OCT images to diagnose eye disorders according to certain embodiments.

[0010] Figure 3 is a flowchart of a method for training a machine learning model to perform comprehensive analysis of MSI images and OCT images to diagnose eye disorders according to certain embodiments.

[0011] Figure 4A A system for capturing both OCT images and MSI images, as well as spectral information, according to certain embodiments is presented.

[0012] Figure 4B An alternative system for capturing both OCT and MSI images, as well as spectral information, in accordance with certain embodiments is presented.

[0013] Figure 4C is a more detailed illustration of a system for capturing both OCT images and MSI images, as well as spectral information, in accordance with certain embodiments.

[0014] Figure 5A and Figure 5B A system for diagnosing eye disorders using a machine learning model according to certain embodiments is presented.

[0015] Figure 6 An example computing device is shown that at least partially implements one or more functions for performing comprehensive analysis of images from multiple imaging modalities in accordance with certain embodiments.

[0016] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the figures. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further recitation. DETAILED DESCRIPTION

[0017] The various embodiments described herein provide a framework for using artificial intelligence to process information obtained from MSI images and OCT images. The advantage of MSI is that MSI images contain rich information about the retina over a wide spectral band range, and this information is a feature that cannot be seen using human vision or fundus cameras. The wide spectral band range of MSI further provides a high degree of depth penetration into the retina. However, MSI images do not provide structural information. In contrast, OCT images do provide structural information about the retina. However, interpreting OCT images requires a high degree of expertise. Using the methods described herein, the rich details and high depth penetration of MSI can be combined with the structural information of OCT to identify biomarkers for various pathologies and perform early disease diagnosis.

[0018] Figure 1 A system 100 is shown for performing integrated analysis of an MSI image 102 and an OCT image 104. The system 100 may include three main stages: a feature extraction stage using machine learning models 106a, 106b, a feature enhancement stage using machine learning models 110, and a biomarker and prediction stage using machine learning models 114, 116. Through these three stages, the system 100 processes the MSI image 102 and the OCT image 104 separately for feature extraction and then combines the extracted features to obtain a meaningful interpretation.

[0019] The MSI image 102 may be captured using any method known in the art for performing MSI, including so-called hyperspectral imaging (HSI). Likewise, the OCT image 104 may be obtained using any method known in the art for performing OCT.

[0020] The MSI image 102 is obtained by illuminating the patient's eye using a multi-spectral band illumination source (e.g., a narrow-band illumination source, a narrow-band filter, etc.) and / or measuring the reflected light using a multi-spectral band camera (e.g., an imaging sensor capable of sensing multiple spectral bands beyond the RGB spectral band). Therefore, each MSI image 102 represents the reflected light within a specific spectral band. The differences between the MSI images 102 are due to the different reflectivities of different structures within the eye for different spectral bands. Therefore, when considered as a whole, the MSI image 102 provides more information about the eye structure than a single broadband image. In some embodiments, the MSI image 102 is a frontal image of the retina for detecting retinal pathology. However, MSI images 102 of other parts of the eye (such as the vitreous or anterior chamber) may also be used.

[0021] Optical coherence tomography (OCT) is a technique that uses light waves from a coherent light source (i.e., a laser) to generate two-dimensional (2D) and three-dimensional (3D) images of the eye. An OCT image is typically a cross-sectional image of the eye for a plane parallel to and colinear with the optical axis of the eye. However, a 3D image can be constructed using OCT images of multiple cross-sectional planes, from which a 2D image can be generated for a cross-sectional plane that is not parallel to the optical axis. For example, a frontal image of the retina can be obtained from the 3D image. In some embodiments, the OCT image 104 is such a frontal image of the retina. OCT is capable of imaging the retina to a certain depth, so that in some embodiments, the OCT image 104 is a collection of frontal images of image planes from at or above the surface of the retina to a depth within or below the retina.

[0022] Although the examples described herein relate to the use of MSI images 102 and OCT images 104, images from any pair of imaging modalities, or images from three or more different imaging modalities, may be used in a similar manner. For example, the additional imaging modalities may include a scanning laser ophthalmoscope (SLO), a fundus camera, and / or a broadband visible light camera.

[0023] In the system 100, the MSI image 102 is processed by a machine learning model 106a, and the OCT image 104 is processed by a machine learning model 106b. The machine learning models 106a, 106b may be implemented as a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of neural networks.

[0024] The result of the machine learning models 106a, 106b processing the images 102, 104 is feature maps 108a, 108b, respectively. For example, the feature maps 108a, 108b may be the output of one or more hidden layers of the machine learning models 106a, 106b. The feature maps 108a, 108b may be two-dimensional or three-dimensional arrays of values. In the case where the feature maps 108a, 108b are two-dimensional arrays, the feature maps 108a, 108b may include the same dimensions in two dimensions, or they may be different. In the case where one or both of the feature maps 108a, 108b are three-dimensional arrays, the feature maps 108a, 108b may include the same dimensions in at least two dimensions, or they may be different in any of the three dimensions.

[0025] The feature maps 108a, 108b and possibly the images 102, 104 are processed by a machine learning model 110. The machine learning model 110 may be implemented as a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of neural networks. The result of the machine learning model 110 processing the feature maps 108a, 108b and possibly the images 102, 104 is a feature map 112. For example, the feature map 112 may be the output of one or more hidden layers of the machine learning model 110, as discussed in more detail below.

[0026] The feature map 112 and possibly the images 102, 104 are then processed by a machine learning model 114 and a machine learning model 116, which then outputs one or more biometric segmentation maps 118 that label eye features represented in the images 102, 104 that correspond to one or more pathologies. Each biometric segmentation map 118 may be in the form of an image having the same dimensions as the images 102, 104, and in which non-zero pixels correspond to pixels in the images 102, 104 that are identified as corresponding to a particular pathology represented by the biometric segmentation map. The biometric segmentation map 118 may include a separate map for each of a plurality of pathologies, or a single map in which all pixels representing any of a plurality of pathologies are non-zero.

[0027] The machine learning model 114 can be implemented as a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of neural networks. For example, the machine learning model 114 can be implemented as a U-shaped network.

[0028] The machine learning model 116 outputs a disease diagnosis 120 and a possible severity score 122 corresponding to the disease diagnosis. The machine learning model 116 can be implemented as a long short-term memory (LSTM) machine learning model, a generative adversarial network (GAN) machine learning model, or other types of machine learning models. The disease diagnosis 120 can be output in the form of text naming the pathology, a numeric code corresponding to the pathology, or some other representation. The severity score 122 can be a numeric value, such as a value from 1 to 10 or a value within some other range. The severity score 122 can be limited to a set of discrete values ​​(e.g., an integer from 1 to 10), or it can be any value within the precision limit of the number of digits used to represent the severity score 122.

[0029] Pathologies for which biometric segmentation maps 118 may be generated and for which diagnoses 120 and severity scores 122 may be generated include at least those pathologies that cause perceptible changes to the retina, such as at least the following:

[0030] Retinal tear

[0031] Retinal detachment

[0032] Diabetic retinopathy

[0033] Hypertensive retinopathy

[0034] Sickle cell retinopathy

[0035] Central retinal vein occlusion

[0036] Epiretinal membrane

[0037] Macular hole

[0038] Macular degeneration (including age-related macular degeneration)

[0039] Retinitis pigmentosa

[0040] ·glaucoma

[0041] Alzheimer's disease

[0042] Parkinson's disease

[0043] The biometric segmentation map 118 may, for example, mark vascular features corresponding to pathologies. Examples of vascular features that may be used to diagnose pathologies are described in the following references, both of which are incorporated herein by reference in their entirety:

[0044] Segmenting Retinal Vessels Using a Shallow Segmentation Network to Aid Ophthalmic Analysis, M. Arsalan et al., Mathematics, 2022, vol. 10, p. 1536.

[0045] PVBM: A Python Vasculature Biomarker Toolbox Based on Retinal BloodVessel Segmentation, J. Fhima et al., Cornell University (July 31, 2022).

[0046] Figure 2A Example methods for training machine learning models 106a, 106b, 110, 114, 116 are shown. In particular, Figure 2A A supervised machine learning method is shown using a plurality (e.g., hundreds, thousands, tens of thousands, hundreds of thousands, or more) of training data entries 200. Each training data entry 200 may include as input an MSI image 102 and an OCT image 104. Each of the MSI images 102 represents an image obtained by detecting light in a different spectral band relative to the other MSI images 102.

[0047] The MSI image 102 and the OCT image 104 images of the training data entry 200 may be images of the same eye of the patient and may be captured substantially simultaneously, such that the anatomical structures represented in the images 102, 104 are substantially the same. For example, "substantially simultaneously" may mean within 1 second and 1 hour of each other. However, "substantially simultaneously" may depend on the pathology being detected: those pathologies that progress very slowly may use images 102, 104 that were captured at a longer time difference (such as less than a day, less than a week, or some other time difference). The MSI image 102 and the OCT image 104 are preferably aligned and scaled relative to each other such that a given pixel coordinate in the MSI image 102 represents substantially the same location in the eye as the same pixel coordinate in the OCT image 104 (e.g., within 0.1 mm, within 1 μm, or within 0.01 μm). Such alignment and scaling may be achieved for the entire image 102, 104 or for at least a portion of one or both of the images 102, 104 showing the anatomical structure of interest (e.g., the macula of the retina).

[0048] The alignment and scaling of the images 102, 104 relative to each other may be achieved by aligning the optical axes of the instruments used to capture the images 102, 104 and calibrating the magnification of the instruments to achieve substantially the same (e.g., within + / - 0.1%, within 0.01%, or within 0.001%) scaling. Alternatively, the alignment and scaling of the images 102, 104 may be achieved by analyzing the anatomical structures represented in the images 102, 104. For example, where the MSI image 102 and the OCT image 104 represent the retina of an eye, the vascular pattern represented in each image 102, 104 may be used to align and scale one or both of the images 102, 104. Where the images 102, 104 have different sizes or do not completely overlap, such as after registration and scaling, non-overlapping portions of one or both of the images 102, 104 may be trimmed and / or one or both of the images 102, 104 may be padded so that the images 102, 104 have the same size and completely overlap each other.

[0049] Each training data entry 200 may include as desired outputs some or all of one or more biomarker segmentation maps 118, disease diagnoses 120, and severity scores 122. Multiple pathologies may be present in the same patient, such that a segmentation map 118, disease diagnoses 120, and severity scores 122 may be included for each pathology present or a subset of the most prevalent pathologies. The desired outputs are generated by a human expert based on an evaluation of the images 102, 104 and other health information of the patient that may be obtained before or after the images 102, 104 are captured. In particular, a biomarker segmentation map 118 for a pathology may include pixels in one or both of the images 102, 104 that are labeled by the human expert as corresponding to that pathology.

[0050] For each training data entry 200, the machine learning model 106a receives the MSI image 102 to generate one or more estimated biomarker segmentation maps. For example, the output of the machine learning model 106a can be a three-dimensional array in which each two-dimensional array along the third dimension is an estimated biometric segmentation map corresponding to a pathology.

[0051] The training algorithm 202 compares the one or more estimated biomarker segmentation maps to the one or more biomarker segmentation maps 118 of the training data entry 200. The training algorithm 202 then updates one or more parameters of the machine learning model 106a based on the difference between each estimated biomarker segmentation map of the pathology and the corresponding biomarker segmentation map 118 of the pathology in the training data entry 200.

[0052] The machine learning model 106b can be trained in a similar manner as the machine learning model 106a. For each training data entry 200, the machine learning model 106b receives the OCT image 104 and generates one or more estimated biomarker segmentation maps. For example, the output of the machine learning model 106a can be a three-dimensional array in which each two-dimensional array along the third dimension is an estimated biometric segmentation map corresponding to a pathology.

[0053] The training algorithm 202 (which may be the same or different than the training algorithm used to train the machine learning model 106a) compares the one or more estimated biomarker segmentation maps to the one or more biomarker segmentation maps 118 of the training data entry 200. The training algorithm 202 then updates one or more parameters of the machine learning model 106b based on the difference between each estimated biomarker segmentation map for the pathology and the corresponding biomarker segmentation map 118 for the pathology in the training data entry 200.

[0054] In the case of using three or more imaging modes, additional machine learning models may be present and trained in a similar manner. Each machine learning model corresponds to an imaging mode and processes the corresponding images of the imaging mode in the training data entry. The machine learning model generates one or more estimated biomarker segmentation maps, which are compared with one or more biomarker segmentation maps 118 of the training data entry by the training algorithm, and the training algorithm then updates the machine learning model based on the comparison. The hidden layer of each machine learning model can generate outputs, which are used as feature maps of the imaging mode corresponding to the machine learning model.

[0055] As mentioned above about Figure 1 As described, the machine learning model 110 takes the feature maps 108a, 108b of the machine learning models 106a, 106b as input. The machine learning models 106a, 106b may be trained with some or all of the training data entries 200 before training the machine learning model 110.

[0056] For each training data entry 200, the machine learning model 110 receives feature maps 108a, 108b obtained by processing the MSI image 102 and the OCT image 104 of the training data entry 200 with the machine learning models 106a, 106b. As described above, the feature maps 108a, 108b can be the output of the hidden layers (i.e., the layers other than the final layer that outputs one or more estimated biomarker segmentation maps) of the machine learning models 106a, 106b, respectively. The machine learning model 110 can also receive the MSI image 102 and the OCT image 104 as input, although in other embodiments, only the feature maps 108a, 108b are used.

[0057] The machine learning model 110 processes the feature maps 108a, 108b and possibly the MSI image 102 and the OCT image 104 and produces one or more estimated biomarker segmentation maps. For example, the output of the machine learning model 110 may be a three-dimensional array in which each two-dimensional array along the third dimension is an estimated biomarker segmentation map corresponding to a pathology. Although two feature maps 108a, 108b are shown, the machine learning model 110 may process any number of feature maps, and any number of possible images for generating feature maps, in a similar manner for any number of imaging modalities.

[0058] The training algorithm 204 compares the one or more estimated biomarker segmentation maps to the one or more biomarker segmentation maps 118 of the training data entry 200. The training algorithm 204 then updates one or more parameters of the machine learning model 110 based on the difference between each estimated biomarker segmentation map for the pathology and the corresponding biomarker segmentation map 118 for the pathology.

[0059] As mentioned above about Figure 1 As described, the machine learning models 114, 116 take as input the feature map 112 of the machine learning model 110. The machine learning model 110 may be trained with some or all of the training data entries 200 before training the machine learning models 114, 116.

[0060] For each training data entry 200, the machine learning model 114, 116 receives a feature map 112 obtained by processing the MSI image 102 and the OCT image 104 of the training data entry 200 with the machine learning model 106a, 106b, 110. As described above, the feature map 112 can be the output of a hidden layer (i.e., a layer other than the final layer that outputs one or more estimated biomarker segmentation maps) of the machine learning model 110. The machine learning model 114, 116 can also take the MSI image 102 and the OCT image 104 as input, although in other embodiments, only the feature map 112 is used.

[0061] The machine learning model 114 processes the feature maps 112 and possibly the images 102, 104 from the training data entries 200 and produces one or more estimated biomarker segmentation maps. In the case where three or more imaging modalities are used, images from the training data entries 200 according to the three or more imaging modalities may be processed by the machine learning model 114 along with the feature maps 112 obtained from these images. The output of the machine learning model 114 may be a three-dimensional array in which each two-dimensional array along the third dimension is an estimated biomarker segmentation map corresponding to a pathology.

[0062] The training algorithm 206a compares the one or more estimated biomarker segmentation maps to the one or more biomarker segmentation maps 118 of the training data entry 200. The training algorithm 206a then updates one or more parameters of the machine learning model 114 based on the difference between each estimated biomarker segmentation map of the pathology and the corresponding biomarker segmentation map 118 of the pathology.

[0063] The machine learning model 116 processes the feature map 112 and possibly the MSI image 102 and the OCT image 104 and generates one or more estimated diagnoses and an estimated severity score for each estimated diagnosis. Where three or more imaging modalities are used, images from the training data entry 200 according to the three or more imaging modalities may be processed by the machine learning model 116 along with the feature maps 112 obtained for those images.

[0064] The output of the machine learning model 116 may be a vector, where each element of the vector (if non-zero) indicates the estimated presence of a pathology. The output of the machine learning model 116 may also be text listing one or more major pathologies estimated to be present. The output of the machine learning model 116 may further include a severity score for each pathology estimated to be present, such as a vector in which each element corresponds to a pathology and the value of the element indicates the severity of the corresponding pathology.

[0065] The training algorithm 206b compares the estimated diagnosis and corresponding severity score to the disease diagnosis 120 and severity score 122 of the training data entry 200. The training algorithm 206b then updates one or more parameters of the machine learning model 116 based on the difference between the estimated diagnosis and corresponding severity score and the disease diagnosis 120 and severity score 122 of the training data entry 200.

[0066] refer to Figure 2B In some embodiments, training of one or both of the machine learning models 106a, 106b may be performed via unsupervised training algorithms 210a, 210b, respectively. Figure 2B While training using images 102 , 104 is shown, it should be understood that one or more machine learning models for additional or alternative imaging modalities may be trained in the same manner.

[0067] for Figure 2B In an embodiment, the machine learning models 110, 114, 116 may be as described above with respect to Figure 2A In some embodiments, only one of the machine learning models 106a, 106b is trained using the unsupervised training algorithm 210a, 210b, while the Figure 2A The described supervised training algorithm 202 is used to train another machine learning model. For the unsupervised machine learning algorithms 210a, 210b, no labeled training data entries are used.

[0068] The machine learning model 106a may be trained using a corpus of a collection of MSI images 102. The corpus may be curated to include a large set of MSI images (e.g., retinal images) of healthy eyes that are free of pathology, as well as a small portion (e.g., less than 5% of the corpus or less than 1 percent of the corpus) that correspond to one or more pathologies. The collection of MSI images 102 may or may not be labeled with whether the collection of images 102 represents a pathology and / or the specific pathology represented.

[0069] The unsupervised training algorithm 210a uses the machine learning model 106a to process the corpus and train the machine learning model 106a to identify and classify anomalies detected in the collection of MSI images 102 in the corpus. The unsupervised training algorithm 210a may be implemented using any method for performing anomaly detection or other unsupervised machine learning known in the art. The output of the machine learning model 106a may be an image of the same size as the individual MSI images 102, with pixels representing anomalies labeled.

[0070] The machine learning model 106b may be trained using a corpus of OCT images 104. The corpus may be curated to include a large number of OCT images of healthy eyes (e.g., retinal images) that are free of pathology, as well as a small portion (e.g., less than 5% of the corpus or less than 1 percent of the corpus) that correspond to one or more pathologies. The OCT images 104 may or may not be labeled with whether the set of images 102 represents a pathology and / or the specific pathology represented.

[0071] The unsupervised training algorithm 210b uses the machine learning model 106b to process the corpus and train the machine learning model 106a to identify and classify abnormalities detected in the OCT images 104 of the corpus. The unsupervised training algorithm 210b can be implemented using any method for performing anomaly detection or other unsupervised machine learning known in the art. The output of the machine learning model 106b can be an image of the same size as each OCT image 104, in which pixels representing abnormalities are labeled.

[0072] The set of MSI images 102 and OCT images 104 used to train the machine learning model by the unsupervised training algorithm 210a, 210b may include images 102, 104 from the training data entry 200 used to train other machine learning models 110, 114, 116. The set of MSI images 102 and OCT images 104 can be further enhanced with images of healthy eyes to facilitate the identification of abnormalities corresponding to pathology. The set of MSI images 102 and OCT images 104 can be constrained to be the same size and can be aligned with each other. For example, although the images 102, 104 are images of multiple different eyes, the images 102, 104 can be aligned to place the representation of the center of the fovea of ​​the retina substantially at the center of each image 102, 104 (e.g., within 1, 2, or 3 pixels). Certain other features (such as the fundus) can be used for alignment. Although the images 102, 104 are images of multiple different eyes, the images 102, 104 can also be scaled so that the size of the anatomical structures represented in the images is substantially the same. For example, the images 102, 104 may be scaled so that the fovea, fundus, or one or more other anatomical features are the same size.

[0073] Once trained, the machine learning models 106a, 106b can provide information to the machine learning model 110 (see Figure 1 and Figure 2A ) provide outputs, which are the outputs of one or more hidden layers of the machine learning models 106a, 106b. Alternatively, the final outputs of the machine learning models 106a, 106b (e.g., images with abnormal labels) can be used as inputs to the machine learning model 110.

[0074] refer to Figure 2C , in the right Figure 2B In an improvement of the unsupervised machine learning approach, the supervised training algorithm 212b may compare the output of the machine learning model 106a with the output of the machine learning model 106b for a given set of MSI images 102 and OCT images 104 of the same patient's eye captured substantially simultaneously as defined above. The supervised training algorithm 212b may then adjust the parameters of the machine learning model 106b based on the comparison so as to train the machine learning model 106b to recognize the same abnormalities detected by the machine learning model 106a. Note that the reverse approach may alternatively or additionally be used: the output of the machine learning model 106b may be used by the supervised training algorithm 212b or a different supervised training algorithm 212a to train the machine learning model 106a to recognize the abnormalities identified by the machine learning model 106b.

[0075] In some embodiments, training can be performed in stages, each stage using the above Figure 2A , Figure 2B and Figure 2C In the first example, we first use Figure 2A The supervised machine learning method of can be used to train the machine learning models 106a and 106b; then, the supervised machine learning method of can be used to train the machine learning models 106a and 106b; Figure 2B The unsupervised method is used to train the machine learning models 106a, 106b; and then, according to Figure 2C The method further trains the machine learning model 106b based on the output of the machine learning model 106a (and / or vice versa). In the second example, only unsupervised learning is used: using Figure 2B The unsupervised method of training machine learning models 106a and 106b separately is then used to Figure 2C The method further trains the machine learning model 106b based on the output of the machine learning model 106a, and / or trains the machine learning model 106a based on the output of the machine learning model 106b.

[0076] Figure 2C While training using images 102, 104 is shown, it should be understood that one or more machine learning models for additional or alternative imaging modalities may be trained in the same manner. In particular, the output of a machine learning model according to one imaging modality may be used to train one or more other machine learning models according to one or more other imaging modalities in the same manner. The outputs of two or more first machine learning models of one or more first imaging modalities may be concatenated or otherwise combined and used to train one or more other machine learning models according to one or more other imaging modalities. Figure 2C Method to train one or more second machine learning models of one or more second machine learning models.

[0077] refer to Figure 3 The method 300 shown may be performed by a computer system (e.g. Figure 6 The method 300 includes, in step 302, training a first input machine learning model using training images of a first imaging mode. For example, step 302 may include the following steps: FIG. 2A to FIG. 2C A machine learning model 106a trained using MCI images 102 of any of the methods described.

[0078] Method 300 includes, in step 304, training a second input machine learning model using training images of a second imaging mode. For example, step 304 may include training a second input machine learning model using training images of a second imaging mode. FIG. 2A to FIG. 2C A machine learning model 106b trained using OCT images 104 of any of the methods described.

[0079] Method 300 includes, in step 306, processing the image according to the first imaging mode using the first input machine learning model to obtain an input feature map F1, and processing the image according to the second imaging mode using the second input machine learning model to obtain an input feature map F2. Feature maps F1 and F2 may be outputs of hidden layers of the first input machine learning model and the second input machine learning model, respectively. Step 304 may include processing the MSI image 102 using the machine learning model 106a and processing the OCT image 104 using the machine learning model 106b to obtain feature maps 108a, 108b, as described above with respect to Figure 1 and Figure 2A As described above, the MSI image 102 and the OCT image 104 may be part of a common training data entry 200, such that the MSI image 102 and the OCT image 104 are from the eye of the same patient and are captured substantially simultaneously.

[0080] The method 300 includes, in step 308, training an intermediate machine learning model using the feature maps F1 and F2. Specifically, multiple pairs of feature maps F1 and F2 can each be processed by the intermediate machine learning model, and the output of the intermediate machine learning model can be used to train the intermediate machine learning model. Each pair of feature maps F1 and F2 can be obtained for images of the first mode and the second mode, which are from the eyes of the same patient and are captured substantially at the same time. Step 308 may include using the intermediate machine learning model to process the images used to obtain each pair of feature maps F1 and F2. Step 308 may include training the machine learning model 110 using the feature maps 108a, 108b and the training data entries 200, as described above with respect to Figure 2A as described.

[0081] The method 300 includes, in step 310, processing the feature pairs of feature maps F1 and F2 and possible training images used to obtain each pair of feature maps F1 and F2 using the intermediate machine learning model to obtain a final feature map F. The final feature map F can be obtained from the output of the hidden layer of the intermediate machine learning model. Step 310 can include using the machine learning model 110 to process the feature maps 108a, 108b and possible corresponding images 102, 104 to obtain the feature map 112, as described above with respect to Figure 1 and Figure 2A as described.

[0082] The method 300 includes, in step 312, training one or more output machine learning models using the feature map F. The one or more output machine learning models may be trained to output, for a given feature map F, estimated representations of pathologies represented in training images used to generate the feature map F using the first input machine learning model and the second input machine learning model and the intermediate machine learning model. The one or more output machine learning models may take as input images used to generate the feature map F according to the first imaging mode and the second imaging mode. Step 312 may include training one or both of the machine learning models 114, 116 using the feature map 112 and possibly corresponding images 102, 104 to output some or all of the biomarker segmentation map 118, the disease diagnosis 120, and the severity score 122.

[0083] Method 300 may include, in step 314, processing utilization images according to the first imaging mode and the second imaging mode according to a pipeline of a first input machine learning model and a second input machine learning model, an intermediate machine learning model, and one or more output machine learning models. Specifically, one or more of the utilization images are processed according to the first imaging mode using the first input machine learning model to obtain a feature map F1; one or more of the utilization images are processed according to the second imaging mode using the second input machine learning model to obtain a feature map F2; feature maps F1 and F2 and possible utilization images are processed using the intermediate machine learning model to obtain a feature map F; and feature map F and possible utilization images are processed by one or more output machine learning models to obtain an estimated representation of the pathology represented in the utilization image. The estimated representation may be output to a display device or stored in a storage device for later use or subsequent processing. The feature maps (F1, F2, F) may be displayed or stored additionally.

[0084] For example, step 314 may include processing the utilization images 102, 104 (i.e., images 102, 104 that are not part of the training data entry 200) using the machine learning models 106a, 106b, respectively, to obtain feature maps 108a, 108b, respectively, as described above with respect to Figure 1 The feature maps 108a, 108b and possibly the images 102, 104 may be processed using a machine learning model 110 to obtain a feature map 112. The feature map 112 and possibly the images 102, 104 may be processed by one or both of the machine learning models 114, 116 to obtain a biomarker segmentation map 118, a disease diagnosis 120, and a severity score 122.

[0085] Steps 302 to 314 may be performed sequentially, i.e., the first input machine learning model and the second input machine learning model are trained, then the intermediate machine learning model is trained, then one or more output machine learning models are trained, and finally the utilization image is trained. Steps 302 to 314 may additionally or alternatively be interleaved, i.e., the first input machine learning model and the second input machine learning model, the intermediate machine learning model, and the one or more output machine learning models are trained as a group. For example, in a first stage, the first input machine learning model and the second input machine learning model, the intermediate machine learning model, and the one or more output machine learning models are trained individually in the order listed, and in a second stage, the training continues as a group, i.e., after an iteration including processing a set of images according to a pipeline, some or all of the first input machine learning model and the second input machine learning model, the intermediate machine learning model, and the one or more output machine learning models may be updated as part of the iteration by a training algorithm according to the outputs of the first input machine learning model and the second input machine learning model, the intermediate machine learning model, and the one or more output machine learning models, respectively. Training individually or as a group may be performed in utilizing step 314 (particularly as described with respect to Figure 2B and / or Figure 2C The unsupervised learning described above continues during the

[0086] Step 314 may be performed by a different computer system than the computer system used to perform steps 302 to 312. For example, the pipeline including the first output machine learning model, the second output machine learning model, the third output machine learning model, and the one or more output machine learning models may be installed on one or more other computer systems for use by a surgeon or other medical professional.

[0087] Although method 300 is described with respect to two imaging modes, three or more imaging modes may be used in a similar manner. i , i=1 to N, where N is greater than or equal to two. For a given training data entry, the imaging mode IM i The images can meet the requirements of substantially the same scaling, alignment, and simultaneous imaging of the same eye of the patient. There may be an input machine learning model ML i , i = 1 to N, each machine learning model ML i Corresponding to imaging mode IM i , and each machine learning model processes the imaging mode IM i One or more corresponding images are used to generate the corresponding feature map F i . The corresponding imaging mode IM iThe images are trained according to any of the methods described above for training machine learning models 106a, 106b. i .

[0088] Therefore, in this embodiment, the intermediate machine learning model converts the N feature maps F i (i=1 to N) and used to generate the feature map F i The possible training images are taken as input. The output machine learning model will generate the final feature map F and the i The possible training images are taken as input. The intermediate machine learning model and the output machine learning model are trained as described above with respect to the machine learning model 110 and the machine learning models 114 and 116.

[0089] Figure 4A , Figure 4B and Figure 4C Example systems 400a, 400b, 400c are shown that can be used to capture images of multiple imaging modes substantially simultaneously and detect some or nearly all known retinal pathologies. For example, diabetic retinopathy is becoming increasingly common. Age-related macular degeneration (AMD) and glaucoma are also relatively common in the elderly. The elderly typically have at least some reduction in vision due to one or more retinal pathologies. Because many retinal diseases are progressive, early detection is critical to improving quality of life and reducing blindness.

[0090] Typically, an ophthalmology clinic setting has a variety of tools, including individual instruments such as OCT, fundus cameras, scanning laser ophthalmoscopes (SLO), etc. In the early stages of retinal diseases, it is difficult to distinguish between different diseases. Images from multiple imaging modalities help differentiate between diseases.

[0091] OCT images provide high-resolution structural information of all layers of the retina at the micrometer scale. However, in the early stages of the disease, structural changes are unlikely. The metabolic state is often abnormal before any structural changes are detected by OCT.

[0092] Color fundus cameras provide color fundus images but do not cover the fluorescence range of 500 nm to 600 nm, which is important for detecting fluorophores deposited in the retinal pigment epithelium (RPE) that are present in the early stages of AMD and may develop into drusen or atrophy.

[0093] Fundus autofluorescence (FAF) enables imaging of indicators of retinal degeneration as well as focal hypo- and hyperpigmentation at the level of the RPE. FAF is the most reliable imaging modality for detecting, depicting, quantifying, and monitoring the progression of outer retinal atrophy. BlinD (basal linear deposits) and BLamD (basal lamellar deposits) in the RPE are precursors to AMD and can be visualized using FAF rather than OCT.

[0094] Currently, no single tool can reliably perform early detection, risk assessment, and progression monitoring of retinal disease. Figure 4A , Figure 4B and Figure 4C Example systems 400a, 400b, 400c are shown that can provide such functionality. Machine learning models can be used to process images obtained using systems 400a, 400b, 400c to further facilitate early detection, risk assessment, and monitoring of retinal disease progression.

[0095] Specific reference Figure 4A , system 400a includes OCT 402. OCT 402 can be implemented as any OCT known in the art, which can include a light source 404, an output optical device 406, and a detector 408. Light source 404 can be a coherent light source, such as a laser or a laser diode. Light source 404 can also be a low-coherence broadband light source. Output optical device 406 can include scanning mirrors, focusing optical devices, and a mechanism for translating the focal depth of OCT 402 (such as time-domain OCT). Detector 408 can include a spectrometer, such as a diffraction grating and a charge-coupled device (CCD), a complementary metal oxide semiconductor (CMOS) sensor, or other detector. The output of detector 408 is an image, or a sample stream that can be organized into an image based on the state of OCT 402 when each sample is detected.

[0096] Output optics 406 directs light from light source 404 onto retina 412 of patient's eye 414. Light from light source 404 may pass through one or more lenses 416 to focus the light on retina 412. Light reflected from retina 412 returns to output optics 406 and at least a portion thereof is directed to detector 408.

[0097] An actuated reflector 418 (e.g., a toggle switch reflector) can be positioned or actuated to position 420 as shown. When the OCT 402 is in use, the actuated reflector 418 is placed in position 420 and allows light reflected from the retina 412 to reach the OCT 402 without interacting with the reflector 418. The reflector 418 can be positioned or actuated to position 420 as shown. Figure 4AMirror 418 may be positioned as shown to capture images according to one or more other imaging modes. For example, the reflective surface of mirror 418 may be oriented at an angle of approximately 45 degrees (e.g., + / - 2 degrees) relative to the optical axis of lens 416 and / or eye 414. Mirror 418 may be manually switched between the positions shown or may be coupled to an electrical actuator 418a.

[0098] Light source 422 can be used to illuminate retina 412 when imaging is performed according to one or more other imaging modes. The one or more other imaging modes can include some or all of MSI, HSI, fundus autofluorescence (FAF), FAF spectrum, infrared, ultraviolet, or other imaging modes. Light source 422 can include (a) a single light source suitable for one or more other imaging modes, (b) a single light source that can operate in different ways (intensity and / or spectrum) for different imaging modes, or (c) multiple light sources, each for a different imaging mode. Light source 422 can be implemented as a broadband light source embodied as one or more LEDs.

[0099] Light from light source 422 may be incident on a reflector 424, such as an annular aperture reflector, a beam splitter, or other reflector capable of partial reflection and transmission. Reflector 424 directs at least a portion of the light from light source 422 to reflector 418, which directs the light to retina 412, such as through lens 416.

[0100] Mirror 418 directs light reflected from retina 412 back to mirror 424, which allows at least a portion of the reflected light to pass therethrough. At least a portion of the reflected light may be directed to spectrometer 428. Spectrometer 428 captures spectral information of the reflected light.

[0101] Additionally or alternatively, a portion of the light reflected from the retina 412 may be reflected by the filter 430 onto the camera 432. The camera 432 may be a monochrome camera or a color camera. The filter 430 may be a filter wheel including a plurality of filters, each corresponding to a different wavelength band. Multiple images of the reflected light may be captured, each image being captured with a different filter from the plurality of filters inserted between the camera 432 and the retina 412. Thus, the plurality of images may constitute an MSI image or an HSI image. The filter 430 may include an electrically controlled actuator for selecting among the plurality of filters, or may be manually adjusted.

[0102] In the case where both the spectrometer 428 and the camera 432 are used, the beam splitter 426 can be positioned so that light reflected from the retina 412 is both (a) transmitted through the beam splitter 426 to one of the camera 432 and the spectrometer 428, and (b) reflected from the beam splitter 426 to the other of the spectrometer 428 and the camera 432.

[0103] The OCT image output by the detector 408, the FAF image and / or FAF spectral information obtained using the spectrometer 428, and the MSI image or HSI image obtained using the camera 432 can be input into the machine learning model 438. The machine learning model 438 is trained for some or all of the following: (a) identifying anatomical structures and features corresponding to pathologies of the retina, (b) diagnosing one or more pathologies of the retina, and (c) estimating the severity of one or more pathologies.

[0104] The machine learning model 438 may be embodied as the system 100 as described above, which has been trained according to any of the embodiments described above. In the case of implementing three or more imaging modes using the system 400a, the system 100 may be configured as described above to include three or more machine learning models 106a, 106b that generate three or more feature maps 108a, 108b, which are input to the intermediate machine learning model 110. In the case of implementing three or more imaging modes using the system 400a, the training data entries 200 for training the machine learning models 106a, 106b, 110, 114, 116 may include images of the three or more imaging modes in addition to or in place of the MSI images 102 and OCT images 104 as described above. The machine learning models 106a, 106b, 110, 114, 116 may be trained in the same manner as described above, wherein each machine learning model 106a, 106b is trained with images of the imaging mode corresponding to the machine learning model 106a, 106b.

[0105] When used in conjunction with system 100, system 400a has a number of advantages. System 400a can easily capture images of multiple imaging modes substantially simultaneously without having to move the patient to another device. System 400a can also be configured so that multiple images according to multiple imaging modes are scaled substantially the same (e.g., within 0.01 percent along each dimension) and are substantially aligned, for example, within 0.1 mm, 0.01 mm, or 1 μm measured relative to retinal features represented in the multiple images. Substantially identical scaling can be achieved by calibrating the magnification of each imaging mode. Substantial alignment can be achieved by substantially aligning the optical axis of each imaging mode with the center of the image obtained using each imaging mode.

[0106] Images from multiple imaging modalities acquired using system 400a can be processed by machine learning model 438 as they are captured. With readily available computing power, processing results can be obtained within minutes. Thus, an ophthalmologist, surgeon, or other medical professional can immediately provide a patient with a diagnosis of nearly any retinal disease.

[0107] refer to Figure 4B , system 400b can be modified relative to system 400a by using an additional camera 436. Light reflected from the retina 412 can be directed to two cameras 432, 436 via a beam splitter 434. Cameras 432, 436 can capture different types of images. For example, cameras 432, 436 can capture two or more of the following types of images: color (RGB), infrared, ultraviolet, MSI, HSI, and FAF. Images from camera 436 and other images from other imaging modes provided by system 400b can be processed using a machine learning model 438, as described above with reference to Figure 4A as described.

[0108] refer to Figure 4C , spectral domain OCT (SD-OCT) can be modified to implement the presented system 400c to capture images according to a variety of other imaging modalities (MSI, HSI, FAF, color, infrared, ultraviolet).

[0109] SD-OCT includes a light source 440. Light source 440 may be a low coherence light source such as a broadband light source that generates light across the visible spectrum (e.g., 380nm to 700nm) and possibly also in the infrared and / or ultraviolet spectrum. Light source 440 may be embodied as one or more light emitting diodes (LEDs).

[0110] Light from light source 440 is input to fiber connector 442. A portion of the light is transmitted to dispersion compensator 444. The light passes through dispersion compensator 444, is incident on reflector 446, and is reflected back to fiber connector 442 by dispersion compensator 444.

[0111] Light from light source 440 is also coupled to optical fiber 448 via fiber optic coupler 442. Optical fiber 448 conducts light to lens 450, which directs light received from optical fiber 448 onto scanning mirror 452. Scanning mirror 452 can be a single mirror actuated along two orthogonal rotation axes to scan light over a two-dimensional area of ​​retina 412. Alternatively, scanning mirror 452 can be embodied as two mirrors, each rotating about one of the two orthogonal rotation axes. For example, scanning mirror 452 can be embodied as a galvanometer and a resonant scanner.

[0112] Light reflected from the scanning mirror 452 passes through a scanning lens 454. The scanning lens 454 is actuated along the optical axis of the scanning lens 454 to change the focal depth of the light from the light source 440 that reaches the retina 412. The scanning lens 454 is thus translated to different positions to image different layers of the retina 412. The scanning lens 454 can be actuated mechanically, or the focal depth can be changed electronically, such as by implementing the scanning lens 454 as an optofluidic lens.

[0113] Light emitted by scanning lens 454 may be directed by one or more other components onto retina 412. For example, one or more mirrors 456 may change the direction of light emitted from lens 454, and one or more lenses 458 may focus light emitted from lens 454 onto retina 412. The position and / or orientation of one or more mirrors 456 and one or more lenses 458 may be adjustable.

[0114] In some embodiments, adaptive optics (AO) 460 may be positioned in the optical path between lens 450 and fiber coupler 442. Adaptive optics 460 improves the quality of images obtained using SD-OCT, and the characteristics of AO 460 may be selected according to any method known in the art of SD-OCT design.

[0115] Light reflected from retina 412 follows a path that is opposite to the path taken by light propagating from fiber optic coupler 442 to retina 412. Upon reaching fiber optic coupler 442, at least a portion of the light reflected from retina 412 is coupled to optical fiber 462 along with at least a portion of the light returned from dispersion compensator 444. Optical fiber 462 directs the light to spectrometer 464, such as through output lens 466, which receives the light output by optical fiber 462 and focuses or collimates the light input to spectrometer 464.

[0116] In some embodiments, fiber connector 442 is coupled to light source 440 and dispersion compensator via multimode optical fiber, and optical fibers 462, 448 are single mode optical fibers.

[0117] Spectrometer 464 can be used for one or more other imaging modes in addition. Therefore, a reflector 468, a beam splitter or other optical element can be used to direct light from multiple light sources into spectrometer 464. In the embodiment shown, reflector 468 is an actuated reflector. Reflector 468 can be placed in the orientation shown to reflect light from output lens 466 into spectrometer 464. Reflector 468 can be moved to orientation 470 to allow light from one or more other light sources to enter spectrometer 464. Reflector 468 can be manually switched between the positions shown or can be connected to an electric actuator 468a.

[0118] Spectrometer 464 can be implemented using any type of spectrometer known in the art. In the illustrated embodiment, spectrometer 464 includes a diffraction grating 472, a lens 474, and a detector 476, such as a CCD or CMOS sensor. The light incident on grating 472 forms a wavelength-dependent fringe pattern, which is focused by lens 474 onto detector 476. Therefore, the light incident on each point of detector 476 can be mapped to a specific wavelength. Therefore, the output of detector 476 can be processed to measure the spectrum of the light entering spectrometer 464. Since the light from light source 440 is scanned on retina 412, the reflectivity spectrum of the light spot on retina 412 can be obtained from each spectral measurement of spectrometer 464.

[0119] Mirrors, beam splitters, or other optical elements may be used to direct light for one or more other imaging modes onto the retina 412. For example, an actuated mirror 480 (such as a toggle switch mirror) may be positioned or actuated to position 482 as shown. When performing SD-OCT imaging, the actuated mirror 480 is placed in position 482 and light is allowed to travel between the light source 440 and the retina 412 without interacting with the mirror 480. The mirror 480 may be as shown in FIG. Figure 4C Mirror 480 may be positioned as shown to capture images according to one or more other imaging modes. For example, the reflective surface of mirror 480 may be oriented at an angle of approximately 45 degrees (e.g., + / - 2 degrees) relative to the optical axis of lens 458 and / or eye 414. Mirror 480 may be manually switched between the positions shown or may be coupled to an electrical actuator 480a.

[0120] Light source 484 can be used to illuminate retina 412 when imaging is performed according to one or more other imaging modes. One or more other imaging modes can include some or all of MSI, HSI, FAF, color, infrared, ultraviolet, or other imaging modes. Light source 484 can include (a) a single light source suitable for one or more other imaging modes, (b) a single light source that can operate in different ways (intensity and / or spectrum) for different imaging modes, or (c) multiple light sources, each light source is used for a different imaging mode. For example, light source 484 can be a broadband light source implemented using one or more LEDs.

[0121] Light from light source 484 may be incident on a reflector 486, such as an annular aperture reflector, a beam splitter, or other reflector capable of partial reflection and transmission. One or more lenses 488 may be inserted between light source 484 and reflector 486 to focus light from light source 484 onto retina 412 or collimate light from light source 484. Reflector 486 directs at least a portion of the light from light source 484 onto reflector 480, which directs the light onto retina 412.

[0122] Mirror 480 directs light reflected from retina 412 back to mirror 486, which allows at least a portion of the reflected light to pass therethrough. A portion of the reflected light may be directed to spectrometer 464.

[0123] Additionally or alternatively, a portion of the light reflected from the retina 412 may be reflected onto the camera 492 by the filter 490. In some embodiments, one or more lenses 494 may be inserted between the filter 490 and the camera 492. The camera 492 may be a monochrome, color, infrared, ultraviolet, or other type of camera. The filter 490 may be a filter wheel including a plurality of filters, each corresponding to a different wavelength band. The filter wheel may be manually adjustable or include an electrically controlled actuator. Multiple images of the reflected light may be captured, each image being captured using a different filter from the plurality of filters inserted between the camera 492 and the retina 412. Thus, the plurality of images may constitute an MSI image or an HSI image. The system 400c may be modified similarly to the system 400b to image the light passing through the filter 490 using a plurality of cameras.

[0124] In the case where both the spectrometer 464 and the camera 492 are used to image light from the light source 484, the beam splitter 496 can be positioned so that light reflected from the retina 412 is both (a) transmitted through the beam splitter 496 to one of the camera 492 and the spectrometer 464, and (b) reflected from the beam splitter 496 to the other of the spectrometer 464 and the camera 492.

[0125] When the system 400c is operated as an SD-OCT, the output of the spectrometer 464 is used to form an OCT image as known in the art of SD-OCT. When the reflector 480 is used to direct light from the light source 484 onto the retina 412, the output of the spectrometer 464 can be used to form a FAF image, and the output of the camera 492 can be used to form an MSI image or an HSI image. Any other camera used can produce FAF, color, infrared, ultraviolet or other types of images. The machine learning model 438 described above can be used to process OCT, MSI or HSI, FAF, and any other images obtained using the system 400c.

[0126] Figure 5A and Figure 5B Additional machine learning models 500a, 500b are shown, which can be used as machine learning model 438 in addition to system 100 or in place of system 100.

[0127] Specific reference Figure 5A , the machine learning model 500a may take as input one or both of the FAF spectrum 502 and the oxygen map 504. The oxygen map measures the oxygen saturation level within the blood vessels of the retina 412. The oxygen map is obtained by analyzing the spectrum reflected from points within the retina 412. Therefore, the oxygen map can be obtained using the output of the spectrometers 428, 464 using any method known in the art for generating an oxygen map. The FAF spectrum 502 can also be obtained from the spectrometers 428, 464.

[0128] The FAF spectra 502 and the oxygen map 504 can each be input to a plurality of machine learning models 506a to 506d, which each produce a corresponding output 508a to 508d describing a particular metric, anatomical structure, or feature corresponding to a pathology. In the illustrated embodiment, the machine learning models 506a to 506d are convolutional neural networks (CNNs) 506a to 506d. However, the machine learning models 506a to 506d can also each be embodied as a deep neural network (DNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of machine learning models.

[0129] Outputs 508a to 508d may include, for example, metabolism in the upper retinal vessels (e.g., anywhere between Bruch's membrane and the vitreous), metabolism in the choroidal vessels, metabolism in the RPE, fluorophore deposition in the RPE, or other anatomical structures or features corresponding to pathologies that may be represented in the spectral information included in the FAF spectrum 502 and / or oxygen map 504.

[0130] The outputs 508a to 508d may be input to an output machine learning model 510. In the illustrated embodiment, the output machine learning model 510 is a random forest. However, other machine learning models may be used, such as a long short-term memory (LSTM) machine learning model, a generative adversarial network (GAN) machine learning model, a logistic regression machine learning model, or other types of machine learning models.

[0131] Output The machine learning model 510 outputs a metabolic state 512 of the retina 412 represented by the FAF spectrum and the oxygen map 504. The metabolic state 512 can reflect information such as blood flow, oxygen saturation, waste removal, or other information. The metabolic state 512 can be in the form of a numerical value indicating the overall metabolic state of the retina 412, for example, low represents poor health and high represents good health. The metabolic state 512 can be a plurality of numerical values, each corresponding to an aspect of the metabolism of the retina 412. The metabolic state can be stored and / or output to a display device. For example, a representation of some or all of the FAF spectrum 502, the oxygen map 504 (outputs 508a to 508d and the metabolic state) can be stored and / or output to a display device for evaluation by an ophthalmologist, surgeon, or other health professional.

[0132] The training data entries may be used to train the machine learning models 506a-506d and the machine learning model 510. For example, each training data entry may include the FAF spectrum 502 and the oxygen map 504 of the retina 412 as inputs, and the outputs 508a-508d determined by the human labeler and the metabolic state 512 determined by the human labeler as the desired outputs.

[0133] Thus, the training algorithm can process the FAF spectra 502 and oxygen maps 504 of the training data entries using the machine learning models 506a to 506d to obtain estimated outputs 508a to 508d that are compared to the outputs 508a to 508d of the training data entries. The training algorithm then updates the parameters of the machine learning models 506a to 506d based on the comparison.

[0134] The output machine learning model 510 can be trained by processing the outputs 508a to 508d from the training data entries using the machine learning model 510 to obtain an estimated metabolic state 512. The training algorithm then compares the estimated metabolic state 512 with the metabolic state 512 from the training data entries and updates the output machine learning model 510 based on the comparison.

[0135] Specific reference Figure 5B , the machine learning model 500b takes as input one or more images 514 of the retina 412. The one or more images 514 may include OCT images, MSI images, HSI images, FAF images, monochrome images or color images, infrared images, ultraviolet images, or other images of the retina 412.

[0136] One or more images 514 can be input to a plurality of machine learning models 516 to 516d, each of which generates a corresponding output 518a to 518d describing a particular metric, anatomical structure, or feature corresponding to a pathology. In the illustrated embodiment, the machine learning models 516a to 516d are convolutional neural networks (CNNs) 516a to 516d. However, the machine learning models 516a to 516d can also each be embodied as a deep neural network (DNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of machine learning models.

[0137] The outputs 518a - 518d may include, for example, representations of exudates, hemorrhages, drusen, lesions, or other anatomical structures or features corresponding to pathologies that may be represented in some or all of the one or more images 514 .

[0138] The outputs 518a to 518d may be input to an output machine learning model 520. In the illustrated embodiment, the output machine learning model 520 is a random forest. However, other machine learning models may be used, such as a long short-term memory (LSTM) machine learning model, a generative adversarial network (GAN) machine learning model, a logistic regression machine learning model, or other types of machine learning models.

[0139] The output machine learning model 520 outputs a diagnosis 522a-522c for each of the one or more pathologies. The output machine learning model 520 may additionally output a stage 524a-524c indicating a level of progression of the pathology for each of the one or more pathologies. For example, the pathologies may include diabetic retinopathy (DRP), age-related macular degeneration (AMD), glaucoma, or other pathologies that may be represented in an image of the retina 412. The stage 524a-524c of each pathology may be one of a set of discrete values ​​for each pathology, such as a number from 1 to 10, or other set of discrete values ​​used by health professionals to represent the progression of a pathology.

[0140] The representations of the diagnoses 522a to 522c and the stages 524a to 524c may be stored and / or output to a display device. The representations of some or all of the one or more images 514 (outputs 518a to 518d) may be stored and / or output to a display device for evaluation by an ophthalmologist, surgeon, or other health professional.

[0141] The training data entries may be used to train the machine learning models 516a-516d and the machine learning model 520. For example, each training data entry may include one or more images 514 of the retina 412 as inputs, and include outputs 518a-108d, diagnoses 522a-522c, and stages 524a-524c of the diagnoses 522a-522c as determined by a human labeler as desired outputs.

[0142] The training algorithm can process one or more images 514 of the training data entry using the machine learning model 516a to 516d to obtain estimated outputs 518a to 518d that are compared to the outputs 518a to 518d of the training data entry. The training algorithm then updates the parameters of the machine learning model 516a to 516d based on the comparison.

[0143] The output machine learning model 520 can be trained by processing the outputs 518a-518d from the training data entries using the output machine learning model 520 to obtain estimated diagnoses 522a-522c and stages 524a-524c. The training algorithm then compares the estimated diagnoses 522a-522c and stages 524a-524c with the diagnoses 522a-522c and stages 524a-524c from the training data entries and updates the output machine learning model 520 based on the comparison.

[0144] The systems and methods disclosed herein provide at least the following advantages:

[0145] Operators can use Figure 4A , Figure 4B or Figure 4C A single device can acquire image and spectral information sufficient to identify nearly all retinal pathologies in a single patient visit.

[0146] The associated machine learning model can also output labeled images and diagnoses based on those images during the same session in which the images were captured.

[0147] An associated machine learning model combines information from multiple imaging modalities to perform early detection of retinal pathology.

[0148] • The ease with which images and diagnostics can be obtained enables testing to be performed earlier and more frequently, which facilitates earlier detection and more frequent testing to be able to observe changes over time to assess the progression of the pathology.

[0149] Figure 6 An example computing system 600 is shown that implements, at least in part, one or more of the functions described herein. FIG. 4A to FIG. 4CThe imaging device of any of the figures may include a computing device having some or all of the attributes of the computing system 600. The computing device may be coupled to and receive images from one or more cameras 432, 436, 492. The computing device further receives the output of the spectrometer 428, 464 and / or the output of the OCT 402, wherein the OCT 402 includes a detector 408 in addition to the spectrometer 428. The computing device may control the operation of the systems 400a, 400b, 400c to capture images and switch between imaging modes as described above, including controlling some or all of the actuators of the OCT 402, any of the light sources 422, 484, one or more mirror actuators (e.g., actuators 418a, 480a, 468a), the scanning mirror 452, one or more actuators of the scanning lens 454, and the filters 430, 490 embodied as filter wheels or other devices capable of selecting among a plurality of filters.

[0150] Integrated into FIG. 4A to FIG. 4C The computing device in any of the imaging devices of FIG. 4 can further execute the machine learning model 438 and / or the machine learning models 500a, 500b. Alternatively, a different computing device having some or all of the properties of the computing system 600 can receive the image and spectral information from the imaging device and process the image and spectral information using the machine learning model 438 and / or the machine learning model 500a, 500b.

[0151] As shown, the computing system 600 includes a central processing unit (CPU) 602, one or more I / O device interfaces 604 that can allow various I / O devices 614 (e.g., keyboard, display, mouse device, pen input, etc.) to be connected to the computing system 600, a network interface 606 through which the computing system 600 is connected to a network 690, a memory 608, a storage device 610, and an interconnect 612.

[0152] Where computing system 600 is an imaging system, computing system 600 may further include one or more optical components for obtaining ophthalmic imaging of a patient's eye, as well as any other components known to one of ordinary skill in the art.

[0153] CPU 602 can retrieve and execute programming instructions stored in memory 608. Similarly, CPU 602 can retrieve and store application data residing in memory 608. Interconnect 612 transmits programming instructions and application data between CPU 602, I / O device interface 604, network interface 606, memory 608, and storage 610. CPU 602 is included to represent a single CPU, multiple CPUs, a single CPU with multiple processing cores, etc.

[0154] The memory 608 represents a volatile memory such as a random access memory and / or a non-volatile memory such as a non-volatile random access memory, a phase change random access memory, etc. As shown, the memory 608 can store a training algorithm 616, such as any of the training algorithms 202, 204, 206a, 206b, 210a, 210b, 212a, 212b described herein, or a training algorithm for training the machine learning models 500a, 500b as described above. The memory 608 can further store a machine learning model 618, such as any of the machine learning models 106a, 106b, 110, 114, 116, 438, 500a, 500b described herein.

[0155] The storage device 610 may be a non-volatile memory, such as a disk drive, a solid-state drive, or a collection of storage devices distributed across multiple storage systems. The storage device 610 may optionally store training data entries 620, such as the training data entries 200 or training data entries used to train the machine learning models 500a, 500b described above. The storage device 610 may store images 622 obtained according to any imaging mode described herein. The storage device 610 may store the results of processing the images using the machine learning model 618, including possibly storing intermediate results of any of the machine learning models in the machine learning model 618.

[0156] Note that the computing system 600 used to train the machine learning model 618 may be different from the computing system 600 that utilizes the trained machine learning model 618. Therefore, the machine learning model 618 and the results 624 of the machine learning model 618 may exist without the corresponding training algorithm 616 and training data entries 620.

[0157] Additional considerations

[0158] The foregoing description is provided to enable any person skilled in the art to practice the various embodiments described herein. Various modifications to these embodiments are clear to those skilled in the art, and the general principles defined herein may be applied to other embodiments. For example, without departing from the scope of the present disclosure, the functions and arrangements of the elements discussed may be changed. Various examples may appropriately omit, replace or add various programs or components. In addition, the features described with respect to some examples may be combined in some other examples. For example, any number of aspects set forth herein may be used to implement a device or practice method. In addition, the scope of the present disclosure is intended to cover such devices or methods practiced using other structures, functions, or structures and functions in addition to or different from the various aspects of the present disclosure set forth herein. It should be understood that any aspect of the present disclosure disclosed herein may be embodied by one or more elements of a claim.

[0159] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single members. For example, "at least one of a, b, or c" is intended to cover any combination of a, b, c, ab, ac, bc, and abc, as well as multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc or any other order of a, b, and c).

[0160] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, database, or other data structure), ascertaining, etc. In addition, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. In addition, "determining" may include parsing, selecting, choosing, establishing, etc.

[0161] The method disclosed herein includes one or more steps or actions for implementing the method. Without departing from the scope of the claims, the method steps and / or actions can be interchangeable with each other. In other words, unless a specific step or action sequence is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims. Further, the various operations of the above-mentioned method can be performed by any suitable device that can perform the corresponding function. These devices may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs) or processors. Generally, in the case of operations shown in the figure, those operations can have corresponding corresponding devices with similar numbers plus functional components.

[0162] The various illustrative logical blocks, modules, and circuits described in conjunction with the present disclosure may be implemented or executed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in an alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0163] The processing system can be implemented with a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnecting buses and bridges. The bus may link together various circuits including a processor, a machine-readable medium, and input / output devices. A user interface (e.g., a keypad, a display, a mouse, a joystick, etc.) may also be connected to the bus. The bus may also link various other circuits such as timing sources, peripherals, voltage regulators, power management circuits, etc., which are well known in the art and therefore will not be described further. The processor may be implemented with one or more general and / or special processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuit systems that can execute software. Those skilled in the art will recognize how to best implement the described functions of the processing system according to the specific application and the overall design constraints imposed on the entire system.

[0164] If implemented in software, the function may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium. Software should be broadly interpreted as instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or other. Computer-readable media include both computer storage media and communication media (such as any medium that facilitates the transfer of computer programs from one place to another). The processor may be responsible for managing the bus and general processing, including the execution of software modules stored on a computer-readable storage medium. A computer-readable storage medium may be connected to a processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative, the storage medium may be integrated into the processor. For example, a computer-readable medium may include a transmission line, a carrier modulated by data, and / or a computer-readable storage medium on which instructions separated from a wireless node are stored, all of which may be accessed by a processor via a bus interface. Alternatively or in addition, a computer-readable medium or any portion thereof may be integrated into a processor, such as a case where a cache and / or a general register file may be provided. For example, examples of machine-readable storage media may include RAM (random access memory), flash memory, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), registers, magnetic disks, optical disks, hard drives, or any other suitable storage media, or any combination thereof. Machine-readable media may be embodied in a computer program product.

[0165] A software module may include a single instruction or multiple instructions, and may be distributed over several different code segments, between different programs, and across multiple storage media. A computer-readable medium may include multiple software modules. A software module includes instructions that, when executed by a device such as a processor, cause a processing system to perform various functions. A software module may include a transmission module and a reception module. Each software module may exist in a single storage device, or may be distributed in multiple storage devices. For example, when a triggering event occurs, a software module may be loaded from a hard drive into a RAM. During the execution of a software module, a processor may load some instructions into a cache to increase access speed. Then, one or more cache lines may be loaded into a general register file for execution by the processor. When referring to the function of a software module, it should be understood that such function is implemented by the processor when executing instructions from the software module.

[0166] The following claims are not intended to be limited to the embodiments shown herein, but are given the full scope consistent with the language of the claims. In the claims, unless otherwise specified, reference to a singular element is not intended to mean "one and only one", but "one or more". Unless otherwise specifically stated, the term "some" refers to one or more. According to 35 U.S.C. § 112 (f), the elements of any claim will not be interpreted unless the phrase "device for..." is used to expressly describe these elements, or in the case of a method claim, the phrase "step for..." is used to describe these elements. All structural and functional equivalents of the elements of the various aspects described throughout the present disclosure that are known or will be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, whether or not such disclosure is explicitly described in the claims, the content disclosed herein is not intended to be dedicated to the public.

Claims

1. A system comprising: a first imaging device configured to capture a first image of a retina of a patient, the first image being an optical coherence tomography (OCT) image; as well as a second imaging device configured to capture a second image of the patient's retina according to an imaging modality other than OCT; Among them, at least one of the following: The sensor of the first imaging device is shared with the second imaging device; or The optical component is configured to select which of the first imaging device and the second imaging device receives light reflected from the retina to the first imaging device or the second imaging device.

2. The system of claim 1, wherein: The imaging modality other than OCT is fundus autofluorescence (FAF).

3. The system of claim 2, wherein: The sensor of the first imaging device is shared with the second imaging device, and the sensor is a spectrometer.

4. The system of claim 3, wherein: The optical component is a first actuated reflector; The first imaging device includes a first light source, and the second imaging device includes a second light source; and The first actuated mirror is further configured to select which of the first light source and the second light source is used to illuminate the retina.

5. The system of claim 4, further comprising a second actuated mirror configured to select between (a) allowing light from a first light source to enter the spectrometer and (b) allowing light from the second light source to enter the spectrometer.

6. The system of claim 1, wherein: The imaging modality other than OCT is a multispectral imaging camera.

7. The system of claim 1, wherein: The first imaging device is a spectral domain OCT (SD-OCT) device.

8. The system of claim 1, further comprising: One or more processing devices and one or more memory devices coupled to the one or more processing devices, the one or more memory devices storing executable code that, when executed by the one or more processing devices, causes the one or more processing devices to: receiving the first image and the second image; and The first image and the second image are processed using a machine learning model to obtain at least one of a diagnosis of a retinal pathology and an identification of a feature represented in at least one of the first image and the second image and corresponding to the retinal pathology.

9. The system of claim 1, further comprising: One or more processing devices and one or more memory devices coupled to the one or more processing devices, the one or more memory devices storing executable code that, when executed by the one or more processing devices, causes the one or more processing devices to: receiving the first image and the second image; processing the first image and the second image using a plurality of first machine learning models to obtain a plurality of outputs representing features of the retina; and The plurality of outputs are processed using a second machine learning model to obtain a diagnosis of retinal pathology.

10. The system of claim 9, wherein: The executable code, when executed by the one or more processing devices, further causes the one or more processing devices to process the plurality of outputs using the second machine learning model to obtain the stage of the retinal pathology.

11. The system of claim 9, wherein: The executable code, when executed by the one or more processing devices, further causes the one or more processing devices to identify a metabolic state of the retina.

12. The system of claim 11, further comprising a display, and wherein the executable code, when executed by the one or more processing devices, further causes the one or more processing devices to output a representation of the metabolic state of the retina to the display.

13. The system of claim 11, wherein: The metabolic state includes at least one of blood flow, oxygen saturation, or waste removal.

14. The system of claim 1, wherein: The first imaging device and the second imaging device are calibrated to include the same magnification.

15. The system of claim 1, wherein: The optical axes of the first imaging device and the second imaging device are aligned.