System for comprehensive analysis of multispectral imaging and optical coherence tomography
By using multiple input and output machine learning models in the system and processing feature maps in combination with intermediate machine learning models, the problem of difficulty in effectively diagnosing eye disorders in the prior art is solved, and accurate diagnosis and early disease recognition of eye disorders is achieved.
Patent Information
- Application Number
- CN202380068465.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-27
- Filing Date
- 2023-09-27
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively utilize multispectral imaging and optical coherence tomography techniques to diagnose ocular disorders, especially in combination with both techniques.
A system is adopted that includes processing devices and memory devices that process images of multiple imaging modes through multiple input machine learning models, and process feature maps in combination with intermediate machine learning models, and ultimately generate pathological estimate representations of the patient's eyes using an output machine learning model.
Comprehensive analysis of multispectral imaging and optical coherence tomography technology is achieved, which can identify biomarkers and conduct early disease diagnosis, improving the diagnostic accuracy of eye disorders.
Smart Images

Figure CN120019405A_ABST
Abstract
Description
Background Art
[0001] Multispectral imaging (MSI) is a technique that involves measuring (or capturing) light from a sample (e.g., ocular tissue / structure) at different wavelengths or spectral bands within the electromagnetic spectrum. MSI can capture more information from a sample that may not be visible through conventional imaging, which typically uses broadband illumination and broadband imaging sensors. The MSI information obtained by an MSI imaging system can be used to diagnose ocular disorders and enable real-time adjustments to the use of instruments (e.g., forceps, lasers, probes, etc.) used to manipulate ocular tissue / structure during surgical procedures.
[0002] Optical coherence tomography (OCT) is a technique that uses light waves to generate two-dimensional (2D) and three-dimensional (3D) images of the eye. 2D OCT can involve the use of time-domain OCT and / or Fourier domain OCT, which involves the use of spectral domain OCT methods and swept source OCT methods. 3D OCT can similarly utilize time-domain OCT imaging techniques and Fourier domain OCT imaging techniques. OCT imaging can also be used to diagnose eye disorders preoperatively or during surgery.
[0003] Better utilization of the capabilities of MSI and OCT to diagnose ocular disorders would be an advancement in the field. Summary of the invention
[0004] In certain embodiments, a system is provided. The system includes one or more processing devices and one or more memory devices, the one or more memory devices being connected to the one or more processing devices. The one or more memory devices store executable code, which when executed by the one or more processing devices causes the one or more processing devices to process one or more images according to each imaging mode for each of a plurality of imaging modes using an input machine learning model corresponding to each imaging mode in a plurality of input machine learning models to obtain an input feature map, the one or more images being images of a patient's eye. The system uses an intermediate machine learning model to process the feature maps of the plurality of imaging modes to obtain a final feature map. The final feature map is processed using one or more output machine learning models to obtain one or more estimated representations of the pathology of the patient's eye. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] In order to be able to understand the above features of the present disclosure in detail, the present disclosure briefly summarized above can be described in more detail with reference to the embodiments, some of which are shown in the accompanying drawings. However, it should be noted that the accompanying drawings only show exemplary embodiments and are therefore not to be considered as limiting the scope thereof, and other equally effective embodiments may be allowed.
[0006] Figure 1An example system for performing combined analysis of MSI images and OCT images to diagnose eye disorders in accordance with certain embodiments is presented.
[0007] Figure 2A is a diagram illustrating a first method for training a machine learning model to perform comprehensive analysis of MSI images and OCT images to diagnose eye disorders according to certain embodiments.
[0008] Figure 2B is a diagram illustrating a second method for training a machine learning model to perform comprehensive analysis of MSI images and OCT images to diagnose eye disorders according to certain embodiments.
[0009] Figure 2C is a diagram illustrating a third method for training a machine learning model to perform comprehensive analysis of MSI images and OCT images to diagnose eye disorders according to certain embodiments.
[0010] Figure 3 is a flowchart of a method for training a machine learning model to perform comprehensive analysis of MSI images and OCT images to diagnose eye disorders according to certain embodiments.
[0011] Figure 4 An example computing device that at least partially implements one or more functions for performing combined analysis of MSI images and OCT images in accordance with certain embodiments is presented.
[0012] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the figures. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further recitation. DETAILED DESCRIPTION
[0013] The various embodiments described herein provide a framework for using artificial intelligence to process information obtained from MSI images and OCT images. The advantage of MSI is that MSI images contain rich information about the retina over a wide spectral band range, and this information is a feature that cannot be seen using human vision or fundus cameras. The wide spectral band range of MSI further provides a high degree of depth penetration into the retina. However, MSI images do not provide structural information. In contrast, OCT images do provide structural information about the retina. However, interpreting OCT images requires a high degree of expertise. Using the methods described herein, the rich details and high depth penetration of MSI can be combined with the structural information of OCT to identify biomarkers for various pathologies and perform early disease diagnosis.
[0014] Figure 1A system 100 is shown for performing integrated analysis of an MSI image 102 and an OCT image 104. The system 100 may include three main stages: a feature extraction stage using machine learning models 106a, 106b, a feature enhancement stage using machine learning models 110, and a biomarker and prediction stage using machine learning models 114, 116. Through these three stages, the system 100 processes the MSI image 102 and the OCT image 104 separately for feature extraction and then combines the extracted features to obtain a meaningful interpretation.
[0015] The MSI image 102 may be captured using any method known in the art for performing MSI, including so-called hyperspectral imaging (HSI). Likewise, the OCT image 104 may be obtained using any method known in the art for performing OCT.
[0016] The MSI image 102 is obtained by illuminating the patient's eye using a multi-spectral band illumination source (e.g., a narrow-band illumination source, a narrow-band filter, etc.) and / or measuring reflected light using a multi-spectral band camera (e.g., an imaging sensor capable of sensing multiple spectral bands beyond the red, green, and blue (RGB) spectral bands). Therefore, each MSI image 102 represents reflected light within a specific spectral band. The differences between the MSI images 102 are due to the different reflectivities of different structures within the eye for different spectral bands. Therefore, when considered as a whole, the MSI image 102 provides more information about the structure of the eye than a single broadband image. In some embodiments, the MSI image 102 is a frontal image of the retina for detecting retinal pathology. However, MSI images 102 of other parts of the eye (such as the vitreous or anterior chamber) may also be used.
[0017] Optical coherence tomography (OCT) is a technique that uses light waves from a coherent light source (i.e., a laser) to generate two-dimensional (2D) and three-dimensional (3D) images of the eye. An OCT image is typically a cross-sectional image of the eye for a plane parallel to and colinear with the optical axis of the eye. However, a 3D image can be constructed using OCT images of multiple cross-sectional planes, from which a 2D image can be generated for a cross-sectional plane that is not parallel to the optical axis. For example, a frontal image of the retina can be obtained from the 3D image. In some embodiments, the OCT image 104 is such a frontal image of the retina. OCT is capable of imaging the retina to a certain depth, so that in some embodiments, the OCT image 104 is a collection of frontal images of image planes from at or above the surface of the retina to a depth within or below the retina.
[0018] Although the examples described herein relate to the use of MSI images 102 and OCT images 104, images from any pair of imaging modalities, or images from three or more different imaging modalities, may be used in a similar manner. For example, the additional imaging modalities may include a scanning laser ophthalmoscope (SLO), a fundus camera, and / or a broadband visible light camera.
[0019] In the system 100, the MSI image 102 is processed by a machine learning model 106a, and the OCT image 104 is processed by a machine learning model 106b. The machine learning models 106a, 106b may be implemented as a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of neural networks.
[0020] The result of the machine learning models 106a, 106b processing the images 102, 104 is feature maps 108a, 108b, respectively. For example, the feature maps 108a, 108b may be the output of one or more hidden layers of the machine learning models 106a, 106b. The feature maps 108a, 108b may be two-dimensional or three-dimensional arrays of values. In the case where the feature maps 108a, 108b are two-dimensional arrays, the feature maps 108a, 108b may include the same dimensions in two dimensions, or they may be different. In the case where one or both of the feature maps 108a, 108b are three-dimensional arrays, the feature maps 108a, 108b may include the same dimensions in at least two dimensions, or they may be different in any of the three dimensions.
[0021] The feature maps 108a, 108b and possibly the images 102, 104 are processed by a machine learning model 110. The machine learning model 110 may be implemented as a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of neural networks. The result of the machine learning model 110 processing the feature maps 108a, 108b and possibly the images 102, 104 is a feature map 112. For example, the feature map 112 may be the output of one or more hidden layers of the machine learning model 110, as discussed in more detail below.
[0022] The feature map 112 and possibly the images 102, 104 are then processed by a machine learning model 114 and a machine learning model 116, which then outputs one or more biometric segmentation maps 118 that label eye features represented in the images 102, 104 that correspond to one or more pathologies. Each biometric segmentation map 118 may be in the form of an image having the same dimensions as the images 102, 104, and in which non-zero pixels correspond to pixels in the images 102, 104 that are identified as corresponding to a particular pathology represented by the biometric segmentation map. The biometric segmentation map 118 may include a separate map for each of a plurality of pathologies, or a single map in which all pixels representing any of a plurality of pathologies are non-zero.
[0023] The machine learning model 114 can be implemented as a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), or other types of neural networks. For example, the machine learning model 114 can be implemented as a U-shaped network.
[0024] The machine learning model 116 outputs a disease diagnosis 120 and a possible severity score 122 corresponding to the disease diagnosis. The machine learning model 116 can be implemented as a long short-term memory (LSTM) machine learning model, a generative adversarial network (GAN) machine learning model, or other types of machine learning models. The disease diagnosis 120 can be output in the form of text naming the pathology, a numeric code corresponding to the pathology, or some other representation. The severity score 122 can be a numeric value, such as a value from 1 to 10 or a value within some other range. The severity score 122 can be limited to a set of discrete values (e.g., an integer from 1 to 10), or it can be any value within the precision limit of the number of digits used to represent the severity score 122.
[0025] Pathologies for which biometric segmentation maps 118 may be generated and for which diagnoses 120 and severity scores 122 may be generated include at least those pathologies that cause perceptible changes to the retina, such as at least the following:
[0026] Retinal tear
[0027] Retinal detachment
[0028] Diabetic retinopathy
[0029] Hypertensive retinopathy
[0030] Sickle cell retinopathy
[0031] Central retinal vein occlusion
[0032] Epiretinal membrane
[0033] Macular hole
[0034] Macular degeneration (including age-related macular degeneration)
[0035] Retinitis pigmentosa
[0036] ·glaucoma
[0037] Alzheimer's disease
[0038] Parkinson's disease
[0039] The biometric segmentation map 118 may, for example, mark vascular features corresponding to pathologies. Examples of vascular features that may be used to diagnose pathologies are described in the following references, both of which are incorporated herein by reference in their entirety:
[0040] Segmenting Retinal Vessels Using a Shallow Segmentation Network to Aid Ophthalmic Analysis, M. Arsalan et al., Mathematics, 2022, vol. 10, p. 1536.
[0041] PVBM: A Python Vasculature Biomarker Toolbox Based on Retinal BloodVessel Segmentation, J. Fhima et al., Cornell University (July 31, 2022).
[0042] Figure 2A Example methods for training machine learning models 106a, 106b, 110, 114, 116 are shown. In particular, Figure 2A A supervised machine learning method is shown using a plurality (e.g., hundreds, thousands, tens of thousands, hundreds of thousands, or more) of training data entries 200. Each training data entry 200 may include as input an MSI image 102 and an OCT image 104. Each of the MSI images 102 represents an image obtained by detecting light in a different spectral band relative to the other MSI images 102.
[0043] The MSI image 102 and the OCT image 104 images of the training data entry 200 may be images of the same eye of the patient and may be captured substantially simultaneously, such that the anatomical structures represented in the images 102, 104 are substantially the same. For example, "substantially simultaneously" may mean within 1 second and 1 hour of each other. However, "substantially simultaneously" may depend on the pathology being detected: those pathologies that progress very slowly may use images 102, 104 that were captured at a longer time difference (such as less than a day, less than a week, or some other time difference). The MSI image 102 and the OCT image 104 are preferably aligned and scaled relative to each other such that a given pixel coordinate in the MSI image 102 represents substantially the same location in the eye as the same pixel coordinate in the OCT image 104 (e.g., within 0.1 mm, within 1 μm, or within 0.01 μm). Such alignment and scaling may be achieved for the entire image 102, 104 or for at least a portion of one or both of the images 102, 104 showing the anatomical structure of interest (e.g., the macula of the retina).
[0044] The alignment and scaling of the images 102, 104 relative to each other may be achieved by aligning the optical axes of the instruments used to capture the images 102, 104 and calibrating the magnification of the instruments to achieve substantially the same (e.g., within + / - 0.1%, within 0.01%, or within 0.001%) scaling. Alternatively, the alignment and scaling of the images 102, 104 may be achieved by analyzing the anatomical structures represented in the images 102, 104. For example, where the MSI image 102 and the OCT image 104 represent the retina of an eye, the vascular pattern represented in each image 102, 104 may be used to align and scale one or both of the images 102, 104. Where the images 102, 104 have different sizes or do not completely overlap, such as after registration and scaling, non-overlapping portions of one or both of the images 102, 104 may be trimmed and / or one or both of the images 102, 104 may be padded so that the images 102, 104 have the same size and completely overlap each other.
[0045] Each training data entry 200 may include as desired outputs some or all of one or more biomarker segmentation maps 118, disease diagnoses 120, and severity scores 122. Multiple pathologies may be present in the same patient, such that a segmentation map 118, disease diagnoses 120, and severity scores 122 may be included for each pathology present or a subset of the most prevalent pathologies. The desired outputs are generated by a human expert based on an evaluation of the images 102, 104 and other health information of the patient that may be obtained before or after the images 102, 104 are captured. In particular, a biomarker segmentation map 118 for a pathology may include pixels in one or both of the images 102, 104 that are labeled by the human expert as corresponding to that pathology.
[0046] For each training data entry 200, the machine learning model 106a receives the MSI image 102 to generate one or more estimated biomarker segmentation maps. For example, the output of the machine learning model 106a can be a three-dimensional array in which each two-dimensional array along the third dimension is an estimated biometric segmentation map corresponding to a pathology.
[0047] The training algorithm 202 compares the one or more estimated biomarker segmentation maps to the one or more biomarker segmentation maps 118 of the training data entry 200. The training algorithm 202 then updates one or more parameters of the machine learning model 106a based on the difference between each estimated biomarker segmentation map of the pathology and the corresponding biomarker segmentation map 118 of the pathology in the training data entry 200.
[0048] The machine learning model 106b can be trained in a similar manner as the machine learning model 106a. For each training data entry 200, the machine learning model 106b receives the OCT image 104 and generates one or more estimated biomarker segmentation maps. For example, the output of the machine learning model 106a can be a three-dimensional array in which each two-dimensional array along the third dimension is an estimated biometric segmentation map corresponding to a pathology.
[0049] The training algorithm 202 (which may be the same or different than the training algorithm used to train the machine learning model 106a) compares the one or more estimated biomarker segmentation maps to the one or more biomarker segmentation maps 118 of the training data entry 200. The training algorithm 202 then updates one or more parameters of the machine learning model 106b based on the difference between each estimated biomarker segmentation map for the pathology and the corresponding biomarker segmentation map 118 for the pathology in the training data entry 200.
[0050] In the case of using three or more imaging modes, additional machine learning models may be present and trained in a similar manner. Each machine learning model corresponds to an imaging mode and processes the corresponding images of the imaging mode in the training data entry. The machine learning model generates one or more estimated biomarker segmentation maps, which are compared with one or more biomarker segmentation maps 118 of the training data entry by the training algorithm, and the training algorithm then updates the machine learning model based on the comparison. The hidden layer of each machine learning model can generate outputs, which are used as feature maps of the imaging mode corresponding to the machine learning model.
[0051] As mentioned above about Figure 1 As described, the machine learning model 110 takes the feature maps 108a, 108b of the machine learning models 106a, 106b as input. The machine learning models 106a, 106b may be trained with some or all of the training data entries 200 before training the machine learning model 110.
[0052] For each training data entry 200, the machine learning model 110 receives feature maps 108a, 108b obtained by processing the MSI image 102 and the OCT image 104 of the training data entry 200 with the machine learning models 106a, 106b. As described above, the feature maps 108a, 108b can be the output of the hidden layers (i.e., the layers other than the final layer that outputs one or more estimated biomarker segmentation maps) of the machine learning models 106a, 106b, respectively. The machine learning model 110 can also receive the MSI image 102 and the OCT image 104 as input, although in other embodiments, only the feature maps 108a, 108b are used.
[0053] The machine learning model 110 processes the feature maps 108a, 108b and possibly the MSI image 102 and the OCT image 104 and produces one or more estimated biomarker segmentation maps. For example, the output of the machine learning model 110 may be a three-dimensional array in which each two-dimensional array along the third dimension is an estimated biomarker segmentation map corresponding to a pathology. Although two feature maps 108a, 108b are shown, the machine learning model 110 may process any number of feature maps, and any number of possible images for generating feature maps, in a similar manner for any number of imaging modalities.
[0054] The training algorithm 204 compares the one or more estimated biomarker segmentation maps to the one or more biomarker segmentation maps 118 of the training data entry 200. The training algorithm 204 then updates one or more parameters of the machine learning model 110 based on the difference between each estimated biomarker segmentation map for the pathology and the corresponding biomarker segmentation map 118 for the pathology.
[0055] As mentioned above about Figure 1 As described, the machine learning models 114, 116 take as input the feature map 112 of the machine learning model 110. The machine learning model 110 may be trained with some or all of the training data entries 200 before training the machine learning models 114, 116.
[0056] For each training data entry 200, the machine learning model 114, 116 receives a feature map 112 obtained by processing the MSI image 102 and the OCT image 104 of the training data entry 200 with the machine learning model 106a, 106b, 110. As described above, the feature map 112 can be the output of a hidden layer (i.e., a layer other than the final layer that outputs one or more estimated biomarker segmentation maps) of the machine learning model 110. The machine learning model 114, 116 can also take the MSI image 102 and the OCT image 104 as input, although in other embodiments, only the feature map 112 is used.
[0057] The machine learning model 114 processes the feature maps 112 and possibly the images 102, 104 from the training data entries 200 and produces one or more estimated biomarker segmentation maps. In the case where three or more imaging modalities are used, images from the training data entries 200 according to the three or more imaging modalities may be processed by the machine learning model 114 along with the feature maps 112 obtained from these images. The output of the machine learning model 114 may be a three-dimensional array in which each two-dimensional array along the third dimension is an estimated biomarker segmentation map corresponding to a pathology.
[0058] The training algorithm 206a compares the one or more estimated biomarker segmentation maps to the one or more biomarker segmentation maps 118 of the training data entry 200. The training algorithm 206a then updates one or more parameters of the machine learning model 114 based on the difference between each estimated biomarker segmentation map of the pathology and the corresponding biomarker segmentation map 118 of the pathology.
[0059] The machine learning model 116 processes the feature map 112 and possibly the MSI image 102 and the OCT image 104 and generates one or more estimated diagnoses and an estimated severity score for each estimated diagnosis. Where three or more imaging modalities are used, images from the training data entry 200 according to the three or more imaging modalities may be processed by the machine learning model 116 along with the feature maps 112 obtained for those images.
[0060] The output of the machine learning model 116 may be a vector, where each element of the vector (if non-zero) indicates the estimated presence of a pathology. The output of the machine learning model 116 may also be text listing one or more major pathologies estimated to be present. The output of the machine learning model 116 may further include a severity score for each pathology estimated to be present, such as a vector in which each element corresponds to a pathology and the value of the element indicates the severity of the corresponding pathology.
[0061] The training algorithm 206b compares the estimated diagnosis and corresponding severity score to the disease diagnosis 120 and severity score 122 of the training data entry 200. The training algorithm 206b then updates one or more parameters of the machine learning model 116 based on the difference between the estimated diagnosis and corresponding severity score and the disease diagnosis 120 and severity score 122 of the training data entry 200.
[0062] refer to Figure 2B In some embodiments, training of one or both of the machine learning models 106a, 106b may be performed via unsupervised training algorithms 210a, 210b, respectively. Figure 2B While training using images 102 , 104 is shown, it should be understood that one or more machine learning models for additional or alternative imaging modalities may be trained in the same manner.
[0063] for Figure 2B In an embodiment, the machine learning models 110, 114, 116 may be as described above with respect to Figure 2A In some embodiments, only one of the machine learning models 106a, 106b is trained using the unsupervised training algorithm 210a, 210b, while the Figure 2A The described supervised training algorithm 202 is used to train another machine learning model. For the unsupervised machine learning algorithms 210a, 210b, no labeled training data entries are used.
[0064] The machine learning model 106a may be trained using a corpus of a collection of MSI images 102. The corpus may be curated to include a large set of MSI images (e.g., retinal images) of healthy eyes that are free of pathology, as well as a small portion (e.g., less than 5% of the corpus or less than 1 percent of the corpus) that correspond to one or more pathologies. The collection of MSI images 102 may or may not be labeled with whether the collection of images 102 represents a pathology and / or the specific pathology represented.
[0065] The unsupervised training algorithm 210a uses the machine learning model 106a to process the corpus and train the machine learning model 106a to identify and classify anomalies detected in the collection of MSI images 102 in the corpus. The unsupervised training algorithm 210a may be implemented using any method for performing anomaly detection or other unsupervised machine learning known in the art. The output of the machine learning model 106a may be an image of the same size as the individual MSI images 102, with pixels representing anomalies labeled.
[0066] The machine learning model 106b may be trained using a corpus of OCT images 104. The corpus may be curated to include a large number of OCT images of healthy eyes (e.g., retinal images) that are free of pathology, as well as a small portion (e.g., less than 5% of the corpus or less than 1 percent of the corpus) that correspond to one or more pathologies. The OCT images 104 may or may not be labeled with whether the set of images 102 represents a pathology and / or the specific pathology represented.
[0067] The unsupervised training algorithm 210b uses the machine learning model 106b to process the corpus and train the machine learning model 106a to identify and classify abnormalities detected in the OCT images 104 of the corpus. The unsupervised training algorithm 210b can be implemented using any method for performing anomaly detection or other unsupervised machine learning known in the art. The output of the machine learning model 106b can be an image of the same size as each OCT image 104, in which pixels representing abnormalities are labeled.
[0068] The set of MSI images 102 and OCT images 104 used to train the machine learning model by the unsupervised training algorithm 210a, 210b may include images 102, 104 from the training data entry 200 used to train other machine learning models 110, 114, 116. The set of MSI images 102 and OCT images 104 can be further enhanced with images of healthy eyes to facilitate the identification of abnormalities corresponding to pathology. The set of MSI images 102 and OCT images 104 can be constrained to be the same size and can be aligned with each other. For example, although the images 102, 104 are images of multiple different eyes, the images 102, 104 can be aligned to place the representation of the center of the fovea of the retina substantially at the center of each image 102, 104 (e.g., within 1, 2, or 3 pixels). Certain other features (such as the fundus) can be used for alignment. Although the images 102, 104 are images of multiple different eyes, the images 102, 104 can also be scaled so that the size of the anatomical structures represented in the images is substantially the same. For example, the images 102, 104 may be scaled so that the fovea, fundus, or one or more other anatomical features are the same size.
[0069] Once trained, the machine learning models 106a, 106b can provide information to the machine learning model 110 (see Figure 1 and Figure 2A ) provide outputs, which are the outputs of one or more hidden layers of the machine learning models 106a, 106b. Alternatively, the final outputs of the machine learning models 106a, 106b (e.g., images with abnormal labels) can be used as inputs to the machine learning model 110.
[0070] refer to Figure 2C , in the right Figure 2B In an improvement of the unsupervised machine learning approach, the supervised training algorithm 212b may compare the output of the machine learning model 106a with the output of the machine learning model 106b for a given set of MSI images 102 and OCT images 104 of the same patient's eye captured substantially simultaneously as defined above. The supervised training algorithm 212b may then adjust the parameters of the machine learning model 106b based on the comparison so as to train the machine learning model 106b to recognize the same abnormalities detected by the machine learning model 106a. Note that the reverse approach may alternatively or additionally be used: the output of the machine learning model 106b may be used by the supervised training algorithm 212b or a different supervised training algorithm 212a to train the machine learning model 106a to recognize the abnormalities identified by the machine learning model 106b.
[0071] In some embodiments, training can be performed in stages, each stage using the above Figure 2A , Figure 2B and Figure 2C In the first example, we first use Figure 2A The supervised machine learning method of Figure 2B The unsupervised method is used to train the machine learning models 106a, 106b; and then, according to Figure 2C The method further trains the machine learning model 106b based on the output of the machine learning model 106a (and / or vice versa). In the second example, only unsupervised learning is used: using Figure 2B The unsupervised method of training machine learning models 106a and 106b separately is then used to Figure 2C The method further trains the machine learning model 106b based on the output of the machine learning model 106a, and / or trains the machine learning model 106a based on the output of the machine learning model 106b.
[0072] Figure 2C While training using images 102, 104 is shown, it should be understood that one or more machine learning models for additional or alternative imaging modalities may be trained in the same manner. In particular, the output of a machine learning model according to one imaging modality may be used to train one or more other machine learning models according to one or more other imaging modalities in the same manner. The outputs of two or more first machine learning models of one or more first imaging modalities may be concatenated or otherwise combined and used to train one or more other machine learning models according to one or more other imaging modalities. Figure 2C Method to train one or more second machine learning models of one or more second machine learning models.
[0073] refer to Figure 3 The method 300 shown may be performed by a computer system (e.g. Figure 4 The method 300 includes, in step 302, training a first input machine learning model using training images of a first imaging mode. For example, step 302 may include the following steps: FIG. 2A to FIG. 2C A machine learning model 106a trained using MCI images 102 of any of the methods described.
[0074] Method 300 includes, in step 304, training a second input machine learning model using training images of a second imaging mode. For example, step 304 may include training a second input machine learning model using training images of a second imaging mode. FIG. 2A to FIG. 2C A machine learning model 106b trained using OCT images 104 of any of the methods described.
[0075] Method 300 includes, in step 306, processing the image according to the first imaging mode using the first input machine learning model to obtain an input feature map F1, and processing the image according to the second imaging mode using the second input machine learning model to obtain an input feature map F2. Feature maps F1 and F2 may be outputs of hidden layers of the first input machine learning model and the second input machine learning model, respectively. Step 306 may include processing the MSI image 102 using the machine learning model 106a and processing the OCT image 104 using the machine learning model 106b to obtain feature maps 108a, 108b, as described above with respect to Figure 1 and Figure 2A As described above, the MSI image 102 and the OCT image 104 may be part of a common training data entry 200, such that the MSI image 102 and the OCT image 104 are from the eye of the same patient and are captured substantially simultaneously.
[0076] The method 300 includes, in step 308, training an intermediate machine learning model using the feature maps F1 and F2. Specifically, multiple pairs of feature maps F1 and F2 can each be processed by the intermediate machine learning model, and the output of the intermediate machine learning model can be used to train the intermediate machine learning model. Each pair of feature maps F1 and F2 can be obtained for images of the first mode and the second mode, which are from the eyes of the same patient and are captured substantially at the same time. Step 308 may include using the intermediate machine learning model to process the images used to obtain each pair of feature maps F1 and F2. Step 308 may include training the machine learning model 110 using the feature maps 108a, 108b and the training data entries 200, as described above with respect to Figure 2A as described.
[0077] The method 300 includes, in step 310, processing the feature pairs of feature maps F1 and F2 and possible training images used to obtain each pair of feature maps F1 and F2 using the intermediate machine learning model to obtain a final feature map F. The final feature map F can be obtained from the output of the hidden layer of the intermediate machine learning model. Step 310 can include using the machine learning model 110 to process the feature maps 108a, 108b and possible corresponding images 102, 104 to obtain the feature map 112, as described above with respect to Figure 1 and Figure 2A as described.
[0078] The method 300 includes, in step 312, training one or more output machine learning models using the feature map F. The one or more output machine learning models may be trained to output, for a given feature map F, estimated representations of pathologies represented in training images used to generate the feature map F using the first input machine learning model and the second input machine learning model and the intermediate machine learning model. The one or more output machine learning models may take as input images used to generate the feature map F according to the first imaging mode and the second imaging mode. Step 312 may include training one or both of the machine learning models 114, 116 using the feature map 112 and possibly corresponding images 102, 104 to output some or all of the biomarker segmentation map 118, the disease diagnosis 120, and the severity score 122.
[0079] Method 300 may include, in step 314, processing utilization images according to the first imaging mode and the second imaging mode according to a pipeline of a first input machine learning model and a second input machine learning model, an intermediate machine learning model, and one or more output machine learning models. Specifically, one or more of the utilization images are processed according to the first imaging mode using the first input machine learning model to obtain a feature map F1; one or more of the utilization images are processed according to the second imaging mode using the second input machine learning model to obtain a feature map F2; feature maps F1 and F2 and possible utilization images are processed using the intermediate machine learning model to obtain a feature map F; and feature map F and possible utilization images are processed by one or more output machine learning models to obtain an estimated representation of the pathology represented in the utilization image. The estimated representation may be output to a display device or stored in a storage device for later use or subsequent processing. The feature maps (F1, F2, F) may be displayed or stored additionally.
[0080] For example, step 314 may include processing the utilization images 102, 104 (i.e., images 102, 104 that are not part of the training data entry 200) using the machine learning models 106a, 106b, respectively, to obtain feature maps 108a, 108b, respectively, as described above with respect to Figure 1 The feature maps 108a, 108b and possibly the images 102, 104 may be processed using a machine learning model 110 to obtain a feature map 112. The feature map 112 and possibly the images 102, 104 may be processed by one or both of the machine learning models 114, 116 to obtain a biomarker segmentation map 118, a disease diagnosis 120, and a severity score 122.
[0081] Steps 302 to 314 may be performed sequentially, i.e., the first input machine learning model and the second input machine learning model are trained, then the intermediate machine learning model is trained, then one or more output machine learning models are trained, and finally the utilization image is trained. Steps 302 to 314 may additionally or alternatively be interleaved, i.e., the first input machine learning model and the second input machine learning model, the intermediate machine learning model, and the one or more output machine learning models are trained as a group. For example, in a first stage, the first input machine learning model and the second input machine learning model, the intermediate machine learning model, and the one or more output machine learning models are trained individually in the order listed, and in a second stage, the training continues as a group, i.e., after an iteration including processing a set of images according to a pipeline, some or all of the first input machine learning model and the second input machine learning model, the intermediate machine learning model, and the one or more output machine learning models may be updated as part of the iteration by a training algorithm according to the outputs of the first input machine learning model and the second input machine learning model, the intermediate machine learning model, and the one or more output machine learning models, respectively. Training individually or as a group may be performed in utilizing step 314 (particularly as described with respect to Figure 2B and / or Figure 2C The unsupervised learning described above continues during the
[0082] Step 314 may be performed by a different computer system than the computer system used to perform steps 302 to 312. For example, the pipeline including the first output machine learning model, the second output machine learning model, the third output machine learning model, and the one or more output machine learning models may be installed on one or more other computer systems for use by a surgeon or other medical professional.
[0083] Although method 300 is described with respect to two imaging modes, three or more imaging modes may be used in a similar manner. i , i=1 to N, where N is greater than or equal to two. For a given training data entry, the imaging mode IM i The images can meet the requirements of substantially the same scaling, alignment, and simultaneous imaging of the same eye of the patient. There may be an input machine learning model ML i , i = 1 to N, each machine learning model ML i Corresponding to imaging mode IM i , and each machine learning model processes the imaging mode IM i One or more corresponding images are used to generate the corresponding feature map F i . The corresponding imaging mode IM iThe images are trained according to any of the methods described above for training machine learning models 106a, 106b. i .
[0084] Therefore, in this embodiment, the intermediate machine learning model converts the N feature maps F i (i=1 to N) and used to generate the feature map F i The possible training images are taken as input. The output machine learning model will generate the final feature map F and the i The possible training images are taken as input. The intermediate machine learning model and the output machine learning model are trained as described above with respect to the machine learning model 110 and the machine learning models 114 and 116.
[0085] Figure 4 Demonstrates at least partial implementation of the invention Figures 1 to 3 An example computing system 400 that can implement one or more of the functions described herein. The computing system 400 can be integrated with an imaging device that captures images according to one or more of the imaging modes described herein, or can be a separate computing device.
[0086] As shown, the computing system 400 includes a central processing unit (CPU) 402, one or more I / O device interfaces 404 that can allow various I / O devices 414 (e.g., keyboard, display, mouse device, pen input, etc.) to be connected to the computing system 400, a network interface 406 through which the computing system 400 is connected to a network 490, a memory 408, a storage device 410, and an interconnect 412.
[0087] Where computing system 400 is an imaging system, such as an SLO, OCT, or fundus camera, computing system 400 may further include one or more optical components for obtaining ophthalmic imaging of the patient's eye, as well as any other components known to those of ordinary skill in the art.
[0088] CPU 402 can retrieve and execute programming instructions stored in memory 408. Similarly, CPU 402 can retrieve and store application data residing in memory 408. Interconnect 412 transmits programming instructions and application data between CPU 402, I / O device interface 404, network interface 406, memory 408, and storage device 410. CPU 402 is included to represent a single CPU, multiple CPUs, a single CPU with multiple processing cores, etc.
[0089] The memory 408 represents a volatile memory such as a random access memory and / or a non-volatile memory such as a non-volatile random access memory, a phase change random access memory, etc. As shown, the memory 408 can store a training algorithm 416, such as any of the training algorithms 202, 204, 206a, 206b, 210a, 210b, 212a, 212b described herein. The memory 408 can further store a machine learning model 418, such as any of the machine learning models 106a, 106b, 110, 114, 116 described herein.
[0090] The storage device 410 may be a non-volatile memory, such as a disk drive, a solid-state drive, or a collection of storage devices distributed across multiple storage systems. The storage device 410 may optionally store the training data entries 200 or other collections of MSI images 102 and OCT images 104 for training and / or utilization according to the systems and methods described herein. The storage device 410 may optionally store intermediate results processed by any of the machine learning models 106a, 106b, 110, 114, 116, such as feature maps 108a, 108b, 112.
[0091] Additional considerations
[0092] The foregoing description is provided to enable any person skilled in the art to practice the various embodiments described herein. Various modifications to these embodiments are clear to those skilled in the art, and the general principles defined herein may be applied to other embodiments. For example, without departing from the scope of the present disclosure, the functions and arrangements of the elements discussed may be changed. Various examples may appropriately omit, replace or add various programs or components. In addition, the features described with respect to some examples may be combined in some other examples. For example, any number of aspects set forth herein may be used to implement a device or practice a method. In addition, the scope of the present disclosure is intended to cover such devices or methods practiced using other structures, functions, or structures and functions in addition to or different from the various aspects of the present disclosure set forth herein. It should be understood that any aspect of the present disclosure disclosed herein may be embodied by one or more elements of a claim.
[0093] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single members. For example, "at least one of a, b, or c" is intended to cover any combination of a, b, c, ab, ac, bc, and abc, as well as multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc or any other order of a, b, and c).
[0094] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, database, or other data structure), ascertaining, etc. In addition, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. In addition, "determining" may include parsing, selecting, choosing, establishing, etc.
[0095] The method disclosed herein includes one or more steps or actions for implementing the method. Without departing from the scope of the claims, the method steps and / or actions can be interchangeable with each other. In other words, unless a specific step or action sequence is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims. Further, the various operations of the above-mentioned method can be performed by any suitable device capable of performing the corresponding function. These devices may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs) or processors. Typically, where operations are shown in the figure, those operations may have corresponding corresponding devices with similar numbers plus functional components.
[0096] The various illustrative logical blocks, modules, and circuits described in conjunction with the present disclosure may be implemented or executed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in an alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0097] The processing system can be implemented with a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnecting buses and bridges. The bus may link together various circuits including a processor, a machine-readable medium, and input / output devices. A user interface (e.g., a keypad, a display, a mouse, a joystick, etc.) may also be connected to the bus. The bus may also link various other circuits such as timing sources, peripherals, voltage regulators, power management circuits, etc., which are well known in the art and therefore will not be described further. The processor may be implemented with one or more general and / or special processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuit systems that can execute software. Those skilled in the art will recognize how to best implement the described functions of the processing system according to the specific application and the overall design constraints imposed on the entire system.
[0098] If implemented in software, the function may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium. Software should be broadly interpreted as instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or other. Computer-readable media include both computer storage media and communication media (such as any medium that facilitates the transfer of computer programs from one place to another). The processor may be responsible for managing the bus and general processing, including the execution of software modules stored on a computer-readable storage medium. A computer-readable storage medium may be connected to a processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative, the storage medium may be integrated into the processor. For example, a computer-readable medium may include a transmission line, a carrier modulated by data, and / or a computer-readable storage medium on which instructions separated from a wireless node are stored, all of which may be accessed by a processor via a bus interface. Alternatively or in addition, a computer-readable medium or any portion thereof may be integrated into a processor, such as a case where a cache and / or a general register file may be provided. For example, examples of machine-readable storage media may include RAM (random access memory), flash memory, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), registers, magnetic disks, optical disks, hard drives, or any other suitable storage media, or any combination thereof. Machine-readable media may be embodied in a computer program product.
[0099] A software module may include a single instruction or multiple instructions, and may be distributed over several different code segments, between different programs, and across multiple storage media. A computer-readable medium may include multiple software modules. A software module includes instructions that, when executed by a device such as a processor, cause a processing system to perform various functions. A software module may include a transmission module and a reception module. Each software module may exist in a single storage device, or may be distributed in multiple storage devices. For example, when a triggering event occurs, a software module may be loaded from a hard drive into a RAM. During the execution of a software module, a processor may load some instructions into a cache to increase access speed. Then, one or more cache lines may be loaded into a general register file for execution by the processor. When referring to the function of a software module, it should be understood that such function is implemented by the processor when executing instructions from the software module.
[0100] The following claims are not intended to be limited to the embodiments shown herein, but are given the full scope consistent with the language of the claims. In the claims, unless otherwise specified, reference to a singular element is not intended to mean "one and only one", but "one or more". Unless otherwise specifically stated, the term "some" refers to one or more. According to 35 U.S.C. § 112 (f), the elements of any claim will not be interpreted unless the phrase "device for..." is used to expressly describe these elements, or in the case of a method claim, the phrase "step for..." is used to describe these elements. All structural and functional equivalents of the elements of the various aspects described throughout this disclosure that are known or will be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, whether or not such disclosure is explicitly described in the claims, the content disclosed herein is not intended to be dedicated to the public.
Claims
1. A system comprising: One or more processing devices and one or more memory devices coupled to the one or more processing devices, the one or more memory devices storing executable code that, when executed by the one or more processing devices, causes the one or more processing devices to: For each of the multiple imaging modes: processing one or more images according to each imaging modality using an input machine learning model corresponding to each imaging modality from a plurality of input machine learning models to obtain an input feature map, the one or more images being images of an eye of a patient; Processing the input feature maps of the plurality of imaging modalities using an intermediate machine learning model to obtain a final feature map; and The final feature map is processed using one or more output machine learning models to obtain one or more estimated representations of pathology of the patient's eye, the one or more estimated representations of pathology of the patient's eye comprising a diagnosis of retinal tear and a severity score for the diagnosis.
2. A system comprising: One or more processing devices and one or more memory devices coupled to the one or more processing devices, the one or more memory devices storing executable code that, when executed by the one or more processing devices, causes the one or more processing devices to: For each of the multiple imaging modes: processing one or more images according to each imaging modality using an input machine learning model corresponding to each imaging modality from a plurality of input machine learning models to obtain an input feature map, the one or more images being images of an eye of a patient; Processing the input feature maps of the plurality of imaging modalities using an intermediate machine learning model to obtain a final feature map; and The final feature map is processed using one or more output machine learning models to obtain one or more estimated representations of the pathology of the patient's eye.
3. The system of claim 2, wherein: The plurality of imaging modalities includes at least one of multi-spectral imaging (MSI) or optical coherence tomography (OCT).
4. The system of claim 2, wherein: The multiple imaging modalities include multispectral imaging (MSI) and optical coherence tomography (OCT).
5. The system of claim 2, wherein: Each input machine learning model is one of a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), and an autoencoder (AE).
6. The system of claim 2, wherein: The input feature map for each imaging mode is the output of the hidden layer of the input machine learning model for each imaging mode.
7. The system of claim 2, wherein: The intermediate machine learning model is one of a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), and an autoencoder (AE).
8. The system of claim 2, wherein: The final feature map is the output of the hidden layer of the intermediate machine learning model.
9. The system of claim 2, wherein: The one or more output machine learning models are one of a neural network, a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a region-based CNN (R-CNN), an autoencoder (AE), a long short-term memory (LSTM) machine learning model, and a generative adversarial network (GAN) machine learning model.
10. The system of claim 2, wherein: The one or more estimated representations of a pathology of the patient's eye include a diagnosis of the pathology.
11. The system of claim 10, wherein: The one or more estimated representations of pathology in the patient's eye comprise a severity score for the diagnosis.
12. The system of claim 2, wherein: The one or more estimated representations of pathology of the patient's eye include one or more biomarker segmentation maps.