Use of multiple subvolumes, thicknesses, and curvatures for OCT / OCTA data alignment and retinal landmark detection.
By using multiple en-face and vascular maps with machine learning and curvature maps, the method enhances the alignment of OCT datasets, addressing challenges in pathological changes and data quality for accurate disease progression analysis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CARL ZEISS MEDITEC INC
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-01
AI Technical Summary
Aligning multiple pairs of optical coherence tomography (OCT) datasets acquired at different times is challenging due to pathological changes, segmentation errors, variations in data quality, and subtle changes over time, which affect the accuracy of landmark detection and alignment.
Utilize multiple pairs of en-face and vascular maps derived from OCT volumes to identify a broader distribution of unique features across the 2D space, employing machine learning techniques and curvature maps to enhance landmark matching and alignment, even with low-quality data.
Improves the reliability and accuracy of aligning OCT datasets by identifying more consistent unique features, overcoming issues with low-quality data and pathological changes, leading to better disease progression analysis.
Smart Images

Figure 2026074063000001_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to an optical coherence tomography (OCT) system. More particularly, the present invention relates to the alignment of corresponding OCT data sets.
Background Art
[0002] Optical coherence tomography (OCT) has become an important modality for eye examinations. The two-dimensional (2D) representation of three-dimensional (3D) OCT volume data is one of the OCT visualization techniques that has benefited significantly from technological advancements in OCT technology. Examples of 2D representations of 3D OCT volume data (e.g., 2D maps) can include, by way of example and not limitation, layer thickness maps, retinal curvature maps, OCT en face images (e.g., en face structural images), and OCTA vascular system maps (e.g., en face OCT angiography, or functional OCT images).
[0003] Multilayer segmentation is often used to generate retinal layer thickness maps, en face images, and (2D) vascular system maps. Thickness maps are based on the measured differences in thickness between retinal layer boundaries. Vascular system maps and OCT en face images can be generated, for example, by projecting subvolumes between two layer boundaries using means such as average, sum, percentile, etc. Thus, the creation of these 2D maps (or 2D representations of 3D volumes or subvolumes) often depends on the effectiveness of an automatic segmentation algorithm for identifying the layers that form the basis of the 2D map.
[0004] To track disease progression over time, it is desirable to compare two corresponding OCT datasets (e.g., OCT volume and / or two corresponding 2D maps) of the same tissue taken at different times, for example, at different examinations. This requires measuring changes between two corresponding OCT volumes and / or 2D maps (e.g., thickness map, en-face image, vascular map, etc.) over time, which may require time-series data alignment (obtained from the same subject). However, aligning two corresponding OCT datasets can be a challenging task due to factors such as pathological changes in the retina between examinations, layer segmentation errors in the OCT data from one or both examinations due to severe pathology which may result in partially inaccurate 2D maps that cannot be precisely aligned, pathological changes affecting one or more subretinal volume data associated with a specific disease (e.g., superficial and deeper retinal layers may be affected by diabetic retinopathy), variations in the quality of OCT data between examinations (e.g., different OCT systems or different OCT techniques / modalities with different imaging quality used at different examinations, or inconsistent imaging conditions), or large lateral or other types of motion. Even if the above problems are avoided, subtle changes over time can be difficult to detect, especially in the short term, as the magnitude of the changes can be very small.
[0005] Effective temporal comparison of two OCT datasets directly depends on the effectiveness of the alignment (e.g., alignment of corresponding parts) of the two OCT datasets. Therefore, the alignment problem mainly involves aligning corresponding landmarks (e.g., characteristic features) so that they are confirmed / identifiable in a pair of (corresponding) 2D maps (e.g., a pair of thickness maps, a pair of en-face images, a pair of vascular maps). Landmark matching can be a difficult problem due to the temporal changes in images as described above.
[0006] Landmark-based alignment can work if a sufficient number of well-distributed landmark matches are found across each image (e.g., each of a pair of OCT datasets) and an appropriate transformation model is selected. However, one or more of the problems described can affect the quality of the OCT dataset, which can result in an insufficient number of identifiable landmarks (e.g., unique features) or identifiable landmarks that are not well-distributed across the 2D map. For example, if a portion of an en-face image is of low quality, it is likely that there are no identifiable landmarks (or only an insufficient number of landmarks) within that low-quality portion, leading to alignment failure of the OCT dataset from which the en-face image was generated.
[0007] Current methods for addressing the alignment problem primarily focus on improving the quality of individual en faces or vascular maps, such as by forming a single en face based on a combination of multiple en faces or vascular maps. Thus, the alignment of two OCT datasets is still based on the alignment of a single en face or vascular map. Examples of this method can be found in Non-Patent Documents 1 and 2. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] Andrew Lang et al., "Combined Registration and Motion Correction of Longitudinal Retinal OCT Data," Proceedings of the SPIE International Society for Optical Engineering, February 27, 2016, p. 9784. [Non-Patent Document 2] Jing Wu et al., "Stable Registration of Pathological 3D-OCT Scans Using Retinal Vessels," Proceedings of the Ophthalmic Medical Image Analysis International Workshop, September 14, 2014. [Overview of the project] [Problems that the invention aims to solve]
[0009] The object of the present invention is to provide a method / system for improving the reliability of aligning multiple pairs of corresponding OCT / OCTA datasets that may be acquired at different times.
[0010] Another objective of the present invention is to improve the alignment of low-quality corresponding OCT / OCTA datasets. [Means for solving the problem]
[0011] The above objectives are achieved in an OCT method / system that applies one or more techniques to improve the quality of the alignment. The present invention improves the alignment of two corresponding OCT volumes (e.g., OCT datasets) by better utilizing more available information, particularly when the OCT volume is of low quality, or when the tissue is necrotic, or when there is any of the other problems described above. The applicants noted that the prior art method of aligning OCT volumes by always using the same target slab (e.g., a subvolume bounded by the same two target layers) can affect the accuracy of the alignment because the above problems can disproportionately affect one or both of these target layers compared to other layers in the OCT volume. In contrast, the main idea of the present method is to use landmark matching of multiple pairs of images created from the two OCT volumes to be aligned (e.g., multiple 2D maps formed from multiple corresponding subvolumes of the two OCT volumes). The images to be aligned may be an OCT en face or an OCTA vascular map, or a combination of an OCT en face and an OCTA vascular map, or a thickness map, or other types of 2D maps. Landmarks can be detected in a pair of 3D OCT / OCTA volumes, and the XY position of the landmark match (such as determined from a aligned 2D map) can be used for alignment.
[0012] In addition, or alternatively, multiple separate 2D maps can also be derived from a single (e.g., parent or base) 2D map (e.g., a macular thickness map). For example, multiple curvature maps (e.g., derived maps or one / multiple child maps) can be derived from each thickness map (e.g., a base map). The thickness maps and the multiple curvature maps corresponding to the thickness maps can then be used for landmark matching (e.g., alignment) to define multiple (alignment) transformation parameters / features (or transformation matrices) that define a custom transformation model for aligning two OCT datasets.
[0013] The applicants also found that multiple maps derived from a single (base or parent) map (e.g., multiple curvature maps derived from a macular thickness map) can be used to identify physiological landmarks of the retina that would otherwise be undetectable in OCT data with low image quality, necrotic tissue, or numerous image artifacts (e.g., motion artifacts). For example, low-quality OCT volumes may produce low-quality en-face or vascular images that are generally unsuitable for detecting target physiological landmarks such as the fovea. However, good quality thickness maps can be obtained even from such low-quality OCT volumes, and multiple curvature maps derived from the thickness map (e.g., each curvature map is based on a different curvature characteristic of the thickness map) can help identify the fovea. That is, the (base or parent) thickness map and the (derived or multiple child) curvature maps corresponding to the thickness map can be used, for example, with various machine learning models / techniques for foveal location estimation.
[0014] Other objects and achievements of the present invention will become apparent and understandable by referring to the following description and claims, which are to be interpreted in conjunction with the accompanying drawings, along with a more thorough understanding of the present invention.
[0015] To facilitate understanding of the present invention, several publications are cited or referenced herein. All publications cited or referenced herein are incorporated herein in their entirety by reference.
[0016] The embodiments disclosed herein are illustrative and the scope of this disclosure is not limited thereto. Any feature of any embodiment described in one claim category, for example, a system, can also be claimed in another claim category, for example, a method. Dependencies or backreferences in the supplementary claims are selected for formal reasons only. However, any subject matter derived from careful backreferences to the earlier claims can also be claimed, thereby disclosing any combination of claims and their features, regardless of the dependencies selected in the supplementary claims. [Brief explanation of the drawing]
[0017] A patent or application file includes at least one drawing made in color. A copy of the published patent or patent application containing the color drawing(s) will be provided by the Patent Office upon request and payment of the required fees.
[0018] In drawings, the same reference number / letter refers to the same component. [Figure 1] This figure provides an example of aligning OCT datasets of multiple corresponding pairs of patients with age-related macular degeneration (AMD) acquired at two different consultations. [Figure 2] This figure provides an example of aligning OCT datasets of multiple corresponding pairs of age-related macular degeneration (AMD) patients, taken at two different consultations. [Figure 3] This figure provides an example of aligning OCT datasets of multiple corresponding pairs of age-related macular degeneration (AMD) patients, taken at two different consultations. [Figure 4]A figure providing a comparison between B-scan and en face images taken with a high-quality OCT system and those taken with a low-cost line-field SD-OCT. [Figure 5] A figure showing a plurality of pairs of thickness maps that can be used for alignment purposes. [Figure 6] A figure showing an example of landmark matching between two thickness maps. [Figure 7] A figure showing one pair of thickness maps and a plurality of curvature maps corresponding to one pair of thickness maps. [Figure 8A] A figure showing two exemplary AMD cases from two examinations using a Cirrus (registered trademark) device. [Figure 8B] A figure showing two exemplary AMD cases from two examinations using a Cirrus (registered trademark) device. [Figure 9A] A figure providing two examples each of a normal case and a central serous retinopathy (CSR) case from the same examination using a low-cost line-field SD-OCT device. [Figure 9B] A figure providing two examples each of a normal case and a central serous retinopathy (CSR) case from the same examination using a low-cost line-field SD-OCT device. [Figure 10] A figure showing a macular thickness map generated by the same low-cost system used to generate the en face image in the second row of FIG. 4, and showing the placement of an Early Treatment Diabetic Retinopathy Study (ETDRS) grid over the fovea. [Figure 11] A figure showing an exemplary network of a suitable U-Net architecture for solving the fovea discovery problem. [Figure 12] A figure showing a general-purpose frequency-domain optical coherence tomography system used to collect 3D image data of an eye suitable for use in the present invention. [Figure 13]This diagram shows an exemplary OCT B scan image of a normal human retina, illustrating the identification of various normal retinal layers and boundaries. [Figure 14] This figure shows an exemplary image of the vascular system in the face. [Figure 15] This figure shows an exemplary B-scan vascular image. [Figure 16] This figure shows an example of a multilayer perceptron (MLP) neural network. [Figure 17] This figure shows a simplified neural network consisting of an input layer, a hidden layer, and an output layer. [Figure 18] This diagram shows an exemplary convolutional neural network architecture. [Figure 19] This diagram shows an exemplary U-Net architecture. [Figure 20] This is a diagram illustrating an exemplary computer system (or computing device or computer). [Modes for carrying out the invention]
[0019] When attempting to align OCT volumes, en-face or vascular (e.g., OCTA) maps or other (e.g., frontal view) 2D representations / maps of corresponding slabs (e.g., sub-volumes) from each of the OCT volumes may be used, and a single pair of corresponding en-face or vascular maps may be used for alignment. That is, the 2D representations provided by the corresponding en-face or vascular maps are compared to identify similarities for alignment purposes. This can be done by identifying unique features (e.g., structures, pixels, image properties, etc., that can be defined in a way that makes them distinguishable from other feature points / regions in the 2D representation), aligning the corresponding unique features from the two 2D representations (determined by the similarity of their definitions), and using one or more (known) techniques to transform the 2D representations (e.g., adjusting their position, angle, shape, etc.) in a way that allows at least a portion of the 2D representations to be aligned by defining (alignment) transformation parameters or (alignment) transformation matrices that can be applied to the 2D representations. To properly align the entire 2D representation, a broad distribution of aligned intrinsic features across the entire (or majority) region of the 2D representation is desirable. Once the (alignment) transformation parameters are defined, they can be ported / applied to their corresponding OCT volumes (e.g., per A-scan) for alignment. If necessary, axial corrections can also be applied, such as by axial alignment / alignment of the corresponding A-scan (from the corresponding OCT volume). Optionally, the 2D representation used for alignment may be formed from a combination of multiple en-face (structure) images or vascular maps (e.g., en-face angiography images). That is, multiple en-face images or vascular maps can be combined to generate a single en-face image for alignment. This can provide a 2D image with additional image clarity, thereby allowing for the rendering of additional intrinsic features.
[0020] The present invention provides an alternative method for aligning OCT volumes. Unless otherwise stated or understood from the context, the term "en face" may be used interchangeably with the term "2D representation" and may include, for example, en face images and / or vascular maps and / or thickness maps and / or other types of 2D maps / images that can be derived therefrom or formed from OCT slabs.
[0021] One embodiment of the present invention uses landmark matching of multiple corresponding (e.g., multiple pairs) en faces from multiple corresponding OCT volumes (e.g., two or more OCT volumes of the same eye acquired at different examinations and / or different scan times). For example, multiple en faces may be created from multiple different OCT subvolumes (e.g., defined between multiple pairs of different tissue layers within an OCT volume (e.g., between two retinal layers, or between a target retinal layer and an axial offset from the target retinal layer)). An en face may be an OCT-based en face image or an OCTA-based vascular map, or a combination of an OCT en face image and an OCTA vascular map, and / or other types of 2D representations or maps (e.g., a thickness map and / or curvature map). Thus, multiple corresponding (e.g., multiple pairs) en faces may be formed from multiple subvolumes of different clarity (e.g., corresponding subvolumes of multiple definitions), and each pair (or set of corresponding) en face renders its own corresponding (local) set of unique features. Next, the unique features (landmarks) of a global set consisting of unique features from all different en faces (or 2D representations / maps) can be aligned together as a group. This makes it possible to identify more consistent unique features across multiple en faces and discard unique features that may be due to errors in one or more en faces but are not found in other en faces (e.g., in all or a majority of en faces, or in a predetermined minimum number or percentage of en faces). Furthermore, since unique features are identified in multiple independent en faces, considering all (valid) unique features across all en faces as a whole / group improves the probability of identifying a wider distribution of unique features across the entire overlapping 2D space. In this way, multiple landmarks (unique features) can be detected in a pair of 3D OCT volumes, and the XY positions of the landmark matches can be used for 2D alignment.
[0022] This method can use one or more method steps outlined below (for example, applied to one or more pairs of OCT / OCTA volumes to be aligned) to identify matching unique features and define a transformation model (e.g., consisting of or based on alignment transformation parameters) for aligning a pair of OCT / OCTA volumes. 1) Select two (e.g., corresponding) layers (e.g., the outer boundary of the ILM [internal limiting membrane] and the IPL [internal reticular layer]) and generate 2D images for each OCT / OCTA volume, such as an en-face image and / or a vascular map and / or a thickness map (e.g., a 2D representation of the slab formed by the two selected layers). 2) A step of selecting one of several 2D images as the reference image. 3) The step of identifying / defining landmarks (unique features) in a reference image, for example, by randomly sampling the reference image or by using a feature discovery algorithm to find a sufficient number of landmarks. 4) A step in which a feature discovery / matching algorithm (e.g., template matching) is used to identify / define corresponding reference landmarks in one or more other 2D images. The feature discovery / matching algorithm may be the same as or similar to the landmark definition method used in step 3. 5) Repeat steps 1-4 above to form multiple corresponding 2D images (e.g., en-face images and / or vascular maps) of the OCT / OCTA volume to be aligned. 6) A step of defining a comprehensive set of unique features (landmarks) for landmark-based alignment using all (e.g., valid) landmark matches from all 2D images (en-face images and / or vascular maps). For example, landmark-based alignment is the following: a. Use an appropriate transformation (e.g., rigid, affine, nonlinear). b. Select the best landmark matches of a subset using RANSAC (Random Sample Consensus) or other feature matching algorithms. c. Selecting some of the best landmark matches (by exhaustive search or other appropriate methods, etc.) d. This may include maximizing the distribution of landmarks across multiple 2D images.
[0023] Figures 1, 2, and 3 provide an example of aligning multiple pairs of corresponding OCT datasets (of age-related macular degeneration (AMD) patients) acquired at two different consultations. For illustrative purposes, Figures 1, 2, and 3 each show macular thickness maps from the first consultation (upper right) and the second consultation (lower right). The thickness maps demonstrate their use in disease progression analysis. For example, Figures 1 and 3 show how much the foveal region may become obscured over time, which can indicate disease progression. In Figures 1, 2, and 3, the first three rows (from the top) show the alignment of individual en-face images (defined from different subvolumes) by matching landmarks (unique features) of a subset of localized sets selected by this algorithm ("Landmark Alignment" in the fourth column from the left). For illustrative purposes, in Figures 1, 2, and 3, the first row of OCT en face images is generated using mean projection based on the ILM layer and the 0.5*(ILM+RPE) layer; the second row of OCT en face images is generated using mean projection based on the 0.5*(ILM+RPE) layer and the RPE layer; and the third row of OCT en face images is generated using mean projection based on the RPE layer and the RPE+100 micron layer. These layers are proposed because they are often the easiest / quickest to identify, even with low-quality OCT data. In these examples, the last row (the fourth row from the top) shows the alignment of the en face image in the third row, but uses the intrinsic features of a comprehensive set of intrinsic features (over 100 landmark matches) selected from intrinsic features (landmarks) of multiple local sets that may include landmarks (features) from the first two rows (and other matched 2D representations, though not shown). Thus, the last row can be considered the most accurate alignment. Some of the landmarks selected by the algorithm are shown in the figure (fourth column from the left).
[0024] The example in Figure 1 shows better landmark matching in en-face images generated based on RPE layers and RPE+100 micron layers. Using all landmarks from all en-face images creates a comprehensive set of unique features (landmarks) with the largest number of landmarks having the greatest distribution (across the 2D space of the image), which can result in more accurate alignment.
[0025] The example in Figure 2 demonstrates the performance of this alignment method in the presence of segmentation errors and large saccade motion during the second examination. By using all landmarks from all en-face images, the maximum number of landmarks with the maximum distribution is generated, which again leads to more accurate alignment.
[0026] The example in Figure 3 shows better landmark matching in en-face images generated based on a 0.5*(ILM+RPE) layer and an RPE layer. By using all landmarks from all en-face images, the maximum number of landmarks with the maximum distribution is generated, which leads to more accurate alignment, as shown in the fourth row from the top.
[0027] This method can be extended to 3D OCT / OCTA volumes. One approach involves the following steps: 1) A step of selecting one of several volumes as the reference volume, 2) A step of finding landmarks within a reference volume by finding a sufficient number of 3D landmarks using a 3D feature discovery algorithm. 3) A step of finding a corresponding reference landmark in the other volume using a feature matching algorithm (e.g., 3D template matching). 4) A step to perform 2D alignment using a 3D landmark with matching XY coordinates, 5) This may include a step of 3D alignment using matching 3D landmarks in XYZ coordinates.
[0028] In addition to (or alternatively to) forming multiple 2D representations using multiple sub-volume definitions, multiple additional 2D representations can also be extracted / derived from existing 2D maps or images. For example, multiple image texture maps or color maps can be extracted from an en-face image (e.g., OCT data) or a vascular map (e.g., OCTA data), and each of the extracted pieces of information may form a new 2D representation. However, another example is deriving multiple new 2D representations from an existing thickness map. A thickness map inherently contains depth information due to its thickness (e.g., axial / depth) information, and since this information is typically displayed in a 2D color format (or grayscale if color is not available), thickness maps are included in the description of these 2D representations as described herein.
[0029] It has been found that even with low-quality OCT data (e.g., when en-face images and / or vascular maps do not provide enough information / detail to generate a sufficient number of unique features), it is still possible to generate a good thickness map. Because retinal thickness does not change much, extracting a sufficient number of unique features from a sufficiently clear thickness map can still be challenging. However, while the thickness of a healthy retina hardly changes, the retina has a significant curvature because the eyeball is round. Multiple curvature maps can be derived from the thickness map (preferably of the macula), and the clearer features of the curvature maps can be used to extract a sufficient number of unique features for proper alignment of the OCT dataset, even if they are of low quality.
[0030] As explained above, the 2D representation of 3D OCT volumes is one of the OCT visualization techniques that has benefited significantly from technological advancements in OCT technology. An example of a 2D representation of a 3D OCT volume is a layer thickness map. Multilayer segmentation is used to generate layer thicknesses. The thickness is measured based on the difference between the boundaries of one or more retinal layers.
[0031] [Image alignment using thickness maps and curvature maps] To measure changes in thickness over time, time-series data (obtained from the same subject) must be aligned. Two OCT volumes can be laterally aligned using OCT en face (and vascular maps). Typically, similarity between two images is used in image alignment algorithms. However, detecting areas of similarity in lower-quality OCT data, such as those provided by two low-cost line-field OCT en face images, can be difficult due to the low contrast of the en face images, which can render these images unusable for alignment purposes.
[0032] Figure 4 provides a comparison of B-scan and en-face images acquired using a high-quality OCT system (e.g., the Cirrus® system by Carl Zeiss Meditec, Inc. (CZMI)®, Dublin, California, USA) with those of a low-cost line-field SD-OCT. As can be seen from the figure, the high-quality OCT system provides clear contrast and sharpness of images in its B-scan and in each of its three distinct subvolumes, which are identified herein as superficial, deep, and choroidal. In contrast, the low-cost line-field B-scan loses much of the detail variation provided by the Cirrus system, and the en-face images of the B-scan provide little to no distinction between different en-face images (from the corresponding different subvolumes of the B-scan (superficial, deep, and choroidal)).
[0033] In contrast, similar regions can be detected in macular thickness maps of the same eye (acquired using a low-cost OCT system). Figure 5 shows several pairs of thickness maps that can be used for alignment purposes. However, using only thickness maps for alignment is not advisable for the following reasons: 1) Thickness maps represent smooth surfaces, and finding a sufficient number of point correspondences between two thickness maps can be a challenging task, especially in normal cases where thickness map variability is minimal compared to the thickness map of diseased cases. 2) This may be a challenging task because pathological changes may affect one or more regions of subretinal volume data associated with a particular disease (for example, diabetic retinopathy may affect the superficial and deep retinal layers). Therefore, using only a thickness map may affect the accuracy of the alignment.
[0034] The alignment problem primarily involves aligning corresponding landmarks that are found in a pair of thickness maps. Landmark matching can be a challenging problem due to changes over time in the two maps, as mentioned above. However, landmark-based alignment can work well if sufficiently distributed landmark matches can be found across multiple maps and an appropriate transformation model (e.g., one composed of or based on alignment transformation parameters) is selected.
[0035] Figure 6 shows an example of landmark matching between two thickness maps. In this embodiment, landmark matching is not well distributed across the entire 2D space defined by the thickness map, which can lead to low-quality alignment.
[0036] Therefore, a solution is needed to overcome the above-mentioned problems that may affect alignment quality. This embodiment uses landmark matching of multiple pairs of maps (e.g., 2D maps). By using multiple pairs of maps, it is possible to generate additional point correspondences at different locations and enhance the distribution of point correspondences. Examples of maps that may be used are as follows: 1) Multiple thickness maps of different retinal layers, 2) Thickness map and corresponding multiple curvature maps, 3) This may include a combination of steps (1) and (2) described above.
[0037] Generally, curvature measures how much a surface is curved in different directions and by different amounts at a given surface point. The advantage of using multiple curvature maps (of the retina) is that these maps have higher variability and contrast than the thickness map alone.
[0038] Figure 7 shows a pair of thickness maps and a set of curvature maps corresponding to (derived from) the pair of thickness maps. In this embodiment, three different curvature maps (mean curvature, maximum curvature, and minimum curvature) are derived from each of the set of thickness maps. Other examples of curvature maps that can be derived from thickness maps, such as Gaussian, are known in the literature, and in this embodiment, any one or more such maps may be used.
[0039] [Alignment] A partial overview of this method using multiple maps is provided herein. To align two OCT volumes of the same eye, 1) Using two layers (e.g., ILM and outer RPE), generate a macular thickness map for each OCT volume. 2) Generate / derive one or more (e.g., a series of) curvature maps (such as Gaussian curvature maps, mean curvature maps, maximum curvature maps, minimum curvature maps, etc.) for each macula thickness map. 3) Select one macular thickness map and a corresponding set of multiple curvature maps as reference maps. 4) For example, find landmarks in a reference map by randomly sampling the reference map or by using a feature discovery algorithm to find a sufficient number of landmarks. 5) Using a feature matching / discovery algorithm (e.g., template matching), find the corresponding reference landmark in the second macula thickness map and the corresponding multiple curvature maps. 6) For landmark-based alignment, use matching landmarks from all (or combinations of) maps. Landmark-based alignment is performed by: a. Use an appropriate transformation (e.g., rigid, affine, nonlinear). b. Select some of the best landmark matches using RANSAC (Random Sample Consensus) or other methods. c. Use exhaustive search to select some of the best landmark matches. d. Ensuring the maximum distribution of landmarks across the entire image, including.
[0040] Figures 8A and 8B show two exemplary AMD cases from two consultations using the Cirrus instrument. In each case, the first four rows (from the top) show the alignment of individual maps using unique features (landmarks) from a separate local set and showing matches of some of the landmarks selected by this algorithm. In each case, the bottom (last) row (from the top) shows the alignment of macular thickness maps using (e.g., valid) unique features from a comprehensive set selected from all landmarks (e.g., more than 100 initial landmark matches per map). Some of the (e.g., valid) landmarks selected by the algorithm are shown in the rightmost column of each figure. As illustrated, using all landmarks from all maps creates the maximum number of landmarks with the maximum distribution, which can lead to more accurate alignment, as shown in the bottom (last) row.
[0041] Figures 9A and 9B provide two examples, one normal and one with central serous retinopathy (CSR), from the same examination using a low-cost Rheinfield SD-OCT instrument. In each example, the first four rows (from the top) show the alignment of individual maps using unique features (landmarks) of a separate local set and indicating matches of some of the landmarks selected by the algorithm. The last row (e.g., the bottom) shows the alignment of macular thickness maps using (e.g., valid) unique features (landmarks) of a comprehensive set selected from all landmarks (e.g., more than 100 initial landmark matches per map). In both Figures 9A and 9B, some of the landmarks selected by the algorithm are shown in the far right column.
[0042] The distinction between this method and the method of the above embodiment is that this method relies more heavily on multiple maps (curvature maps) created / derived from a single (base or parent) map (thickness map), whereas the above embodiment relies more heavily on multiple independent OCT / OCTA slabs for alignment. However, it should be understood that both methods can be used jointly / in combination (both contributing to landmarks / unique features) for the purpose of alignment. For example, landmarks can be detected using OCT / OCTA en face images and vascular maps as well as thickness maps and corresponding multiple curvature maps.
[0043] Some of the embodiments described above utilize mathematical tools such as surface curvature derived from differential geometry. One advantage of using thickness maps and corresponding curvature maps is that alignment and fovea detection are independent of OCT data obtained from different OCT techniques. This can be a convenient way to solve the multimodal alignment and fovea detection problem, for example, by using different OCT systems with different OCT qualities and different techniques (SD-OCT, SS-OCT, etc.).
[0044] Typically, en-face images and / or vascular maps are used to identify different physiological features of the eye, such as the fovea. However, as described above, low-cost OCT systems may not be able to provide en-face images or vascular maps with sufficient contrast / clarity for this purpose. It is proposed herein that thickness maps (and / or multiple curvature maps derived therefrom) may be used to identify distinct physiological features of the eye, such as the fovea.
[0045] [Discovery of the fovea using thickness maps and curvature maps] The location of the fovea can be found using thickness maps and corresponding curvature maps. The location of the fovea is clinically important because it is the location of highest visual acuity. In automated analysis of retinal diseases, the location of the fovea is used as a reference point. The fovea has several prominent anatomical features (e.g., vascular pattern and FAZ) that can be used to identify it in OCTA images. The presence of pathologies such as edema, posterior vitreous (posterior vitreous) membrane traction, CNV, and other disease types can disrupt the normal foveal structure. Furthermore, this anatomical structure can be disrupted by various lesions of the outer layers of the retina.
[0046] One use case for foveal location is the placement of the Early Treatment Diabetic Retinopathy Study (ETDRS) grid at foveal locations (see Figure 10). As shown above with reference to Figure 4, the low contrast of OCT en face images acquired using low-cost line-field SD-OCT systems can make them unusable for finding foveal location. In contrast, Figure 10 shows that macular thickness maps (e.g., between the ILM and RPE layers, left image) and corresponding curvature maps (not shown) generated by the same low-cost system contain information in the foveal region that can be used to find foveal location. Thickness maps and corresponding curvature maps can be used in combination with machine learning techniques to develop machine models for identifying foves when given thickness maps or OCT datasets as input. For example, deep learning methods can be used for automated detection of foveal centers using macular thickness maps and corresponding curvature maps. The same idea can be applied to other retinal landmarks such as the optic nerve head (ONH).
[0047] In an exemplary embodiment of this design, a convolutional neural network was used to segment the foveal region within the macular thickness map and corresponding curvature map. A description of the machine learning and neural network is provided below, in particular. Target training images were generated by creating a 3mm binary disk mask around the foveal center identified by a human expert grader using OCT data or a macular thickness map or a combination thereof. An example of a suitable deep learning algorithm / machine / model / system is the U-net architecture (e.g., n-channel), which employs five convergent convolutional layers and five augmented convolutional layers, ReLU activation, maximum pooling, binary cross-entropy loss, and sigmoid activation in the final layer. The input channels can be one or more macular thickness maps and their curvature maps, such as a Gaussian curvature map, mean curvature map, maximum curvature map, and minimum curvature map. In this design, data augmentation (rotation in 3° steps around the center between -9° and 9°) was performed to increase the amount of training data. U-Net predicted the ONH region and then performed template matching using a 3 mm diameter disk to locate the foveal center. Figure 11 shows an exemplary network of the U-Net architecture that can solve the foveal discovery problem. Other examples of neural networks, including another exemplary U-Net architecture suitable for use in conjunction with the present invention to locate the fovea (or other physiological structures of the retina), are provided below.
[0048] The aforementioned en-face image and thickness-mapping techniques can be used to train deep learning machine models (e.g., convolutional neural networks (CNNs)), which can be used to solve / provide motion estimation using optical flow, or to directly estimate transformation parameters (e.g., to provide an alignment model that can be used to transform a video to a reference image). Optical flow is a measure of sub-pixel translation between two images. For example, optical flow can be a description of how fast and in which direction a sub-image region (which may represent a feature or object in the image) moves (e.g., with respect to the pixels of the image). Optical flow can be computed using a network such as FlowNet (e.g., a convolutional (neural) network used to learn optical flow). A CNN can be trained to output transformation parameters used to align a video to match a reference image.
[0049] Various hardware and architectures suitable for the present invention are described below. Optical coherence tomography system Generally, optical coherence tomography (OCT) uses low-coherence light to generate two-dimensional (2D) and three-dimensional (3D) internal views of biological tissue. OCT enables in vivo imaging of retinal structures. OCT angiography (OCTA) generates flow information, such as blood flow from within the retina. Examples of OCT systems are provided in U.S. Patent No. 6,741,359 and No. 9,706,915, and examples of OCTA systems are provided in U.S. Patent No. 9,700,206 and No. 9,759,544, all of which are incorporated herein by reference in their entirety. Exemplary OCT / OCTA systems are provided herein.
[0050] Figure 12 illustrates a general frequency-domain optical coherence tomography (FD-OCT) system for acquiring 3D image data of the eye suitable for use with the present invention. The FD-OCT system OCT_1 includes a light source LtSrc1. Typical light sources include, but are not limited to, broadband light sources with short time coherence length or swept laser sources. The beam of light from the light source LtSrc1 is typically guided by an optical fiber Fbr1 to illuminate a sample, e.g., an eye E, a typical sample being human intraocular tissue. The light source LrSrc1 may be, for example, a broadband light source with short time coherence length in the case of spectral-domain OCT (SD-OCT) or a tunable laser source in the case of swept light source OCT (SS-OCT). The light may typically be scanned using a scanner Scrnr1 between the output of the optical fiber Fbr1 and the sample E, so that the beam of light (dashed line Bm) is scanned laterally across the region of the sample to be imaged. The light beam from scanner Sconr1 passes through scanning lens SL and ophthalmic lens OL and can be focused onto the sample E to be imaged. Scanning lens SL can receive the light beam from scanner Sconr1 at multiple incidence angles and produce substantially collimated light, which can then be focused onto the sample by ophthalmic lens OL. This example shows a scanning beam that needs to be scanned in two lateral directions (e.g., x and y directions on the Cartesian plane) to scan a desired field of view (FOV). This example is a point-field OCT that uses a point-field beam to scan across the sample. Thus, scanner Sconr1 is shown exemplarily to include two subscanners, namely a first subscanner Xscn for scanning the point-field beam across the sample in a first direction (e.g., horizontal x direction) and a second subscanner Yscn for scanning the point-field beam on the sample in a second intersecting direction (e.g., vertical y direction). If the scanning beam is a line-field beam (e.g., line-field OCT) and can sample the entire line portion of the sample at once, then only one scanner may be required to scan the line-field beam across the sample to cover the desired field of view (FOV).If the scanning beam is a full-field beam (e.g., full-field OCT), a scanner may not be required, and the full-field light beam may illuminate the entire desired FOV at once.
[0051] Regardless of the type of beam used, scattered light from the sample (e.g., sample light) is collected. In this embodiment, scattered light returning from the sample is collected in the same optical fiber Fbr1 used to route the light for illumination. Reference light derived from the same light source LtSrc1 travels along a different path, in this case including an optical fiber Fbr2 and a back reflector RR1 with an adjustable optical delay. As those skilled in the art will know, a transmissive reference path can also be used, and the adjustable delay can be placed in the sample or the reference arm of the interferometer. The focused sample light is coupled with the reference light in, for example, a fiber coupler Cplr1 to form an optical interference in an OCT photodetector Dtctr1 (e.g., a photodetector array, digital camera, etc.). Although it is shown that one fiber port reaches the detector Dtctr1, as those skilled in the art will know, various designs of interferometers can be used for balancing or unbalancing the interference signal. The output from detector Dtctr1 is fed to a processor (e.g., an internal or external computing device) Cmp1, which converts the observed interference into sample depth information. Depth information may be stored in memory associated with processor Cmp1 and / or displayed on a display (e.g., computer / electronic display / screen) Scin1. Processing and storage functions may be localized within the OCT device, or the functions may be offloaded to an external processor (e.g., an external computer system) to which the collected data is transferred (e.g., executed on the external processor). Figure 15 shows an example of a computing device (or computer system). This unit may be dedicated solely to data processing, or it may perform other common tasks not specific to the OCT device.The processor (computing device) Cmp1 may include, for example, a field-programmable gate array (FPGA), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a graphics processing unit (GPU), a system-on-a-chip (SoC), a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), or a combination thereof, which can perform some or all of the processing steps in a serial and / or parallel manner with one or more host processors and / or one or more external computing devices.
[0052] The sample arm and reference arm within the interferometer can consist of a bulk optical system, a fiber optical system, or a hybrid bulk optical system, and can have different architectures, such as Michelson, Mach-Zehnder, or common path system designs, as is known to those skilled in the art. The term "light beam," as used herein, should be interpreted as any carefully directed optical path. Instead of mechanically scanning a beam, the light field can illuminate a one- or two-dimensional area of the retina to generate OCT data (e.g., U.S. Patent No. 9332902, D. Hillmann et al., "Holoscopy-holographic optical coherence tomography," Optics Letters, Vol. 36(13), p. 2290, 2011; Y. Nakamura et al., "High-Speed three-dimensional human retinal imaging by line field spectral domain optical coherence tomography," Optics Express). (See Applied Optics, Vol. 44(36), p. 7722 (2005)), Blazkiewicz et al., "Signal-to-noise ratio study of full-field Fourier-domain optical coherence tomography," Express, Vol. 15(12), p. 7103, 2007. In time-domain systems, the reference arm must have an adjustable optical delay to produce interference. Balance detection systems are typically used in TD-OCT and SS-OCT systems, and spectrometers are used in detection ports for SD-OCT systems. The inventions described herein can be applied to any type of OCT system.Various aspects of the present invention can be applied to any type of OCT system, or to multiple types of ophthalmic diagnostic systems and / or ophthalmic diagnostic systems including, but not limited to, fundus imaging systems, visual field testing devices, and scanning laser polarimeters.
[0053] In Fourier-domain optical coherence tomography (FD-OCT), each measurement is a real-value spectrally controlled interference figure (Sj(k)). The real-value spectral data typically undergoes several post-processing steps, including background removal and dispersion correction. The Fourier transform of the processed interference figure yields the complex OCT signal output Aj(z) = |Aj|eiφ. From the absolute value of this complex OCT signal, |Aj|, the scattering intensity at different path lengths, and therefore the scattering profile with respect to depth (z-direction) within the sample, is revealed. Similarly, the phase φj can also be extracted from the complex OCT signal. The scattering profile with respect to depth is called the axial scan (A-scan). A collection of A-scans measured at adjacent locations within the sample generates a cross-sectional image (tomographic image or B-scan) of the sample. A collection of B-scans collected at different lateral locations on the sample constitutes a data volume or cube. For a given data volume, the fast axis refers to the scanning direction along a single B-scan, and the slow axis refers to the axis along which multiple B-scans are collected. The term "cluster scan" may refer to a unit or block of data generated by repeated acquisitions at the same (or substantially the same) location (or region) for the purpose of analyzing motion contrast that may be used to identify blood flow. A cluster scan can consist of multiple A-scans or B-scans collected at relatively short time intervals at approximately the same location on the sample. Because the scans in a cluster scan are from the same region, static structures remain relatively unchanged between scans in the cluster scan, while motion contrast between scans that meet certain criteria may be identified as blood flow.
[0054] Various methods for generating B-scans are known in the industry, including, but are not limited to, those along the horizontal or x-direction, along the vertical or y-direction, along the x and y diagonals, or in circular or spiral patterns. B-scans may be in the xz dimension, but may also be any cross-sectional image including the z dimension. An exemplary OCT B-scan image of a normal human retina is shown in Figure 13. An OCT B-scan of the retina provides a view of the structure of the retinal tissue. For illustrative purposes, Figure 13 identifies various normal retinal layers and layer boundaries. The identified retinal boundary layers include (from top to bottom) the inner limiting membrane (ILM) layer 1, the retinal nerve fiber layer (RNFL) layer 2, the ganglion cell layer (GCL) layer 3, the inner plexiform layer (IPL) layer 4, the inner nuclear layer (INL) layer 5, the outer plexiform layer (OPL) layer 6, the outer nuclear layer (ONL) layer 7, the junction between the outer segments (OS) and inner segments (IS) of the photoreceptor cells (indicated by reference code layer 8), the outer limiting membrane or outer limiting membrane (ELM) layer 9, the retinal pigment epithelium (RPE) layer 10, and the Bruch's membrane (BM) layer 11.
[0055] In OCT angiography or functional OCT, the analysis algorithm may be applied to OCT data collected at different times (e.g., cluster scans) at the same or nearly the same sample location on the sample to analyze motion or flow (see, for example, U.S. Patent Publication Nos. 2005 / 0171438, 2012 / 0307014, 2010 / 0027857, 2012 / 0277579, and U.S. Patent No. 6549801, all of which are incorporated herein by reference). The OCT system may use any one of many OCT angiography processing algorithms (e.g., motion contrast algorithms) to identify blood flow. For example, a motion contrast algorithm can be applied to intensity information derived from image data (intensity-based algorithm), phase information from image data (phase-based algorithm), or complex image data (complex-based algorithm). An en-face image is a 2D projection of 3D OCT data (e.g., by averaging the intensity of each individual A-scan, thereby defining each A-scan as a pixel in the 2D projection). Similarly, an en-face vascular image is an image that displays motion contrast signals, in which the data dimension corresponding to depth (e.g., the z-direction along the A-scan) is typically represented as a single representative value (e.g., a pixel in the 2D projection image) by adding or aggregating all or isolated portions of the data (see, for example, U.S. Patent No. 7301644, which is incorporated herein by reference in its entirety). An OCT system that provides angiography capabilities may be called an OCT angiography (OCTA) system.
[0056] Figure 14 shows an example of an en-face vascular image. After processing the data and highlighting the motion contrast using one of the motion contrast methods known in the art, a pixel range corresponding to a certain tissue depth from the surface of the internal limiting membrane (ILM) of the retina may be added to generate an en-face (e.g., frontal view) image of the vascular system. Figure 15 shows an exemplary B-scan of an vascular (OCTA) image. As illustrated, structural information may not be clear because blood flow traversing multiple retinal layers can obscure multiple retinal layers more than in a structural OCT B-scan as shown in Figure 13. Nevertheless, OCTA provides a non-invasive technique for imaging the microvascular system of the retina and choroid, which can be important for diagnosing and / or monitoring a variety of lesions. For example, OCTA can be used to identify diabetic retinopathy by identifying microaneurysms, neovascular complexes, and quantifying foveal avascular zones and non-perfusion areas. Furthermore, OCTA has been shown to have good agreement with fluorescein angiography (FA), a more traditional but less invasive technique that requires dye injection to observe vascular flow in the retina. Additionally, in atrophic age-related macular degeneration, OCTA is used to monitor the overall decrease in choroidal capillary lamina flow. Similarly, in exudative age-related macular degeneration, OCTA can provide qualitative and quantitative analysis of choroidal neovascularization. OCTA is also used to study vascular occlusion, for example, to evaluate non-perfusion areas and assess the integrity of the superficial and deep plexuses.
[0057] Neural Network As mentioned above, the present invention may use neural network (NN) machine learning (ML) models. For the sake of clarity, neural networks will be outlined herein. The invention may use any of the following neural network architectures, either alone or in combination. A neural network, or neural net, is a network of interconnected neurons (through nodes), where each neuron represents a node in the network. The set of neurons may be arranged in layers, and the output of one layer is forward-facing to the next layer in a multilayer perceptron (MLP) configuration. An MLP may be understood as a feedforward neural network that maps a set of input data to a set of output data.
[0058] Figure 16 illustrates an example of a multilayer perceptron (MLP) neural network. Its structure may include multiple hidden (e.g., inner) layers HL1 to HLn, which map the input layer InL (which receives a set of inputs (or vector inputs) in_1 to in_3) to the output layer OutL, which generates a set of outputs (or vector outputs), e.g., out_1 and out_2. Each layer may have any number of nodes, which are shown here as circles within each layer for illustrative purposes. In this example, the first hidden layer HL1 has two nodes, and the hidden layers HL2, HL3, and HLn each have three nodes. Generally, the deeper the MLP (e.g., the more hidden layers in the MLP), the greater its learning capacity. The input layer InL receives a vector input (shown here as a 3D vector consisting of in_1, in_2, and in_3 for illustrative purposes) and may feed the received vector input to the first hidden layer HL1 in the sequence of hidden layers. The output layer OutL receives the output from the last hidden layer in the multilayer model, for example HLn, and generates a vector output result (shown for illustrative purposes as a two-dimensional vector consisting of out_1 and out_2).
[0059] Typically, each neuron (i.e., node) produces one output, which is then fed forward to the neurons in the immediately following layer. However, each neuron in a hidden layer may receive multiple inputs from the input layer or from the output of a neuron in the immediately preceding hidden layer. In general, each node may apply a function to its input to produce an output for that node. Nodes in a hidden layer (e.g., a learning layer) may apply the same function to each input to produce their respective outputs. However, some nodes, for example, nodes in the input layer InL, may receive only one input and be passive, meaning they simply relay the value of that single input to their output, for example, providing a copy of that input to their output, which is illustrated by the dashed arrows within the nodes in the input layer InL for illustrative purposes.
[0060] For illustrative purposes, Figure 17 shows a simplified neural network consisting of an input layer InL', a hidden layer HL1', and an output layer OutL'. The input layer InL' is shown to have two input nodes i1 and i2, which receive input Input_1 and Input_2 respectively (for example, the input nodes of layer InL' receive a 2D input vector). The input layer InL' feeds forward to one hidden layer HL1' which has two nodes h1 and h2, which in turn feeds forward to the output layer OutL' with two nodes o1 and o2. The interconnections, or links, between neurons (shown by solid arrows for illustrative purposes) have weights w1 to w8. Typically, except for the input layer, a node (neuron) may receive the output of the node in the immediately preceding layer as input. Each node may calculate its output by multiplying each of its inputs by the corresponding interconnection weights of each input, adding the products of those inputs, adding (or multiplying) a constant defined by other weights or biases that may be associated with that particular node (e.g., node weights w9, w10, w11, w12 corresponding to nodes h1, h2, o1, and o2, respectively), and then applying a nonlinear or logarithmic function to the result. The nonlinear function may be called an activation function or transfer function. Several activation functions are known in the art, and the choice of a particular activation function is not important to this explanation. However, it should be noted that in ML models, the behavior of the neural network depends on the values of the weights, which the neural network may be learned to provide a desired output for a given input.
[0061] During the training or learning phase, a neural network learns appropriate weight values to achieve a desired output for a given input (e.g., it is trained to identify it). Before the neural network is trained, each weight may be individually assigned an initial (e.g., random, arbitrarily non-zero) value, such as a random seed. Various methods for assigning initial weights are known in the art. The weights are then trained (optimized) so that for a given training vector input, the neural network produces an output close to the desired (given) training vector output. For example, the weights may be gradually adjusted over thousands of iterations by a method called backpropagation. In each cycle of backpropagation, the training input (e.g., a vector input or training input image / sample) is forward-passed through the neural network to provide its actual output (e.g., a vector output). Subsequently, the error of each output neuron, or output node, is calculated based on the actual output of the neuron and the teacher-defined training output for that neuron (e.g., the training output image / sample corresponding to the current training input image / sample). This then propagates backward through the neural network (from the output layer to the input layer), updating the weights based on how much each weight contributes to the overall error, thereby bringing the neural network's output closer to the desired training output. This cycle is then repeated until the neural network's actual output falls within an acceptable error range for the desired training output for its training input. As you may understand, each training input may require many backpropagation iterations to achieve the desired error range. Typically, an epoch refers to one backpropagation iteration for all training samples (e.g., one forward pass and one backward pass), and training a neural network may require many epochs. Generally, the larger the training set, the better the performance of the ML model being trained, so various data augmentation methods may be used to increase the size of the training set.For example, if the training set contains corresponding pairs of training input and training output images, the training images may be divided into multiple corresponding image segments (or patches). The corresponding patches from the training input and training output images may be paired, and multiple training patch pairs may be defined from one input / output image pair, thereby expanding the training set. However, training a large training set increases the demands on computing resources, such as memory and data processing resources. This computational demand may be mitigated by dividing the large training set into multiple minibatches, the size of which determines the number of training samples in a single forward / backward pass. In this case, and a single epoch may contain multiple minibatches. Another problem is the possibility that the neural network may overfit the training set, reducing its ability to generalize from one input to a different input. The problem of overfitting may be mitigated by creating an ensemble of neural networks or by randomly dropping out nodes in the neural network during training, which effectively removes dropped reads from the neural network. Various dropout adjustment methods, such as inverse dropout, are known in the industry.
[0062] It should be noted that the computation of a trained NN machine model is not a simple algorithm of computation / analysis steps. In fact, when a trained NN machine model receives an input, that input is not analyzed in the conventional sense. Rather, regardless of the purpose or nature of the input (e.g., a vector defining a live image / scan, or a vector defining any other entity such as a population structure description or activity record), the input is subject to the same architecture construction of the trained neural network (e.g., the same node / layer arrangement, trained weights and bias values, given convolution / deconvolution operations, activation functions, pooling operations, etc.), and how the architecture construction of the trained network generates its output may not be clear. Furthermore, the values of trained weights and biases are not deterministic and depend on many factors, including the amount of time allocated to training the neural network (e.g., the number of epochs in training), the random starting values of the weights before training began, the computer architecture of the machine on which the NN is trained, the selection of training samples, the distribution of training samples across multiple minibatches, the selection of the activation function, the selection of the error function that modifies the weights, and even whether training was interrupted on one machine (e.g., with the first computer architecture) and completed on another machine (e.g., with a different computer architecture). The point is that the reasons why a trained ML model arrives at a particular output are not obvious, and much research is currently being done to identify the elements on which the ML model's output is based. Therefore, processing of live data with a neural network cannot be reduced to a simple step-by-step algorithm. Rather, the computation depends on the training architecture, the training sample set, the training sequence, and the various circumstances in which the ML model is trained.
[0063] In summary, the configuration of a neural network (NN) machine learning model may include a learning (or training) stage and a classification (or computation) stage. In the learning stage, the neural network may be trained for a specific purpose, and a set of training examples may be provided, which include training (sample) inputs and training (sample) outputs, and optionally a set of validation examples to test the progress of training. During this learning process, various weights related to the nodes and node interconnections in the neural network are gradually adjusted to reduce the error between the actual output of the neural network and the desired training output. In this way, a multilayer feedforward neural network (e.g., the one described above) may be able to approximate any measurable function to any desired accuracy. The result of the learning stage is a learned (e.g., trained) (neural network) machine learning (ML). In the computation stage, a set of test inputs (or live inputs) may be provided to the learned (trained) ML model, which may apply what it has learned to generate output predictions based on the test inputs.
[0064] Similar to the conventional neural networks in Figures 16 and 17, a convolutional neural network (CNN) also consists of neurons with learnable weights and biases. Each neuron receives an input, performs an operation (e.g., a dot product), and is then, by choice, subjected to a nonlinear transformation. However, a CNN takes raw image pixels at one end (e.g., the input end) and provides a classification (or class) score at the other end (e.g., the output end). Since a CNN anticipates an image as input, these are optimized to handle volume (e.g., the pixel height and width of the image, and the depth of the image, such as color depth, e.g., RGB depth defined by three colors: red, green, and blue). For example, the layers of a CNN may be optimized for neurons arranged in three dimensions. Neurons within a CNN layer may be connected to a small region of the previous layer rather than all of the fully connected neurons of the NN. The final output layer of a CNN may reduce the full image to a single vector (classification) arranged along the depth dimension.
[0065] Figure 18 provides an exemplary convolutional neural network architecture. A convolutional neural network may be defined as a sequence of two or more layers (e.g., layers 1 to N), where each layer may include an (image) convolution step, a (resulting) weighted sum step, and a nonlinear function step. Convolution may be performed on its input data by, for example, applying a filter (or kernel) on a moving window across the input data to generate a feature map. Each layer and its components may have different predetermined filters (from a filter bank), weights (or weighting parameters), and / or function parameters. In this example, the input data is an image with a certain pixel height and width, and may be the raw pixel values of this image. In this example, the input image is drawn to have a depth of three color channels RGB (red, green, blue). Optionally, the input image may be preprocessed in various ways, and the results of the preprocessing may be input instead of, or in addition to, the raw image data. Some examples of image processing may include retinal vascular map segmentation, color space conversion, adaptive histogram equalization, and connection component generation. Within a layer, the dot product may be calculated between a weight and a small region in the input volume to which they are connected. Many methods for constructing CNNs are known in the industry, but as an example, layers may be constructed to apply element-wise activation functions, such as a max(0,x) threshold at zero. Pooling functions may be applied to downsample the volume (e.g., along the x and y directions). Fully connected layers may be used to identify classification outputs and generate one-dimensional output vectors that have proven useful for image recognition and classification. However, for image segmentation, the CNN needs to classify each pixel. Since each CNN layer tends to degrade the resolution of the input image, another stage is needed to upsample the image to its original resolution. This may be achieved by applying a transposed convolution (or deconvolution) stage TC, which typically does not use any given interpolation method and instead has learnable parameters.
[0066] Convolutional neural networks (CNNs) are well-suited to many computer vision problems. As mentioned earlier, training CNNs generally requires large training datasets. The U-Net architecture is based on CNNs and can generally be trained with smaller training datasets than traditional CNNs.
[0067] Figure 19 illustrates an exemplary U-Net architecture. This exemplary U-Net includes an input module (or input layer or stage) which accepts an input U-in of any size (e.g., an input image or image patch). For convenience, the image size in any stage or layer is shown in a box representing the image; for example, in the input module, the numbers "128×128" are enclosed, indicating that the input image U-in consists of 128×128 pixels. The input image may be a fundus image, OCT / OCTA en face, B-scan image, etc. However, it should be understood that the input can be of any size or dimension. For example, the input image may be an RGB color image, a monochrome image, a volume image, etc. The input image passes through a series of processing layers, each illustrated with exemplary sizes, but these sizes are for illustrative purposes only and will depend, for example, on the image size, convolutional filters, and / or pooling stage. This architecture consists of a convergence path (in this specification, including four coding modules as an example) followed by an extension path (in this specification, including four decoding modules as an example), and copy-and-crop links (e.g., CC1-CC4) between the corresponding modules / stages that copy the output of one coding module in the convergence path and combine it with (e.g., append to) the upconverted input of the corresponding decoding module in the extension path. The result is a characteristic U-shape, from which the architecture is named. Optionally, a "bottleneck" module / stage (BN) may be placed between the convergence path and the extension path for computational considerations or other reasons. The bottleneck BN may consist of two convolutional layers (with batch normalization and optional dropout).
[0068] The convergence path is similar to that of an encoder, and typically uses feature maps to capture contextual (or feature) information. In this example, each coding module in the convergence path includes two or more convolutional layers, indicated by the asterisk symbol "*", followed by a maximum pooling layer (e.g., a downsampling layer). For example, the input image U-in is shown passing through two convolutional layers, each with 32 feature maps. It can be understood that each convolutional kernel generates a feature map (e.g., the output from a convolution operation using a given kernel is an image generally called a "feature map"). For example, the input U-in, after the first convolution applying 32 convolutional kernels (not shown), produces an output consisting of 32 individual feature maps. However, as is known in the art, the number of feature maps generated by a convolution operation can be adjusted (up or down). For example, the number of feature maps can be reduced by averaging groups of feature maps, removing some feature maps, or by other known methods of reducing feature maps. In this example, this initial convolution is followed by a second convolution in which the output is limited to 32 feature maps. Another way of thinking about feature maps is to consider the output of the convolutional layer as a 3D image given by XY plane pixel dimensions (e.g., 128 × 128 pixels) with 2D dimensions, and depth given by the number of feature maps (e.g., the depth of 32 plane images). Following this example, the output of the second convolution (e.g., the output of the first coding module in the convergence path) can be described as a 128 × 128 × 32 image. The output from the second convolution is then subjected to a pooling operation, which reduces the 2D dimension of each feature map (e.g., the X and Y dimensions may each be reduced by half). The pooling operation can be embodied within a downsampling process, as indicated by the downward arrow. Several pooling methods, such as max pooling, are known in the art, and no particular pooling method is relevant to this invention.The number of feature maps doubles with each pooling, for example, 32 feature maps in the first coding module (or block), 64 feature maps in the second coding module, and so on. Thus, the convergence path forms a convolutional network consisting of multiple coding modules (or stages or blocks). As is typical of convolutional networks, each coding module may provide at least one convolution stage, followed by an activation function (e.g., a rectified linear unit (ReLU) or sigmoid layer) (not shown), and a maximum pooling operation. Generally, the activation function introduces nonlinearity into the layer (e.g., to avoid the problem of overfitting), receives the results of the layer, and decides whether to "activate" the output (e.g., whether a value in a particular node meets a given criterion for forwarding the output to the next layer / node). In summary, the convergence path generally reduces spatial information and increases feature information.
[0069] The extension path is similar to the decoder, and in particular, provides spatial information for the localization and convergence path results, despite the downsampling and any maximum pooling performed in the condensation stage. The extension path includes multiple decoding modules, each decoding module combining its current upconverted input with the output of the corresponding encoding module. Thus, feature and spatial information are combined in the extension path through a series of upconvolutions (e.g., upsampling or transposed convolution, i.e., deconvolution) and combinations with high-resolution features from the convergence path (e.g., via CC1-CC4). Therefore, the output of the deconvolution layer is combined with the corresponding (optionally cropped) feature map from the convergence path, followed by two convolution layers and an activation function (optionally batch normalization).
[0070] The output from the last extension module in the extension path may be fed to other processing / training blocks or layers, such as a classifier block, which may be trained together with the U-Net architecture. Alternatively, or additionally, the output of the last upsampling block (at the end of the extension path) may be subjected to another convolution operation (e.g., an output convolution), as indicated by the dotted arrow, before generating its output U-out. The kernel size of the output convolution may be chosen to reduce the dimensions of the last upsampling block to a desired size. For example, a neural network may have multiple features on a pixel-by-pixel basis just before reaching the output convolution, which may provide a 1x1 convolution operation to combine these multiple features into a single pixel-by-pixel output value at the pixel level.
[0071] Computing devices / systems Figure 20 illustrates an exemplary computer system (or computing device or computer device). In some embodiments, one or more computer systems may provide the functions described or illustrated herein and / or perform one or more steps of one or more methods described or illustrated herein. The computer system may take any suitable physical form. For example, the computer system may be an embedded computer system, a system-on-a-chip (SOC), or a single-board computer system (SBC) (e.g., a computer-on-a-module (COM) or system-on-a-module (SOM)), a desktop computer system, a laptop or notebook computer system, a computer system mesh, a mobile phone, a portable information terminal (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the computer system may reside in a cloud, which may include one or more cloud components in one or more networks.
[0072] In some embodiments, the computer system may include a processor Cpnt1, memory Cpnt2, storage Cpnt3, an input / output (I / O) interface Cpnt4, a communication interface Cpnt5, and a bus Cpnt6. The computer system may optionally also include a display Cpnt7, such as a computer monitor or screen.
[0073] Processor Cpnt1 includes hardware for executing instructions, such as components of a computer program. For example, processor Cpnt1 may be a central processing unit (CPU) or a general-purpose computing-on-graphics processing unit (GPGPU). Processor Cpnt1 may read (or fetch) instructions from internal registers, internal caches, memory Cpnt2, or storage Cpnt3, decode and execute these instructions, and write one or more results to internal registers, internal caches, memory Cpnt2, or storage Cpnt3. In certain embodiments, processor Cpnt1 may include one or more internal caches for data, instructions, or addresses. Processor Cpnt1 may include one or more instruction caches and one or more data caches, for example, to hold data tables. Instructions in the instruction cache may be copies of instructions in memory Cpnt2 or storage Cpnt3, and the instruction cache may speed up the reading of these instructions by processor Cpnt1. Processor Cpnt1 may include any appropriate number of internal registers and may include one or more arithmetic logic units (ALUs). Processor Cpnt1 may be a multicore processor or may include one or more processors Cpnt1. While this disclosure describes and illustrates a specific processor, this disclosure assumes any appropriate processor.
[0074] Memory Cpnt2 may include main memory that stores instructions for processor Cpnt1 to execute processes or hold intermediate data during processing. For example, a computer system may load instructions or data (e.g., a data table) into memory Cpnt2 from storage Cpnt3 or from another source (e.g., another computer system). Processor Cpnt1 may load instructions and data from memory Cpnt2 into one or more internal registers or internal caches. To execute an instruction, processor Cpnt1 may read and decode the instruction from an internal register or internal cache. During or after the execution of an instruction, processor Cpnt1 may write one or more results (which may be intermediate or final results) to an internal register, internal cache, memory Cpnt2, or storage Cpnt3. Bus Cpnt6 may include one or more memory buses (each of which may include an idle bus and a data bus) that connect processor Cpnt1 to memory Cpnt2 and / or storage Cpnt3. Optionally, one or more memory management units (MMUs) facilitate data transmission between the processor Cpnt1 and memory Cpnt2. Memory Cpnt2 (which may be high-speed volatile memory) may include random-access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM). Storage Cpnt3 may include long-term or high-capacity memory storage for data or instructions. Storage Cpnt3 may be built into or external to the computer system and may include one or more of the following: disk drives (e.g., hard disk drives (HDD) or solid-state drives (SSD)), flash memory, ROM, EPROM, optical disks, magneto-optical disks, magnetic tapes, Universal Serial Bus (USB)-accessible drives, or other types of non-volatile memory.
[0075] The I / O interface Cpnt4 may be software, hardware, or a combination of both, and may include one or more interfaces (e.g., serial or parallel communication ports) for communicating with I / O devices, which may enable communication with a person (e.g., a user). For example, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, table, touchscreen, trackball, video camera, other suitable I / O devices, or two or more combinations thereof.
[0076] The communication interface Cpnt5 may provide a network interface for communicating with other systems or networks. The communication interface Cpnt5 may include a Bluetooth® interface or other types of packet-based communication. For example, the communication interface Cpnt5 may include a network interface controller (NIC) and / or a wireless NIC or wireless adapter for communication with a wireless network. The communication interface Cpnt5 may provide communication with Wi-Fi networks, ad-hoc networks, personal area networks (PANs), wireless PANs (e.g., Bluetooth WPANs), local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), mobile phone networks (e.g., Global System for Mobile Communications (GSM®) networks), the Internet, or a combination of two or more of these.
[0077] Bus Cpnt6 may provide communication links between the aforementioned components of the computing system. For example, bus Cpnt6 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand bus, a Low-Pin-Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or any other suitable bus, or a combination of two or more of these.
[0078] While this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure also assumes any particular computer system having any particular number of particular components in any particular arrangement.
[0079] In this specification, computer-readable non-temporary storage media may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards or drives, or any other suitable computer-readable non-temporary storage media, or any two or more of these in any suitable combination. Computer-readable non-temporary storage media may be volatile, non-volatile, or a combination of volatile and non-volatile, as appropriate.
[0080] Although the present invention has been described in conjunction with several specific embodiments, many other alternatives, improvements, and variations will be apparent to those skilled in the art by referring to the above description. Therefore, the invention described herein is intended to encompass all such alternatives, improvements, applications, and variations that may fall within the spirit and scope of the accompanying claims.
Claims
1. A method for aligning first optical coherence tomography (OCT) volume data with second OCT volume data, A step of generating a plurality of image pairs, wherein each image pair includes a two-dimensional (hereinafter referred to as 2D) representation of a subvolume in the first OCT volume data and a corresponding 2D representation of the corresponding subvolume in the second OCT volume data, For each image pair, the steps include identifying matching unique features of a local set within the corresponding 2D representation of each image pair, The steps include defining a set alignment transformation parameter based on the intrinsic features of a comprehensive set, which are based on the matching intrinsic features of the local sets extracted from all of the aforementioned image pairs, The step of electronically processing, storing, or displaying the alignment of the first and second OCT volume data based on the aforementioned set of alignment transformation parameters, At least a portion of the plurality of image pairs includes a first image pair and one or more derived image pairs based on the first image pair. A method wherein the first pair of images is a corresponding thickness map, and the one or more derived pairs of images are corresponding curvature maps based on the corresponding thickness maps.
2. A method for aligning first optical coherence tomography (OCT) volume data with second OCT volume data, A step of generating a plurality of image pairs, wherein each image pair includes a two-dimensional (hereinafter referred to as 2D) representation of a subvolume in the first OCT volume data and a corresponding 2D representation of the corresponding subvolume in the second OCT volume data, For each image pair, the steps include identifying matching unique features of a local set within the corresponding 2D representation of each image pair, The steps include defining a set alignment transformation parameter based on the intrinsic features of a comprehensive set, which are based on the matching intrinsic features of the local sets extracted from all of the aforementioned image pairs, The step of electronically processing, storing, or displaying the alignment of the first and second OCT volume data based on the aforementioned set of alignment transformation parameters, At least a portion of the plurality of image pairs includes a first image pair and one or more derived image pairs based on the first image pair. The method wherein the first pair of images is a corresponding en-face image, and the one or more derived pair of images is based on one or more of the image texture, color, intensity, contrast, and negative image of the en-face image.
3. A method for aligning first optical coherence tomography (OCT) volume data with second OCT volume data, A step of generating a plurality of image pairs, wherein each image pair includes a two-dimensional (hereinafter referred to as 2D) representation of a subvolume in the first OCT volume data and a corresponding 2D representation of the corresponding subvolume in the second OCT volume data, For each image pair, the steps include identifying matching unique features of a local set within the corresponding 2D representation of each image pair, The steps include defining alignment transformation parameters for one combination set based on the intrinsic features of a comprehensive set based on the combination of matching intrinsic features of local sets extracted from all of the aforementioned image pairs, The step of electronically processing, storing, or displaying the alignment of the first and second OCT volume data based on the aforementioned set of alignment transformation parameters, A method wherein one or more of the corresponding 2D representations of the plurality of image pairs include a curvature map.
4. The method according to any one of claims 1 to 3, wherein the corresponding 2D representation of each image pair is an en-face structural image, an en-face angiographic image, a thickness map, or a curvature map.
5. The method according to any one of claims 1 to 4, wherein the 2D representations of multiple different image pairs are based on different physical measurements of subvolumes corresponding to the 2D representations.
6. The method according to any one of claims 1 to 5, wherein the subvolume of the first image pair among the plurality of image pairs is a different retinal layer from the retinal layer of the second image pair among the plurality of image pairs.
7. The method according to any one of claims 1 to 6, wherein the first and second OCT volumes are OCT structural volumes or OCT angiography volumes.
8. The method according to any one of claims 1 to 7, wherein the first OCT volume data is of a first region of the sample, and the second OCT volume data is of a second region of the sample, the second region at least partially overlapping with the first region.