Using Multiple Subvolumes, Thickness, and Curvature for OCT / OCTA Data Registration and Retinal Landmark Detection

By using multiple en face and curvature maps from OCT volumes and machine learning, the method addresses registration challenges in OCT/OCTA datasets, ensuring accurate alignment and detection of subtle changes, even in low-quality conditions.

JP7812858B2Active Publication Date: 2026-02-10CARL ZEISS MEDITEC INC +1
4 Cites 0 Cited by

Patent Information

Application Number
JP2023533955
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-04
Filing Date
2021-12-01
Publication Date
2026-02-10
Estimated Expiration
2041-12-01

Smart Images

  • Figure 0007812858000001
    Figure 0007812858000001
  • Figure 0007812858000002
    Figure 0007812858000002
  • Figure 0007812858000003
    Figure 0007812858000003
Patent Text Reader

Abstract

A system / method / device for registering two OCT data sets forms corresponding 2D representations of multiple image pairs of one or more corresponding subvolumes in the two OCT data sets. Matching landmarks in the multiple image pairs are identified as a group, and a set of transformation parameters is defined based on the matching landmarks from all image pairs. The two OCT data sets can then be registered based on the set of transformation parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to optical coherence tomography (OCT) systems, and more particularly to the registration of corresponding OCT data sets. [Background technology]

[0002] Optical coherence tomography (OCT) has become an important modality for ophthalmic examinations. Two-dimensional (2D) representations of three-dimensional (3D) OCT volume data are one OCT visualization technique that has significantly benefited from technological advances in OCT technology. Examples of 2D representations (e.g., 2D maps) of 3D OCT volume data may include, by way of example and not limitation, layer thickness maps, retinal curvature maps, OCT en face images (e.g., en face structural images), and OCTA vasculature maps (e.g., en face OCT angiograms, or functional OCT images).

[0003] Multi-layer segmentation is often used to generate retinal layer thickness maps, en face images, and (2D) vasculature maps. Thickness maps are based on measured thickness differences between retinal layer boundaries. Vasculature maps and OCT en face images can be generated by projecting subvolumes between two layer boundaries, using, for example, averages, sums, percentiles, etc. Therefore, the creation of these 2D maps (or 2D representations of 3D volumes or subvolumes) often relies on the effectiveness of automated segmentation algorithms to identify the layers on which the 2D maps are based.

[0004] To track disease progression over time, it is desirable to compare two corresponding OCT datasets (e.g., OCT volumes and / or two corresponding 2D maps) of the same tissue taken at different times, e.g., during different consultations. This requires measuring changes between two corresponding OCT volumes and / or 2D maps (e.g., thickness maps, en face images, vasculature maps, etc.) over time, which may require alignment of time-series data (acquired from the same subject). However, aligning two corresponding OCT datasets can be a challenging task due to, for example, pathological changes in the retina from one visit to another, layer segmentation errors in the OCT data from one or both visits due to severe pathology that can result in partially inaccurate 2D maps that cannot be accurately aligned, pathological changes affecting one or more subretinal volume data related to a particular disease (e.g., superficial and deeper retinal layers can be affected by diabetic retinopathy), changes in the quality of the OCT data between visits (e.g., when different OCT systems or OCT techniques / modalities with different imaging qualities are used during different visits, or when imaging conditions are not constant), or large lateral or other types of motion. Even if the above problems are avoided, subtle changes over time can be difficult to detect, especially over short periods of time, because the magnitude of the changes can be very small.

[0005] Effective comparison of two OCT data sets over time directly depends on the effectiveness of registration (e.g., alignment of corresponding parts) of the two OCT data sets. Thus, the registration problem primarily consists of aligning corresponding landmarks (e.g., characteristic features) as identified / identified in a pair of (corresponding) 2D maps (e.g., a pair of thickness maps, a pair of en face images, a pair of vasculature maps, etc.). Landmark matching can be a challenging problem due to the aforementioned changes in images over time.

[0006] Landmark-based registration can work if a sufficient number of well-distributed landmark matches are found across each image (e.g., each of a pair of OCT datasets) and an appropriate transformation model is selected. However, one or more of the problems described above can affect the quality of the OCT dataset, resulting in an insufficient number of identifiable landmarks (e.g., unique features) or identified landmarks that are not well-distributed across the 2D map. For example, if a portion of an en face image is of low quality, there is likely to be no identifiable landmarks (or an insufficient number of landmarks) within that low-quality portion, leading to a failure to register the OCT dataset from which the en face image was generated.

[0007] Current approaches to addressing the registration problem focus on attempting to improve the quality of individual en faces or vasculature maps, such as by forming a single en face based on a combination of multiple en faces or vasculature maps. Thus, registration of two OCT datasets is still based on the registration of a single en face or vasculature map. Examples of this approach can be found in Non-Patent Document 1 and Non-Patent Document 2. [Prior art documents] [Non-patent literature]

[0008] [Non-Patent Document 1] Andrew Lang et al., "Combined Registration and Motion Correction of Longitudinal Retinal OCT Data," Proceedings SPIE International Society for Optical Engineering 2016 February 27, p. 9784 [Non-patent document 2] Jing Wu et al., "Stable Registration of Pathological 3D-OCT Scans Using Retinal Vessels," Proceedings of the Ophthalmic Medical Image Analysis International Workshop, September 14, 2014. Summary of the Invention [Problem to be solved by the invention]

[0009] It is an object of the present invention to provide a method / system for improving the reliability of registering multiple pairs of corresponding OCT / OCTA data sets, such as those acquired at different times.

[0010] Another object of the present invention is to improve the registration of corresponding OCT / OCTA datasets of low quality. [Means for solving the problem]

[0011] The above objectives are achieved in an OCT method / system that applies one or more techniques to improve registration quality. The present invention improves the registration of two corresponding OCT volumes (e.g., OCT datasets) by better utilizing more available information, particularly when the OCT volumes are of low quality, the tissue is necrotic, or any of the other issues mentioned above. The applicants have noted that prior art methods for registering OCT volumes by always using the same target slab (e.g., subvolumes bounded by the same two target layers) can affect registration accuracy because the above-mentioned issues can disproportionately affect one or both of these target layers compared to other layers in the OCT volume. In contrast, the main idea of ​​the present approach is to use landmark matching of multiple paired images created from the two OCT volumes to be registered (e.g., multiple 2D maps formed from multiple corresponding subvolumes of the two OCT volumes). The images to be registered can be OCT en face or OCTA vasculature maps, or a combination of OCT en face and OCTA vasculature maps, or thickness maps, or other types of 2D maps. Landmarks can be detected in one pair of 3D OCT / OCTA volumes, and the XY positions of the landmark matches (e.g., as determined from the registered 2D maps) can be used for registration.

[0012] Additionally or alternatively, multiple separate 2D maps can be derived from a single (e.g., parent or base) 2D map (e.g., a macular thickness map). For example, multiple curvature maps (e.g., derived maps or one or more child maps) can be derived from each thickness map (e.g., a base map). The thickness maps and the multiple curvature maps corresponding to the thickness maps can then be used for landmark matching (e.g., registration) to define multiple (registration) transformation parameters / features (or transformation matrices) that define a custom transformation model for registering two OCT datasets.

[0013] Applicants have also found that multiple maps derived from a single (base or parent) map (e.g., multiple curvature maps derived from a macular thickness map) can be used to identify retinal physiological landmarks that would otherwise not be detectable in OCT data with low image quality, necrotic tissue, or many image artifacts (e.g., motion artifacts). For example, low-quality OCT volumes may generally produce low-image-quality en face or vascular images that are not suitable for detecting target physiological landmarks such as the fovea. However, even from such low-quality OCT volumes, a good-quality thickness map can be obtained, and multiple curvature maps derived from the thickness map (e.g., each curvature map based on a different curvature characteristic of the thickness map) can assist in identifying the fovea. That is, the (base or parent) thickness map and its corresponding (derived or multiple child) curvature maps can be used, for example, to estimate the location of the fovea using various machine learning models / techniques.

[0014] Other objects and achievements of the present invention, together with a fuller understanding of the invention, will become apparent and appreciated by reference to the following description and claims taken in conjunction with the accompanying drawings.

[0015] To facilitate the understanding of the present invention, several publications are cited or referenced herein. All publications cited or referenced herein are incorporated by reference in their entirety.

[0016] The embodiments disclosed herein are merely examples, and the scope of the present disclosure is not limited thereto. Features of any embodiment described in one claim category, e.g., a system, may also be claimed in other claim categories, e.g., a method. Dependencies or back-references in the appended claims are selected for formality reasons only. However, any subject matter available from a careful back-reference to a previous claim may also be claimed, thereby disclosing any combination of claims and their features and may be claimed regardless of the dependencies selected in the appended claims. [Brief explanation of the drawings]

[0017] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0018] In the drawings, like reference numbers / letters refer to like elements. [Figure 1] FIG. 1 provides an example of registering corresponding pairs of OCT data sets of an Age-Related Macular Degeneration (AMD) patient taken at two different visits. [Figure 2] FIG. 1 provides an example of registering corresponding pairs of OCT data sets of an age-related macular degeneration (AMD) patient taken at two different visits. [Figure 3] FIG. 1 provides an example of registering corresponding pairs of OCT data sets of an age-related macular degeneration (AMD) patient taken at two different visits. [Figure 4]FIG. 1 provides a comparison of B-scans and en face images taken with a high-quality OCT system and those taken with a low-cost line-field SD-OCT. [Figure 5] FIG. 10 illustrates several pairs of thickness maps that can be used for registration purposes. [Figure 6] FIG. 10 illustrates an example of landmark matching between two thickness maps. [Figure 7] FIG. 10 shows a pair of thickness maps and a plurality of curvature maps corresponding to the pair of thickness maps. [Figure 8A] FIG. 1 shows two exemplary AMD cases from two visits using the Cirrus® device. [Figure 8B] FIG. 1 shows two exemplary AMD cases from two visits using the Cirrus® device. [Figure 9A] 1A-1C provide two examples, one normal and one central serous retinopathy (CSR) case, respectively, from the same examination using a low-cost linefield SD-OCT instrument. [Figure 9B] 1A-1C provide two examples, one normal and one central serous retinopathy (CSR) case, respectively, from the same consultation using a low-cost linefield SD-OCT instrument. [Figure 10] Figure 4 shows macular thickness maps generated by the same low-cost system used to generate the en face images in the second row of Figure 4, illustrating the placement of the Early Treatment Diabetic Retinopathy Study (ETDRS) grid on the fovea. [Figure 11] FIG. 1 illustrates an exemplary network of U-Net architecture suitable for solving the fovea finding problem. [Figure 12] FIG. 1 illustrates a generalized frequency-domain optical coherence tomography system used to collect 3D image data of the eye suitable for use in the present invention. [Figure 13]FIG. 1 shows an exemplary OCT B-scan image of a normal retina of a human eye, illustratively identifying various normal retinal layers and boundaries. [Figure 14] FIG. 1 shows an exemplary en face vasculature image. [Figure 15] FIG. 1 shows an exemplary B-scan vascular image. [Figure 16] FIG. 1 illustrates an example of a multi-layer perceptron (MLP) neural network. [Figure 17] FIG. 1 illustrates a simplified neural network consisting of an input layer, a hidden layer, and an output layer. [Figure 18] FIG. 1 illustrates an exemplary convolutional neural network architecture. [Figure 19] FIG. 1 illustrates an exemplary U-Net architecture. [Figure 20] FIG. 1 illustrates an exemplary computer system (or computing device or computer). DETAILED DESCRIPTION OF THE INVENTION

[0019] When attempting to align OCT volumes, en face or vascular (e.g., OCTA) maps or other (e.g., en face) 2D representations / maps of corresponding slabs (e.g., subvolumes) from each of the OCT volumes may be used, and a single pair of corresponding en face or vasculature maps may be used for alignment. That is, the 2D representations provided by the corresponding en face or vasculature maps are compared to identify similarities for alignment purposes. This can be done by using one or more (known) techniques to identify unique features (e.g., structures, pixels, image characteristics, etc. that can be defined in a manner that is distinguishable from other feature points / regions in the 2D representations), align corresponding unique features (determined from the similarity of their definitions) from the two 2D representations, and define (alignment) transformation parameters or (alignment) transformation matrices that can be applied to the 2D representations to transform the 2D representations (e.g., adjust their positions, angles, shapes, etc.) in a manner that allows at least portions of the 2D representations to be aligned. To properly align the entire 2D representation, a wide distribution of aligned intrinsic features across the entire (or most) area of ​​the 2D representation is desirable. Once the (registration) transformation parameters are defined, they can be ported / applied (e.g., A-scan by A-scan) to their corresponding OCT volumes for registration. If necessary, axial correction can also be applied, such as by axial alignment / registration of corresponding A-scans (from the corresponding OCT volumes). Optionally, the 2D representation used for registration can be formed from a combination of multiple en face (structural) images or vasculature maps (e.g., en face angiography images). That is, multiple en face images or vasculature maps can be combined to generate a single en face image for registration. This can provide a 2D image with additional image clarity, allowing additional intrinsic features to be rendered.

[0020] The present invention provides an alternative method for registering OCT volumes. Unless otherwise stated or understood from the context, hereinafter the term en face may be used interchangeably with the term 2D representation and may include, for example, en face images and / or vessel maps and / or thickness maps and / or other types of 2D maps / images that may be derived therefrom or formed from OCT slabs.

[0021] One embodiment of the present invention uses landmark matching of multiple corresponding (e.g., multiple pairs of) en faces from multiple corresponding OCT volumes (e.g., two or more OCT volumes of the same eye acquired at different examinations and / or different scan times). For example, multiple en faces can be created from multiple different OCT sub-volumes (e.g., defined between different pairs of tissue layers within the OCT volume (e.g., between two retinal layers, or between a target retinal layer and an axial offset from the target retinal layer)). The en faces can be OCT-based en face images or OCTA-based vasculature maps, or a combination of OCT en face images and OCTA vasculature maps, and / or other types of 2D representations or maps (e.g., thickness maps and / or curvature maps). Thus, multiple corresponding (e.g., multiple pairs of) en faces can be formed from multiple sub-volumes of different definitions (e.g., multiple definitions of corresponding sub-volumes), with each pair (or set of corresponding) en faces rendering its own corresponding (local) set of unique features. The unique features (landmarks) of a global set of unique features from all different en faces (or 2D representations / maps) can then be aligned together as a group. This allows for the identification of more consistent unique features across multiple en faces and for discarding unique features that may be due to errors in one or more en faces but are not found in other en faces (e.g., in all or a majority of the en faces, or in a predetermined minimum number or percentage of the en faces). Furthermore, because unique features are identified in multiple independent en faces, considering all (valid) unique features across all en faces as a whole / group improves the probability of identifying a wider distribution of unique features across the overlapping 2D space. In this way, multiple landmarks (unique features) can be detected in a pair of 3D OCT volumes, and the XY locations of landmark matches can be used for 2D registration.

[0022] The technique can use one or more of the following outlined method steps (e.g., as applied to a pair (or more) of OCT / OCTA volumes to be aligned) to identify matching unique features and define a transformation model (e.g., comprised of or based on alignment transformation parameters) for aligning the pair of OCT / OCTA volumes. 1) Selecting two (e.g., corresponding) layers (e.g., the outer boundary of the ILM [internal limiting membrane] and IPL [inner plexiform layer]) and generating 2D images (e.g., a 2D representation of the slab formed by the two selected layers) such as en face images and / or vasculature maps and / or thickness maps for each OCT / OCTA volume. 2) Selecting one of the plurality of 2D images as a reference image. 3) Identifying / defining landmarks (unique features) in the reference image, for example, by randomly sampling the reference image or by using a feature discovery algorithm to find a sufficient number of landmarks. 4) Using a feature discovery / matching algorithm (e.g., template matching) to identify / define corresponding reference landmarks in one or more other 2D images. The feature discovery / matching algorithm can use the same or similar landmark definition techniques as in step 3. 5) Repeating steps 1-4 above to form multiple corresponding 2D images (e.g., en face images and / or vasculature maps) of the OCT / OCTA volumes to be registered. 6) Using all (e.g., valid) landmark matches from all 2D images (en face images and / or vasculature maps), define a comprehensive set of unique features (landmarks) for landmark-based registration. For example, landmark-based registration can involve: a. Using appropriate transformations (e.g., rigid, affine, nonlinear); b. Using RANSAC (random sample consensus) or other feature matching algorithms to select a subset of the best landmark matches; c. selecting some of the best landmark matches (such as by exhaustive search or other suitable method); d. maximizing the distribution of landmarks across the multiple 2D images.

[0023] Figures 1, 2, and 3 provide an example of registering corresponding pairs of OCT datasets (of an age-related macular degeneration (AMD) patient) acquired at two different visits. For illustrative purposes, each of Figures 1, 2, and 3 shows macular thickness maps from the first visit (top right) and the second visit (bottom right). The thickness maps demonstrate their utility for disease progression analysis. For example, Figures 1 and 3 show the extent to which the foveal region may become obscured over time, which may indicate disease progression. In Figures 1, 2, and 3, the first three rows (from the top) show the registration of individual en face images (defined from different subvolumes) by matching a local set of landmarks (unique features) selected by the algorithm ("Landmark Registration," fourth column from the left). For illustrative purposes, in each of Figures 1, 2, and 3, the OCT en face images in the first row are generated based on the ILM and 0.5*(ILM+RPE) layers using mean projection; the OCT en face images in the second row are generated based on the 0.5*(ILM+RPE) and RPE layers using mean projection; and the OCT en face images in the third row are generated based on the RPE and RPE+100 micron layers using mean projection. These layers are suggested because they are often the easiest / quickest to identify, even with low-quality OCT data. In these examples, the last row (fourth from the top) shows the alignment of the en face images in the third row, but using a comprehensive set of unique features (more than 100 landmark matches) selected from multiple local sets of unique features, which may include landmarks from the first two rows (and other matched 2D representations, not shown). Therefore, the last row can be considered the most accurate alignment. Some landmarks selected by the algorithm are shown in the figure (fourth column from the left).

[0024] The example in Figure 1 shows better landmark matching in en face images generated based on the RPE layer and the RPE+100 micron layer. Using all landmarks from all en face images creates a comprehensive set of unique features (landmarks) with the greatest number of landmarks with the greatest distribution (across the 2D space of the image), which can result in more accurate registration.

[0025] The example in Figure 2 shows the performance of our registration method in the presence of segmentation errors and large saccade motion on the second visit. Using all landmarks from all en face images produces the largest number of landmarks with the largest distribution, which again leads to a more accurate registration.

[0026] The example in Figure 3 shows better landmark matching in en face images generated based on the 0.5*(ILM+RPE) layer and the RPE layer. Using all landmarks from all en face images generates the largest number of landmarks with the greatest distribution, which leads to more accurate registration, as shown in the fourth row from the top.

[0027] The method can be extended to 3D OCT / OCTA volumes. One approach involves the following steps: 1) selecting one of a plurality of volumes as a reference volume; 2) finding landmarks within the reference volume by finding a sufficient number of 3D landmarks using a 3D feature discovery algorithm; 3) using a feature matching algorithm (e.g., 3D template matching) to find corresponding reference landmarks in the other volume; 4) 2D registration using 3D landmarks with matching XY coordinates; 5) 3D registration using matching 3D landmarks in XYZ coordinates.

[0028] In addition to (or instead of) forming multiple 2D representations using multiple subvolume definitions, multiple additional 2D representations can also be extracted / derived from an existing 2D map or image. For example, multiple image texture maps, color maps, etc. can be extracted from an en face image (e.g., OCT data) or a vascular map (e.g., OCTA data), and a new 2D representation can be formed from each of the extracted information. However, another example is deriving multiple new 2D representations from an existing thickness map. Because a thickness map inherently contains depth information due to its thickness (e.g., axial / depth) information, but it is customary to display this information in a 2D color format (or grayscale if color is not available), thickness maps are included in the description of this 2D representation herein.

[0029] We have found that even in the case of low-quality OCT data (e.g., when the en face image and / or vascular map do not provide enough information / detail to generate a sufficient number of unique features), it may still be possible to generate a good thickness map. Because retinal thickness does not vary significantly, extracting enough unique features from a sufficiently clear thickness map can still be challenging. However, while a healthy retina has little variation in thickness, due to the roundness of the eyeball, the retina has significant curvature. By deriving multiple curvature maps from a thickness map (preferably of the macula) and using the more clear features of the curvature maps, we can extract a sufficient number of unique features for proper registration of OCT datasets, even those of low quality.

[0030] As explained above, 2D representation of a 3D OCT volume is one OCT visualization technique that has benefited significantly from technological advances in OCT technology. One example of a 2D representation of a 3D OCT volume is a layer thickness map. Multilayer segmentation is used to generate layer thicknesses. Thickness is measured based on the difference between the boundaries of one or more retinal layers.

[0031] [Image registration using thickness and curvature maps] To measure thickness changes over time, registration of time-series data (acquired from the same subject) is necessary. Using OCT en face (and the vascular map), two OCT volumes can be laterally registered. Typically, the similarity between the two images is used in image registration algorithms. However, detecting areas of similarity in lower-quality OCT data, such as those provided by two low-cost line-field OCT en face images, can be difficult due to the low contrast of the en face images, which can make these images unusable for registration purposes.

[0032] Figure 4 provides a comparison of B-scans and en face images taken with a high-quality OCT system (e.g., the Cirrus® system by Carl Zeiss Meditec, Inc. (CZMI)®, Dublin, California, USA) with those of a low-cost line-field SD-OCT. As can be seen, the high-quality OCT system provides clear contrast and image clarity in its B-scan and in each of its three distinct subvolumes, identified herein as superficial, deep, and choroid. In contrast, the low-cost line-field B-scan loses much of the detailed variation provided by the Cirrus system, and the en face images of the B-scan offer little or no differentiation between the different en face images (from the corresponding different subvolumes of the B-scan: superficial, deep, and choroid).

[0033] In contrast, similar regions can be detected in macular thickness maps of the same eye (acquired using a low-cost OCT system). Figure 5 shows multiple pairs of thickness maps that can be used for registration purposes. However, using only thickness maps for registration is problematic for the following reasons: 1) Thickness maps are smooth surfaces, and finding sufficient point correspondences in two thickness maps can be a challenging task, especially in normal cases where thickness map variation is minimal compared to thickness maps of disease cases; 2) This can be a challenging task due to the fact that pathological changes may affect one or more regions of the subretinal volume data associated with a particular disease (e.g., diabetic retinopathy affects both the superficial and deep retinal layers). Therefore, using only the thickness map may affect the accuracy of the registration.

[0034] The registration problem mainly involves aligning corresponding landmarks as identified in a pair of thickness maps. While landmark matching can be a challenging problem due to temporal changes in the two maps, as discussed above, landmark-based registration can perform well if well-distributed landmark matches are found across multiple maps and an appropriate transformation model (e.g., consisting of or based on registration transformation parameters) is selected.

[0035] Figure 6 shows an example of landmark matching between two thickness maps. In this example, the landmark matches are not well distributed throughout the 2D space defined by the thickness maps, which can lead to poor registration quality.

[0036] Therefore, a solution is needed to overcome the above problems that may affect the registration quality. This embodiment uses landmark matching of multiple pairs of maps (e.g., 2D maps). By using multiple pairs of maps, additional multiple point correspondences can be generated at different locations to enhance the distribution of point correspondences. Examples of maps that can be used are the following: 1) Multiple thickness maps of different retinal layers, 2) a thickness map and corresponding multiple curvature maps; 3) A combination of the above steps (1) and (2).

[0037] In general, curvature measures how much a surface is curved in different directions and by different amounts at a surface point. The advantage of using multiple curvature maps (of the retina) is that these maps have higher variance and contrast than the thickness map itself.

[0038] 7 shows a pair of thickness maps and multiple curvature maps corresponding to the pair of thickness maps. In this example, three different curvature maps (mean curvature, maximum curvature, and minimum curvature) are derived from each of the multiple thickness maps. Other examples of curvature maps that can be derived from thickness maps, such as Gaussian, are known in the literature, and any such map(s) can be used in this embodiment.

[0039] Justification An overview of some of the methods using multiple maps is provided herein: To register two OCT volumes of the same eye, 1) generating a macular thickness map for each OCT volume using two layers (e.g., ILM and outer RPE); 2) generating / deriving one or more (e.g., a series of) curvature maps (e.g., Gaussian curvature map, mean curvature map, maximum curvature map, minimum curvature map, etc.) for each macular thickness map; 3) selecting one macular thickness map and corresponding multiple curvature maps as reference maps; 4) Finding landmarks in the reference map, for example, by randomly sampling the reference map or by using a feature discovery algorithm to find a sufficient number of landmarks; 5) using a feature matching / finding algorithm (e.g., template matching) to find corresponding reference landmarks in the second macular thickness map and the corresponding multiple curvature maps; 6) For landmark-based registration, use all (or a combination of) landmark matches from all (or multiple) maps. a. Using appropriate transformations (e.g., rigid, affine, nonlinear); b. Selecting some best landmark matches using RANSAC (random sample consensus) or other methods; c. Using an exhaustive search to select some of the best landmark matches; d. Ensuring maximum landmark distribution across the image.

[0040] 8A and 8B show two exemplary AMD cases from two visits using the Cirrus device. In each example, the first four rows (from the top) show the alignment of individual maps using individual local sets of unique features (landmarks) and showing the match of some landmarks selected by the algorithm. In each example, the bottom (last) row (from the top) shows the alignment of macular thickness maps using a comprehensive set of (e.g., valid) unique features selected from all landmarks (e.g., matches of more than 100 initial landmarks per map). Some (e.g., valid) landmarks selected by the algorithm are shown in the rightmost column of each figure. As shown, using all landmarks from all maps creates the largest number of landmarks with the greatest distribution, which can lead to a more accurate alignment, as shown in the bottom (last) row.

[0041] Figures 9A and 9B provide two examples, one for a normal case and one for a central serous retinopathy (CSR) case, respectively, from the same examination using a low-cost linefield SD-OCT device. In each example, the first four rows (from the top) show the alignment of individual maps using individual local sets of unique features (landmarks) and showing matches of some landmarks selected by the algorithm. The last (e.g., bottom) row shows the alignment of macular thickness maps using a comprehensive set of (e.g., valid) unique features (landmarks) selected from all landmarks (e.g., matches of over 100 initial landmarks per map). In both Figures 9A and 9B, some landmarks selected by the algorithm are shown in the rightmost column.

[0042] The distinction between this approach and that of the above embodiment is that this approach relies more on multiple maps (curvature maps) created / derived from a single (base or parent) map (thickness map), while the above embodiment relies more on multiple independent OCT / OCTA slabs for registration, but it should be understood that both methods can be used jointly / combined (both contributing landmarks / unique features) for registration purposes. For example, landmarks can be detected using OCT / OCTA en face images and vessel maps as well as thickness maps and corresponding multiple curvature maps.

[0043] Some of the above-described embodiments utilize mathematical tools such as surface curvature derived from differential geometry. One advantage of using a thickness map and corresponding curvature map is that the registration and fovea finding are independent of OCT data acquired from different OCT techniques. This can be a convenient way to solve multimodal registration and fovea finding problems, for example, using different OCT systems with different OCT quality and different techniques (e.g., SD-OCT, SS-OCT).

[0044] Typically, en face images and / or vascular maps are used to identify distinct physiological features of the eye, such as the fovea. However, as explained above, low-cost OCT systems may not be able to provide en face images or vascular maps of sufficient contrast / clarity for this purpose. It is proposed herein that thickness maps (and / or their derived curvature maps) may be used to identify distinct physiological features of the eye, such as the fovea.

[0045] [Finding the fovea using thickness and curvature maps] The location of the fovea can be found using the thickness map and corresponding curvature map. The location of the fovea is clinically important because it is the site of highest visual acuity. The location of the fovea is used as a reference point in automated analysis of retinal disease. The fovea has several prominent anatomical features (e.g., vascular patterns and FAZ) that can be used to identify the fovea in OCTA images. The presence of pathologies such as edema, posthyaloid membrane traction, CNV, and other disease types can disrupt the normal foveal structure. Furthermore, this anatomical structure can be disrupted by various outer retinal lesions.

[0046] One use case for foveal location is the placement of the Early Treatment Diabetic Retinopathy Study (ETDRS) grid at the foveal location (see Figure 10). As shown above with reference to Figure 4, the low contrast of OCT en face images acquired using a low-cost linefield SD-OCT system can make them unusable for locating the fovea. In contrast, Figure 10 shows that a macular thickness map (e.g., between the ILM and RPE layers, left image) and corresponding curvature map (not shown) generated by the same low-cost system contain information in the foveal region that can be used to locate the fovea. The thickness map and corresponding curvature map can be used in combination with machine learning techniques to develop a machine model for identifying the fovea when given a thickness map or OCT dataset as input. For example, deep learning methods can be used for automated detection of the foveal center using a macular thickness map and corresponding curvature map. The same concept applies to other retinal landmarks, such as the optic nerve head (ONH).

[0047] In an exemplary embodiment of this embodiment, a convolutional neural network was used to segment the foveal region within the macular thickness map and corresponding curvature map. A description of the machine learning and neural network is provided below. Target training images were generated by creating a 3 mm binary disk mask around the foveal center identified by a human expert grader using OCT data, macular thickness maps, or a combination thereof. One example of a suitable deep learning algorithm / machine / model / system is a (e.g., n-channel) U-net architecture, which uses five convergent and five dilated convolutional layers, ReLU activation, max pooling, binary cross-entropy loss, and sigmoid activation in the final layer. The input channels can be one or more macular thickness maps and their corresponding curvature maps, such as a Gaussian curvature map, a mean curvature map, a maximum curvature map, and a minimum curvature map. In this example, data augmentation (rotation around the center between -9° and 9° in 3° steps) was performed to increase the number of training data. The U-net predicted the ONH region, followed by template matching using a 3 mm diameter disk to find the foveal center. Figure 11 shows an exemplary network with a U-Net architecture that can solve the foveal center finding problem. Other examples of neural networks, including another exemplary U-Net architecture, suitable for use with the present invention to find the fovea (or other retinal physiological structures) are provided below.

[0048] The above-described multiple en face image and thickness map techniques can be used to train a deep learning machine model (e.g., a convolutional neural network (CNN)), which can be used to solve / provide motion estimation using optical flow or to directly estimate transformation parameters (e.g., provide a registration model that can be used to transform a video image to a reference image). Optical flow is a measure of sub-pixel translation between two images. For example, optical flow can be a velocity that describes (e.g., in terms of image pixels) how fast and in what direction a sub-image region (which may represent a feature or object in the image) moves. For example, a network such as FlowNet (e.g., a convolutional (neural) network used to learn optical flow) can be used to calculate optical flow. The CNN can be trained to output transformation parameters that are used to align the video to match the reference image.

[0049] The following describes various hardware and architectures suitable for the present invention. Optical coherence tomography system Generally, optical coherence tomography (OCT) uses low-coherence light to generate two-dimensional (2D) and three-dimensional (3D) internal views of biological tissues. OCT enables in vivo imaging of retinal structures. OCT angiography (OCTA) generates flow information, such as vascular flow, from within the retina. Examples of OCT systems are provided in U.S. Pat. Nos. 6,741,359 and 9,706,915, and examples of OCTA systems are provided in U.S. Pat. Nos. 9,700,206 and 9,759,544, all of which are incorporated herein by reference in their entireties. An exemplary OCT / OCTA system is provided herein.

[0050] FIG. 12 illustrates a generalized frequency-domain optical coherence tomography (FD-OCT) system for collecting 3D image data of the eye suitable for use with the present invention. The FD-OCT system OCT_1 includes a light source LtSrc1. Typical light sources include, but are not limited to, a broadband light source with a short temporal coherence length or a swept laser source. A beam of light from the light source LtSrc1 is typically guided by an optical fiber Fbr1 to illuminate a sample, such as an eye E, a typical sample being human intraocular tissue. The light source LrSrc1 may be, for example, a broadband light source with a short temporal coherence length in the case of spectral-domain OCT (SD-OCT) or a tunable laser source in the case of swept-source OCT (SS-OCT). The light may typically be scanned using a scanner Scnr1 between the output of the optical fiber Fbr1 and the sample E, such that the beam of light (dashed line Bm) is scanned laterally across the region of the sample to be imaged. The light beam from scanner Scnr1 passes through a scan lens SL and an ophthalmic lens OL and can be focused onto the sample E to be imaged. The scan lens SL can receive the light beam from scanner Scnr1 at multiple angles of incidence to generate substantially collimated light, which the ophthalmic lens OL can then focus onto the sample. This example shows a scanning beam that needs to be scanned in two lateral directions (e.g., the x and y directions on a Cartesian plane) to scan a desired field of view (FOV). This example is point-field OCT, which uses a point-field beam to scan across the sample. Thus, scanner Scnr1 is illustratively shown to include two sub-scanners: a first sub-scanner Xscn for scanning the point-field beam across the sample in a first direction (e.g., the horizontal x direction) and a second sub-scanner Yscn for scanning the point-field beam across the sample in an intersecting second direction (e.g., the vertical y direction). If the scanning beam is a line field beam (e.g., line field OCT) and can sample an entire line portion of the sample at a time, only one scanner may be required to scan the line field beam across the sample to span the desired FOV.If the scanning beam is a full-field beam (eg, full-field OCT), a scanner may not be required and the full-field light beam may be illuminated across the entire desired FOV at once.

[0051] Regardless of the type of beam used, light scattered from the sample (e.g., sample light) is collected. In this example, scattered light returning from the sample is collected into the same optical fiber Fbr1 used to route light for illumination. A reference beam derived from the same light source LtSrc1 travels along a separate path, which in this case includes optical fiber Fbr2 and a retroreflector RR1 with an adjustable optical delay. As will be appreciated by those skilled in the art, a transmissive reference path can also be used, with an adjustable delay located in either the sample or the reference arm of the interferometer. The collected sample light is combined with the reference beam, for example, in a fiber coupler Cplr1, to form optical interference within an OCT photodetector Dtctr1 (e.g., a photodetector array, digital camera, etc.). While one fiber port is shown reaching detector Dtctr1, those skilled in the art will appreciate that various interferometer designs can be used for balanced or unbalanced detection of the interference signal. The output from detector Dtctr1 is fed to a processor (e.g., an internal or external computing device) Cmp1, which converts the observed interference into sample depth information. The depth information may be stored in a memory associated with processor Cmp1 and / or displayed on a display (e.g., computer / electronic display / screen) Scn1. The processing and storage functions may be localized within the OCT device, or the functions may be offloaded to (e.g., executed on) an external processor (e.g., an external computer system) to which the collected data is transferred. An example of a computing device (or computer system) is shown in Figure 15. This unit may be dedicated to data processing or may perform other tasks that are quite general and not dedicated to the OCT device.The processor (computing device) Cmp1 may include, for example, a field programmable gate array (FPGA), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a system on a chip (SoC), a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), or combinations thereof, which may perform some or all of the processing steps in a serial and / or parallel manner with one or more host processors and / or one or more external computing devices.

[0052] The sample and reference arms in an interferometer can be constructed of bulk optics, fiber optics, or hybrid bulk optics systems and can have different architectures, such as Michelson, Mach-Zehnder, or common-path designs, as known to those skilled in the art. Light beam, as used herein, should be interpreted as any carefully directed optical path. Instead of mechanically scanning the beam, a light field can illuminate a one-dimensional or two-dimensional area of ​​the retina to generate OCT data (e.g., U.S. Pat. No. 9,332,902; D. Hillmann et al., "Holoscopy-holographic optical coherence tomography," Optics Letters, Vol. 36(13), p. 2290, 2011; Y. Nakamura et al., "High-Speed ​​three dimensional human retinal imaging by line field spectral domain optical coherence tomography," Optics Express, 2011). Express, 15(12), p. 7103, 2007; Blazkiewicz et al., "Signal-to-noise ratio study of full-field Fourier-domain optical coherence tomography," Applied Optics, 44(36), p. 7722 (2005). In time-domain systems, the reference arm must have an adjustable optical delay to create interference. Balanced detection systems are typically used in TD-OCT and SS-OCT systems, while a spectrometer is used at the detection port for SD-OCT systems. The invention described herein can be applied to either type of OCT system.Various aspects of the present invention may be applied to any type of OCT system or other types of ophthalmic diagnostic systems and / or multiple ophthalmic diagnostic systems, including, but not limited to, fundus imaging systems, visual field testing devices, and scanning laser polarimeters.

[0053] In Fourier-domain optical coherence tomography (FD-OCT), each measurement is a real-valued spectrally controlled interferogram (Sj(k)). The real-valued spectral data typically undergoes several post-processing steps, including background subtraction, dispersion correction, etc. A Fourier transform of the processed interferogram yields a complex OCT signal output Aj(z) = |Aj|eiφ. The absolute value of this complex OCT signal, |Aj|, ​​reveals the scattering intensity at different path lengths and, therefore, the scattering profile with respect to depth (z-direction) within the sample. Similarly, the phase φj can also be extracted from the complex OCT signal. The scattering profile with respect to depth is called an axial scan (A-scan). A collection of A-scans measured at adjacent locations within the sample produces a cross-sectional image (tomogram or B-scan) of the sample. A collection of B-scans collected at different lateral locations on the sample constitutes a data volume or cube. For a particular data volume, the fast axis refers to the scan direction along one B-scan, and the slow axis refers to the axis along which multiple B-scans are collected. The term "cluster scan" may refer to a unit or block of data generated by repeated acquisition at the same (or substantially the same) location (or region) to analyze motion contrast, which may be used to identify blood flow. A cluster scan can consist of multiple A-scans or B-scans collected at approximately the same location on the sample at a relatively short time interval. Because the scans in a cluster scan are of the same region, stationary structures remain relatively unchanged between scans in the cluster scan, while motion contrast between scans that meet predetermined criteria may be identified as blood flow.

[0054] Various methods for generating B-scans are known in the art, including, but not limited to, along the horizontal or x-direction, along the vertical or y-direction, along the x and y diagonals, or in a circular or spiral pattern. A B-scan may be in the xz dimension, but may also be any cross-sectional image including the z-dimension. An exemplary OCT B-scan image of a normal retina of a human eye is shown in FIG. 13. An OCT B-scan of the retina provides a view of the structure of the retinal tissue. For illustrative purposes, FIG. 13 identifies the various normal retinal layers and layer boundaries. The identified retinal boundary layers include (from top to bottom) the inner limiting membrane (ILM) layer 1, the retinal nerve fiber layer (RNFL or NFL) layer 2, the ganglion cell layer (GCL) layer 3, the inner plexiform layer (IPL) layer 4, the inner nuclear layer (INL) layer 5, the outer plexiform layer (OPL) layer 6, the outer nuclear layer (ONL) layer 7, the junction between the outer segments (OS) and inner segments (IS) of photoreceptors (indicated by reference numeral 8), the external limiting membrane (ELM or OLM) layer 9, the retinal pigment epithelium (RPE) layer 10, and the Bruch's membrane (BM) layer 11.

[0055] In OCT angiography or functional OCT, analysis algorithms may be applied to OCT data collected at different times (e.g., cluster scans) at the same or nearly the same sample location on the sample to analyze motion or flow (see, e.g., U.S. Patent Application Publication Nos. 2005 / 0171438, 2012 / 0307014, 2010 / 0027857, 2012 / 0277579, and U.S. Patent No. 6,549,801, all of which are incorporated by reference in their entireties). OCT systems may use any one of a number of OCT angiography processing algorithms (e.g., motion contrast algorithms) to identify blood flow. For example, motion contrast algorithms can be applied to intensity information derived from the image data (intensity-based algorithms), phase information from the image data (phase-based algorithms), or complex image data (complex-based algorithms). An en face image is a 2D projection of the 3D OCT data (e.g., by averaging the intensity of each individual A-scan, whereby each A-scan defines a pixel in the 2D projection). Similarly, an en face vasculature image is an image displaying motion contrast signals in which the data dimension corresponding to depth (e.g., the z-direction along the A-scan) is displayed as a single representative value (e.g., a pixel in the 2D projection image), typically by summing or integrating all or isolated portions of the data (see, e.g., U.S. Pat. No. 7,301,644, incorporated herein by reference in its entirety). OCT systems that provide angiography capabilities may be referred to as OCT angiography (OCTA) systems.

[0056] FIG. 14 shows an example of an en face vasculature image. After processing the data and highlighting motion contrast using any of the motion contrast methods known in the art, an en face (e.g., front view) image of the vasculature may be generated by summing pixel ranges corresponding to a tissue depth from the surface of the retinal internal limiting membrane (ILM). FIG. 15 shows an exemplary B-scan of a vasculature (OCTA) image. As shown, structural information may be less clear because blood flow traverses multiple retinal layers, obscuring them more than in a structural OCT B-scan such as that shown in FIG. 13. Nevertheless, OCTA provides a noninvasive technique for imaging the retinal and choroidal microvasculature, which may be important for diagnosing and / or monitoring various pathologies. For example, OCTA can be used to identify diabetic retinopathy by identifying microaneurysms, neovascular complexes, and quantifying the foveal avascular zone and nonperfused areas. Furthermore, OCTA has been shown to show good agreement with fluorescein angiography (FA), a more traditional but less invasive technique that requires the injection of dye to observe vascular flow in the retina. Furthermore, in dry age-related macular degeneration (AMD), OCTA has been used to monitor the overall decrease in choriocapillaris flow. Similarly, in exudative AMD, OCTA can provide qualitative and quantitative analysis of choroidal neovascular membranes. OCTA has also been used to study vascular obstruction, for example, to assess nonperfused areas and the integrity of the superficial and deep plexuses.

[0057] Neural Networks As mentioned above, the present invention may use neural network (NN) machine learning (ML) models. For completeness, neural networks are generally described herein. The invention may use any of the following neural network architectures, alone or in combination: A neural network, or neural net, is a network of interconnected neurons (via nodes), with each neuron representing a node in the network. Collections of neurons may be arranged in layers, with the output of one layer being fed forward to the next layer in a multi-layer perceptron (MLP) arrangement. An MLP may be understood as a feed-forward neural network that maps a set of input data to a set of output data.

[0058] FIG. 16 illustrates an example of a multilayer perceptron (MLP) neural network. The structure may include multiple hidden (e.g., inner) layers HL1 through HLn, which map an input layer InL (which receives a set of inputs (or vector inputs) in_1 through in_3) to an output layer OutL, which generates a set of outputs (or vector outputs), e.g., out_1 and out_2. Each layer may have any number of nodes, which are shown here as circles within each layer for illustrative purposes. In this example, the first hidden layer HL1 has two nodes, and hidden layers HL2, HL3, and HLn each have three nodes. Generally, the deeper the MLP (e.g., the greater the number of hidden layers in the MLP), the greater its learning capacity. The input layer InL may receive vector inputs (shown for illustrative purposes as a three-dimensional vector consisting of in_1, in_2, and in_3) and feed the received vector inputs to the first hidden layer HL1 in the sequence of hidden layers. The output layer OutL receives the output from the last hidden layer in the multi-layer model, say HLn, and produces a vector output result (shown for illustration purposes as a two-dimensional vector consisting of out_1 and out_2).

[0059] Typically, each neuron (i.e., node) generates one output, which is fed forward to neurons in the immediately succeeding layer. However, each neuron in a hidden layer may receive multiple inputs, either from the input layer or from the outputs of neurons in the immediately preceding hidden layer. In general, each node may apply a function to its inputs to generate the output for that node. Nodes in a hidden layer (e.g., the training layer) may apply the same function to each of their inputs to generate their respective outputs. However, some nodes, e.g., nodes in the input layer InL, may receive only one input and be passive, meaning that they simply relay the value of that one input to their output, e.g., they provide a copy of that input to their output; this is indicated by the dashed arrows in the nodes of the input layer InL for illustrative purposes.

[0060] For illustrative purposes, FIG. 17 shows a simplified neural network consisting of an input layer InL′, a hidden layer HL1′, and an output layer OutL′. The input layer InL′ is shown to have two input nodes i1 and i2, which receive inputs Input_1 and Input_2, respectively (e.g., the input nodes of layer InL′ receive a two-dimensional input vector). The input layer InL′ feeds forward into one hidden layer HL1′ with two nodes h1 and h2, which in turn feeds forward into an output layer OutL′ with two nodes o1 and o2. The interconnections, or links, between neurons (shown with solid arrows for illustrative purposes) have weights w1 through w8. Typically, except for the input layer, a node (neuron) may receive as input the output of a node in the previous layer. Each node may calculate its output by multiplying each of its inputs by each input's corresponding interconnection weight, summing the products of the inputs, adding (or multiplying) a constant defined by other weights or biases that may be associated with that particular node (e.g., node weights w9, w10, w11, and w12 corresponding to nodes h1, h2, o1, and o2, respectively), and then applying a nonlinear or logarithmic function to the result. The nonlinear function may be referred to as an activation function or a transfer function. Multiple activation functions are known in the art, and the selection of a particular activation function is not important to this description. However, it should be noted that the operation of an ML model, the behavior of a neural net, depends on the values ​​of the weights, which the neural network may be trained to provide a desired output for a given input.

[0061] During a training, or learning, phase, a neural network learns (e.g., is trained to identify) appropriate weight values ​​to achieve a desired output for a given input. Before a neural network is trained, each weight may be individually assigned an initial (e.g., random, optionally non-zero) value, such as a random number seed. Various methods for assigning initial weights are known in the art. The weights are then trained (optimized) so that, for a given training vector input, the neural network produces an output that approximates a desired (predetermined) training vector output. For example, the weights may be gradually adjusted over thousands of iterative cycles by a method called backpropagation. In each backpropagation cycle, a training input (e.g., a vector input or training input image / sample) is passed forward through the neural network to provide its actual output (e.g., a vector output). The error of each output neuron, or output node, is then calculated based on the actual neuron's output and the supervised training output for that neuron (e.g., a training output image / sample corresponding to the current training input image / sample). It then propagates backward through the neural network (from the output layer back to the input layer), updating the weights based on how much influence each weight has on the overall error, thereby moving the neural network's output closer to the desired training output. This cycle is then repeated until the neural network's actual output is within an acceptable error range of the desired training output for that training input. As will be appreciated, each training input may require many backpropagation iterations to achieve the desired error range. Typically, an epoch refers to one backpropagation iteration (e.g., one forward pass and one backward pass) of all training samples, and training a neural network may require many epochs. Generally, the larger the training set, the better the performance of the trained ML model; therefore, various data augmentation methods may be used to increase the size of the training set.For example, if the training set includes pairs of corresponding training input images and training output images, the training images may be divided into multiple corresponding image segments (or patches). Corresponding patches from the training input images and training output images may be paired to define multiple training patch pairs from one input / output image pair, thereby expanding the training set. However, training a large training set increases the demands on computer resources, such as memory and data processing resources. The computational demands may be reduced by dividing the large training set into multiple mini-batches, the size of which determines the number of training samples in one forward / backward pass. In this case, one epoch may contain multiple mini-batches. Another problem is the possibility that the neural network may overfit the training set, reducing its ability to generalize from a specific input to different inputs. The overfitting problem may be mitigated by creating an ensemble of neural networks or by randomly dropping out nodes in the neural network during training, which effectively removes the dropped leads from the neural network. Various dropout adjustment methods, such as inverse dropout, are known in the art.

[0062] It should be noted that the operation of a trained NN machine model is not a simple algorithm of computation / analysis steps. Indeed, when a trained NN machine model receives an input, the input is not analyzed in the traditional sense. Rather, regardless of the purpose or nature of the input (e.g., vectors defining a live image / scan or vectors defining any other entity such as a demographic description or activity record), the input is subjected to the same architectural construction of the trained neural network (e.g., the same node / layer arrangement, trained weights and bias values, predetermined convolution / deconvolution operations, activation functions, pooling operations, etc.), and it may not be obvious how the architectural construction of the trained network generates its output. Furthermore, the values ​​of the trained weights and biases are not deterministic and depend on many factors, such as the amount of time given to the neural network for training (e.g., the number of epochs in training), the random starting values ​​of the weights before training begins, the computer architecture of the machine on which the NN is trained, the selection of training samples, the distribution of training samples among multiple mini-batches, the selection of activation functions, the selection of error functions that modify the weights, and even whether training is interrupted on one machine (e.g., with a first computer architecture) and completed on another machine (e.g., with a different computer architecture). The point is that the reasons why a trained ML model arrives at a particular output are not clear, and much research is currently being conducted to identify the factors on which ML models base their output. Therefore, the processing of neural networks on live data cannot be reduced to a simple algorithmic step. Rather, the operation depends on the training architecture, training sample set, training sequence, and various circumstances in the training of the ML model.

[0063] In general, constructing a neural network machine learning model may include a learning (or training) stage and a classification (or computation) stage. In the learning stage, a neural network may be trained for a specific purpose and provided with a set of training examples, including training (sample) inputs and training (sample) outputs, and optionally a set of validation examples for testing the training progress. During this learning process, various weights associated with nodes and node interconnections within the neural network are gradually adjusted to reduce the error between the neural network's actual output and the desired training output. In this way, a multi-layer feedforward neural network (such as those described above) may be able to approximate any measurable function to any desired accuracy. The result of the learning stage is a learned (e.g., trained) (neural network) machine learning (ML). In the computation stage, a set of test inputs (or live inputs) may be provided to the learned (trained) ML model, which may apply what it has learned to generate output predictions based on the test inputs.

[0064] Similar to the conventional neural networks of Figures 16 and 17, convolutional neural networks (CNNs) are also composed of neurons with learnable weights and biases. Each neuron receives an input and performs an operation (e.g., a dot product), optionally followed by a nonlinear transformation. However, CNNs receive raw image pixels at one end (e.g., the input) and provide a classification (or class) score at the other end (e.g., the output). Because CNNs expect images as input, they are optimized to handle volumes (e.g., image pixel height and width, and image depth, e.g., color depth, such as RGB depth defined by three colors: red, green, and blue). For example, CNN layers may be optimized for neurons arranged in three dimensions. Neurons in a CNN layer may connect to a small region of the previous layer rather than all of the neurons in a fully connected NN. The final output layer of a CNN may reduce the full image to a single vector (classification) arranged along the depth dimension.

[0065] FIG. 18 provides an exemplary convolutional neural network architecture. A convolutional neural network may be defined as a sequence of two or more layers (e.g., Layer 1 through Layer N), where a layer may include an (image) convolution step, a (resulting) weighted sum step, and a nonlinear function step. The convolution may be performed on the input data by, for example, applying a filter (or kernel) on a moving window over the input data to generate a feature map. Each layer and layer component may have different predetermined filters (from a filter bank), weights (or weighting parameters), and / or function parameters. In this example, the input data may be an image of a certain pixel height and width, or the raw pixel values ​​of this image. In this example, the input image is depicted as having a depth of three color channels, RGB (red, green, blue). Optionally, various preprocessing steps may be performed on the input image, and the results of the preprocessing steps may be input instead of or in addition to the raw image data. Some examples of image processing may include retinal vessel map segmentation, color space conversion, adaptive histogram equalization, connected component generation, etc. Within a layer, a dot product may be calculated between certain weights and their connected small regions within the input volume. While many methods for constructing CNNs are known in the art, by way of example, layers may be configured to apply element-wise activation functions, such as a max(0,x) threshold at zero. Pooling functions may be performed (e.g., along the x and y directions) to downsample the volume. Fully connected layers may be used to identify classification outputs and generate one-dimensional output vectors, which have proven useful in image recognition and classification. However, for image segmentation, CNNs must classify each pixel. Because each CNN layer tends to reduce the resolution of the input image, another stage is required to upsample the image to its original resolution. This may be achieved by applying a transposed convolution (or deconvolution) stage TC, which typically does not use any predetermined interpolation method but instead has learnable parameters.

[0066] Convolutional neural networks have been successfully applied to many problems in computer vision. As mentioned above, training a CNN generally requires a large training dataset. The U-Net architecture is based on a CNN and can generally be trained with a smaller training dataset than traditional CNNs.

[0067] FIG. 19 illustrates an exemplary U-Net architecture. This exemplary U-Net includes an input module (or input layer or stage), which receives an input U-in (e.g., an input image or image patch) of any size. For convenience, the image size at any stage or layer is indicated within a box representing the image; for example, in the input module, the number "128x128" is enclosed, indicating that the input image U-in is composed of 128x128 pixels. The input image may be a fundus image, an OCT / OCTA en face image, a B-scan image, etc. However, it should be understood that the input may be of any size or dimensionality. For example, the input image may be an RGB color image, a monochrome image, a volumetric image, etc. The input image passes through a series of processing layers, each illustrated with exemplary sizes, although these sizes are for illustrative purposes only and may depend, for example, on the size of the image, the convolutional filters, and / or the pooling stage. This architecture consists of a convergent path (exemplary herein including four encoding modules) followed by an augmented path (exemplary herein including four decoding modules), with copy-and-crop links (e.g., CC1-CC4) between corresponding modules / stages that copy the output of one encoding module in the convergent path and connect (e.g., append) it to the upconverted input of the corresponding decoding module in the augmented path. This results in a characteristic U-shape, from which the architecture is named. Optionally, for computational considerations, a "bottleneck" module / stage (BN) can be placed between the convergent path and the augmented path. The bottleneck BN may consist of two convolutional layers (with batch normalization and optional dropout).

[0068] The convergent path is similar to an encoder and typically uses feature maps to capture context (or feature) information. In this example, each encoding module in the convergent path includes two or more convolutional layers, indicated by an asterisk symbol "*," which may be followed by a max-pooling layer (e.g., a downsampling layer). For example, the input image U-in is shown passing through two convolutional layers, each with 32 feature maps. It can be understood that each convolutional kernel produces a feature map (e.g., the output from a convolution operation with a given kernel is an image commonly referred to as a "feature map"). For example, the input U-in passes through a first convolution that applies 32 convolutional kernels (not shown), producing an output consisting of 32 individual feature maps. However, as is known in the art, the number of feature maps produced by a convolution operation can be adjusted (upward or downward). For example, the number of feature maps can be reduced by averaging groups of feature maps, deleting some feature maps, or other known methods of reducing feature maps. In this example, this first convolution is followed by a second convolution whose output is limited to 32 feature maps. Another way to envision the feature maps is to consider the output of the convolution layer as a 3D image whose 2D dimensions are given by the stated XY plane pixel dimensions (e.g., 128x128 pixels) and whose depth is given by the number of feature maps (e.g., the depth of the 32 plane images). Following this illustration, the output of the second convolution (e.g., the output of the first encoding module in the convergence path) can be described as a 128x128x32 image. The output from the second convolution is then subjected to a pooling operation, which reduces the 2D dimensions of each feature map (e.g., the X and Y dimensions can each be reduced by half). The pooling operation can be embodied within a downsampling process, as indicated by the downward arrows. Several pooling methods, such as max pooling, are known in the art, and the particular pooling method is not critical to the present invention.The number of feature maps doubles with each pooling, such as 32 feature maps in the first encoding module (or block), 64 feature maps in the second encoding module, and so on. Thus, the convergence path forms a convolutional network composed of multiple encoding modules (or stages or blocks). As is typical for convolutional networks, each encoding module may provide at least one convolution stage followed by an activation function (e.g., a rectified linear unit (ReLU) or sigmoid layer) (not shown) and a max-pooling operation. Generally, the activation function introduces nonlinearity into the layer (e.g., to avoid overfitting), receives the layer's results, and determines whether to "activate" the output (e.g., whether the value at a particular node meets a predetermined criterion to forward the output to the next layer / node). In summary, the convergence path generally reduces spatial information and increases feature information.

[0069] The extension path is similar to the decoder, notably providing localization and spatial information to the results of the convergence path, despite the downsampling and any max-pooling performed in the contraction stage. The extension path includes multiple decoding modules, each of which combines its current upconverted input with the output of a corresponding encoding module. Thus, features and spatial information are combined in the extension path through a series of upconvolutions (e.g., upsampling or transposed convolutions, i.e., deconvolutions) and combinations (e.g., via CC1-CC4) with high-resolution features from the convergence path. Thus, the output of the deconvolution layer is combined with the corresponding (optionally cropped) feature map from the convergence path, followed by two convolutional layers and activation functions (optionally batch normalized).

[0070] The output from the last augmentation module in the augmentation path may be fed to another processing / training block or layer, such as a classifier block, which may be trained together with the U-Net architecture. Alternatively, or additionally, the output of the last upsampling block (at the end of the augmentation path) may be subjected to another convolution (e.g., output convolution) operation, as indicated by the dotted arrow, before generating its output U-out. The kernel size of the output convolution may be selected to reduce the dimensions of the last upsampling block to a desired size. For example, a neural network may have multiple features on a pixel-by-pixel basis just before reaching the output convolution, which may provide a 1×1 convolution operation to combine these multiple features into a single pixel-by-pixel output value at the pixel-by-pixel level.

[0071] Computing Devices / Systems FIG. 20 illustrates an exemplary computer system (or computing device). In some embodiments, one or more computer systems may provide functionality described or illustrated herein and / or perform one or more steps of one or more methods described or illustrated herein. The computer system may take any suitable physical form. For example, the computer system may be an embedded computer system, a system-on-chip (SOC), or a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, a mesh of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the computer system may reside in a cloud, which may include one or more cloud components within one or more networks.

[0072] In some embodiments, a computer system may include a processor Cpnt1, a memory Cpnt2, a storage Cpnt3, an input / output (I / O) interface Cpnt4, a communication interface Cpnt5, and a bus Cpnt6. The computer system may also optionally include a display Cpnt7, such as a computer monitor or screen.

[0073] The processor Cpnt1 includes hardware for executing instructions, such as those that constitute a computer program. For example, the processor Cpnt1 may be a central processing unit (CPU) or a general-purpose computing-on-graphics processing unit (GPGPU). The processor Cpnt1 may read (or fetch) instructions from an internal register, an internal cache, memory Cpnt2, or storage Cpnt3, decode and execute the instructions, and write one or more results to the internal register, the internal cache, memory Cpnt2, or storage Cpnt3. In particular embodiments, the processor Cpnt1 may include one or more internal caches for data, instructions, or addresses. The processor Cpnt1 may include one or more instruction caches and one or more data caches, for example, to hold data tables. Instructions in the instruction caches may be copies of instructions in memory Cpnt2 or storage Cpnt3, and the instruction caches may speed up retrieval of these instructions by the processor Cpnt1. Processor Cpnt1 may include any suitable number of internal registers and may include one or more arithmetic logic units (ALUs). Processor Cpnt1 may be a multi-core processor or may include one or more processors Cpnt1. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0074] The memory Cpnt2 may include a main memory that stores instructions for the processor Cpnt1 to execute processing or to hold intermediate data during processing. For example, the computer system may load instructions or data (e.g., a data table) from the storage Cpnt3 or from other sources (e.g., another computer system) into the memory Cpnt2. The processor Cpnt1 may load instructions and data from the memory Cpnt2 into one or more internal registers or internal caches. To execute an instruction, the processor Cpnt1 may read and decode the instruction from the internal register or internal cache. During or after the execution of an instruction, the processor Cpnt1 may write one or more results (which may be intermediate or final results) to an internal register, an internal cache, the memory Cpnt2, or the storage Cpnt3. The bus Cpnt6 may include one or more memory buses (each of which may include an ADDRESS bus and a DATA bus) and may couple the processor Cpnt1 to the memory Cpnt2 and / or the storage Cpnt3. Optionally, one or more memory management units (MMUs) facilitate data transfer between the processor Cpnt1 and the memory Cpnt2. The memory Cpnt2 (which may be a high-speed volatile memory) may include random access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM). The storage Cpnt3 may include long-term or high-capacity storage for data or instructions. The storage Cpnt3 may be internal or external to the computer system and may include one or more of a disk drive (e.g., a hard disk drive (HDD) or a solid-state drive (SSD)), flash memory, ROM, EPROM, optical disk, magneto-optical disk, magnetic tape, a universal serial bus (USB)-accessible drive, or other type of non-volatile memory.

[0075] The I / O interface Cpnt4 may be software, hardware, or a combination of both, and may include one or more interfaces (e.g., serial or parallel communication ports) for communicating with I / O devices, which may enable communication with a human (e.g., a user). For example, the I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, table, touch screen, trackball, video camera, other suitable I / O device, or a combination of two or more thereof.

[0076] The communication interface Cpnt5 may provide a network interface for communicating with other systems or networks. The communication interface Cpnt5 may include a Bluetooth interface or other types of packet-based communication. For example, the communication interface Cpnt5 may include a network interface controller (NIC) and / or a wireless NIC or wireless adapter for communication with a wireless network. The communication interface Cpnt5 may provide communication with a Wi-Fi network, an ad-hoc network, a personal area network (PAN), a wireless PAN (e.g., Bluetooth WPAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a cellular network (e.g., a Global System for Mobile Communications (GSM) network), the Internet, or a combination of two or more thereof.

[0077] Bus Cpnt6 may provide a communication link between the above-mentioned components of the computing system. For example, bus Cpnt6 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand bus, a low-pin-count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or any other suitable bus, or a combination of two or more thereof.

[0078] Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0079] As used herein, a computer-readable non-transitory storage medium may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, or any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more thereof, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.

[0080] While the present invention has been described in conjunction with several specific embodiments, as will be apparent to those skilled in the art in light of the foregoing description, many other alternatives, modifications, and variations will be apparent. Accordingly, the invention as described herein is intended to embrace all such alternatives, modifications, applications, and variations that may fall within the spirit and scope of the appended claims.

Claims

1. 1. A method for registering first optical coherence tomography (hereinafter OCT) volume data to second OCT volume data, comprising: generating a plurality of image pairs, each image pair including a generated two-dimensional (hereinafter referred to as 2D) representation of a sub-volume in the first OCT volume data and a corresponding 2D representation of a corresponding sub-volume in the second OCT volume data; for each image pair, identifying a local set of matching unique features in the corresponding 2D representations of each image pair; defining a set of alignment transformation parameters based on a global set of matching characteristic features based on all local sets of matching characteristic features extracted from all of the image pairs; electronically processing, storing, or displaying the registration of the first and second OCT volume data based on the set of registration transformation parameters; A method wherein the different image pairs include a combination of two or more of en face structural images, en face angiographic images, thickness maps, and curvature maps.

2. The method of claim 1 , wherein the corresponding 2D representation of each image pair is an en face structural image, an en face angiographic image, a thickness map, or a curvature map.

3. The method of claim 1 or 2, wherein the 2D representations of different image pairs are based on different physical measurements of the sub-volumes corresponding to the 2D representations.

4. The method of claim 1 , wherein a sub-volume of a first image pair of the plurality of image pairs is different from a sub-volume of a second image pair of the plurality of image pairs.

5. The method of claim 1 , wherein at least some of the plurality of image pairs include a first image pair and one or more derived image pairs based on the first image pair.

6. The method of claim 5 , wherein the first image pair are corresponding thickness maps, and the one or more derived image pairs are corresponding curvature maps based on the corresponding thickness maps.

7. 6. The method of claim 5, wherein the first image pair are corresponding en face images, and the one or more derived image pairs are based on one or more of image texture, color, intensity, contrast, and negative images of the en face images.

8. The method of claim 1 , wherein the first and second OCT volumes are OCT structural volumes or OCT angiography volumes.

9. 9. The method of claim 1, wherein the first OCT volume data is of a first region of a sample and the second OCT volume data is of a second region of the sample, the second region at least partially overlapping the first region.

10. 1. A method for registering optical coherence tomography (OCT) data, comprising: accessing a first OCT volume of data for a first region of a sample; accessing a second OCT volume data of a second region of the sample that at least partially overlaps the first region; generating a first set of property maps based on corresponding property measurements of one or more sub-volumes of the first OCT volumetric data; generating a second set of characteristic maps, each map in the second set having a one-to-one correspondence with a map in the first set, and each map in the second set being based on a corresponding characteristic measurement of a corresponding one or more sub-volumes of the second OCT volume data; registering corresponding maps in the first and second sets to each other as a group to identify registration parameters for the first and second OCT volume data; storing or displaying the alignment of the first and second OCT volume data; The method, wherein the first and second sets of property maps include a combination of two or more of en face structural images, en face angiographic images, thickness maps, and curvature maps.

11. The method of claim 10 , wherein the characteristic maps include one or more thickness maps of the first and second OCT volume data.

12. The method of claim 10 , wherein the property maps include one or more curvature maps of the first and second OCT volume data.

Citation Information

Patent Citations

  • Ophthalmic image processing device and program

    JP2013154121A

  • Ophthalmologic imaging apparatus, control method thereof, program, and recording medium

    JP2019054990A

  • Image processing device, image processing method and program

    JP2020032072A

  • Optical coherence tomography multiple en-face angiographic averaging system, method and apparatus

    JP2020513983A