Real-time tracking of ir fundus images using reference landmarks in presence of artifacts

CN115515474BActive Publication Date: 2026-09-22CARL ZEISS MEDITEC INC +1
View PDF 17 Cites 0 Cited by

Patent Information

Application Number
CN202180032024.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-29
Filing Date
2021-04-29
Publication Date
2026-09-22
Estimated Expiration
2041-04-29

AI Technical Summary

Technical Problem

这些复杂算法的添加可能会阻碍它们在实时应用中的使用,特别是如果需要对高分辨率(例如,大)图像执行追踪算法用于更准确的追踪

Benefits of technology

[0015]与现有技术相比,本方法的优点在于,在实时图像中的特定区域/窗口中搜索和检测界标,其位置和/或尺寸相对于参考界标被限定(界标被独立地跟踪)。因此,可以实现对高分辨率IR图像进行最少预处理的实时追踪,并且在追踪过程中不需要复杂的图像处理技术。由于追踪是相对于参考点限定的,本发明对图像伪影(例如条纹伪影、中央反射伪影、眼睑等)的存在也不太敏感。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115515474B_ABST
    Figure CN115515474B_ABST
Patent Text Reader

Abstract

A system and method for ophthalmic motion tracking. Anchor points and a plurality of secondary points are selected from a reference image. A single live image in a series of images is then searched to match the anchor points and secondary points. First, the anchor points are found, and then the search for individual secondary points is limited to within a search window defined by a known distance and / or direction of the secondary points sought relative to the anchor points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention is generally aimed at motion tracking. More specifically, it targets ocular motion tracking of the anterior and posterior segments of the eye. Background Technology

[0002] Fundus imaging, which can be obtained using a fundus camera, typically provides a frontal view of the fundus seen through the pupil. Fundus imaging can use light of different frequencies, such as white, red, blue, green, infrared (IR), etc., to image tissues, or can use selected frequencies to excite fluorescent molecules in certain tissues (e.g., autofluorescence) or to excite fluorescent dyes injected into the patient (e.g., fluorescein angiography). A more detailed discussion of different fundus imaging techniques is provided below.

[0003] OCT is a non-invasive imaging technique that uses light waves to generate cross-sectional images of retinal tissue. For example, OCT allows people to view unique tissue layers of the retina. Typically, an OCT system is an interferometric imaging system that determines the scattering profile of a sample along an OCT beam by detecting the interference between light reflected from the sample and a reference beam, thus creating a three-dimensional (3D) representation of the sample. Each scattering profile in the depth direction (e.g., z-axis or axial direction) can be reconstructed individually as an axial scan or A-scan. Cross-sections, two-dimensional (2D) images (B-scans), and extended 3D volumes (C-scans or cube scans) can be constructed from multiple A-scans acquired as the OCT beam scans / moves through a set of lateral (e.g., x-axis and y-axis) positions on the sample. OCT also allows the construction of planar frontal views (e.g., frontal) 2D images of selected portions of a tissue volume (e.g., a target tissue plate (sub-volume) or target tissue layer of the retina). OCTA is an extension of OCT that can identify (e.g., presented in image format) the presence or absence of blood flow in tissue layers. OCTA can identify blood flow by recognizing differences (e.g., contrast differences) over time in multiple OCT images of the same retinal region, and designates differences that meet predefined criteria as blood flow. A more in-depth discussion of OCT and OCTA is provided below.

[0004] Real-time and efficient tracking of fundus images (e.g., infrared (IR) fundus images) is important in automated retinal OCT image acquisition. Retinal tracking is particularly important due to involuntary eye movements during image acquisition, especially between OCT and OCT-TA scans.

[0005] Infrared (IR) images can be used to track retinal motion. However, insufficient IR image quality and the presence of various artifacts can affect automated real-time processing, thus reducing success rate and reliability. The quality of IR images varies significantly over time, depending on fixation, focus, vignetting effects, eyelids, streaks, and central reflection artifacts. Therefore, a method is needed that can robustly track the retina in real time using IR images. Figure 1 Exemplary IR fundus images with various artifacts are provided, including stripe artifact 11, central reflection artifact 13, and eyelid 15 (e.g., considered as a shadow).

[0006] Current tracking systems use a reference image with a set of landmarks extracted from the image. The tracking algorithm then tracks the live image by searching for landmarks in each live image using the landmarks extracted from the reference image. Landmark matching between the reference and live images is determined independently. Therefore, matching becomes challenging due to artifacts such as stripes and central reflection artifacts in the images. Typically, complex image processing algorithms are needed to augment the image before landmark detection. If the tracking algorithm needs to be performed on high-resolution images for more accurate tracking, the real-time performance of the tracking algorithm will be affected by using these additional algorithms.

[0007] In summary, existing tracking systems use reference fundus images with a set of landmarks extracted from reference images. The tracking algorithm then tracks a series of live images using the landmarks extracted from the reference images by independently searching for each landmark in each live image. Landmark matching between the reference and live images is determined independently. Therefore, landmark matching becomes challenging due to artifacts in the images (e.g., streaks and central reflection artifacts). Complex image processing algorithms are required to enhance the IR images before landmark detection. The addition of these complex algorithms can hinder their use in real-time applications, especially if a tracking algorithm is needed to perform tracking on high-resolution (e.g., large) images for more accurate tracking.

[0008] The purpose of this invention is to provide a more efficient system / method for ophthalmic motion tracking.

[0009] Another object of the present invention is to provide real-time ophthalmic motion tracking using high-resolution images. Summary of the Invention

[0010] The aforementioned objectives are achieved in a method / system for eye tracking. Unlike prior art tracking systems, this system does not independently search for matching landmarks. Instead, the present invention identifies a reference (anchor) point / template (e.g., a landmark) and matches additional landmarks at positions relative to the reference point. The landmarks can then be detected in a live IR image relative to the reference (anchor) point or template acquired from a reference image.

[0011] Reference points can be selected as prominent anatomical / physical features that are easily identifiable and expected to appear in subsequent live images (e.g., optic nerve head (ONH), lesions, or specific vascular patterns, if the posterior segment of the eye is being imaged). Alternatively, for example, if no prominent and consistent anatomical features are present (e.g., if the anterior segment of the eye is being imaged), reference anchors can be selected from a pool / group of candidate reference points based on the current state of the series of images. As the quality of the series of live images changes, or different prominent features emerge, the reference anchors are modified / changed accordingly. Thus, in this alternative embodiment, reference anchors can change over time based on the captured / collected images.

[0012] It should be understood that a reference point or template may include one or more characteristic features (pixel identifiers) that collectively define (e.g., for identification) the specific landmark used as a reference physical landmark (e.g., ONH, lesion, or specific vascular pattern). The distance between the reference point and the selected landmark in the reference IR image and the live IR image remains constant in both images (or their relative distance remains constant). Therefore, landmark detection in the live IR image becomes a simpler problem by searching small regions (e.g., bound regions or windows of a predefined / specific distance from the reference point) in relation to the reference point. The robustness of landmark detection is improved / enhanced because the distance between the reference point and the landmark location is constant. As the initial landmark is matched, the search for additional landmarks may be further limited to specific directions / orientations / angles relative to already matched landmarks (e.g., in addition to specific distances).

[0013] This method eliminates the need for complex image processing algorithms to enhance IR images, thus improving speed, especially when processing high-resolution images. In other words, it ensures real-time performance of the tracking algorithm on high-resolution images, enabling more accurate tracking.

[0014] In summary, the present invention can begin by detecting salient points (e.g., ONH locations or other points such as lesions) in a selected reference IR image (or other imaging modality) using methods such as deep learning or knowledge-based computer vision. Additional templates / points offset from the center of the reference points are extracted from the reference IR image to increase the number of templates and corresponding landmarks in the IR image. Optionally, multiple reference / anchor points are feasible for general purposes. For example, multiple images of the eye can be captured, including a reference image and one or more real-world images. Multiple reference anchors can then be defined in the reference image, along with one or more auxiliary points. Then, within the selected real-world image, multiple initial matching points that match (all or part) the multiple reference anchors are identified, and the selected real-world image is transformed into a reference image based on (e.g., using) the identified multiple reference anchors, according to coarse registration. After providing this coarse registration, matches of the selected auxiliary points within the region can be searched based on (e.g., defined within a search area / FOV / window-based range) the position of the selected auxiliary points relative to the multiple matched reference anchors. The tracking error between the reference image and the selected live image can then be corrected based on their matching points. This approach can be helpful when there are significant geometric transformations between the reference and live images during tracking, and when used in more complex tracking systems. For example, if there is a large rotation (or affine / projective relationship) between two images, more than two anchor points can be coarsely registered first to more accurately search for other landmarks. The tracking algorithm then uses a template centered on the reference point and additional templates extracted from the reference IR image to track the live IR image (or other corresponding imaging modalities). Suppose a set of templates is extracted from the reference IR image, and their corresponding positions in the live IR image (as a set of landmarks) can be determined by template matching in small regions far from the reference point in the live IR image. Full or partial matching can be used to calculate the transformations (x and y translations, rotations, affines, projections, and / or nonlinear transformations) between the IR reference image and the live IR image. In this way, landmarks (e.g., matching templates or matching points) are detected relative to the reference landmarks. This results in real-time operation (e.g., processing is limited to a small area of ​​the image) and robust tracking (e.g., false alarms are eliminated because the distance between reference landmarks and landmarks is known, and additional checks are provided to verify the validity of candidate matches).

[0015] Compared to existing technologies, the advantage of this method lies in its ability to search for and detect landmarks within specific regions / windows in real-time images, whose positions and / or sizes are defined relative to reference landmarks (the landmarks are tracked independently). Therefore, real-time tracking of high-resolution IR images with minimal preprocessing is possible, and complex image processing techniques are not required during the tracking process. Since the tracking is defined relative to a reference point, this invention is also less sensitive to the presence of image artifacts (e.g., streak artifacts, central reflection artifacts, eyelid artifacts, etc.).

[0016] The invention can be further extended to tracking regions / windows of moving fundus images (e.g., IR images) that have a given FOV that at least partially overlaps with the FOV of an OCT scan. The IR FOV can be moved around (while maintaining overlap with the OCT FOV) until a position is reached that includes a maximum number (or a sufficient number) of easily / robustly identifiable anatomical features (e.g., ONH, lesions, or specific vessels) for robust tracking. This tracking information can then be used to provide motion compensation to the OCT system for OCT scanning.

[0017] Therefore, when the eye is stable (without movement or minimal movement), the reference image can be used to align and trigger automatic capture (e.g., from an OCT system) and robustly track retinal image sequences.

[0018] This invention provides various metrics to quantify the quality of tracked images and the possible causes of poor tracking. Therefore, this invention can extract various statistical data from historical data to identify various characteristic problems affecting tracking. For example, this invention can analyze image sequences used for tracking and determining whether images exhibit characteristics of systematic motion, random motion, or good fixation. The ophthalmological system of this invention can then be used to inform the system operator or patient of potential problems affecting tracking and provide suggested solutions.

[0019] Other objects and achievements, as well as a fuller understanding of the invention, will become apparent and understood by taking into account the accompanying drawings and the following description and claims.

[0020] To facilitate understanding of this invention, several disclosures may be referenced or mentioned herein. All disclosures referenced or mentioned herein are incorporated herein by reference in their entirety.

[0021] The embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited thereto. Any embodiment feature mentioned in one claim class, such as a system, may also be claimed in another claim class, such as a method. Dependencies or references in the appended claims are chosen solely for formal reasons. However, any subject matter arising from an intentional retrospective to any prior claim, thereby disclosing any combination of claims and their features, may also be claimed and may be claimed regardless of the dependency chosen in the appended claims. Attached Figure Description

[0022] In the accompanying drawings, the same reference symbols / characters denote the same parts:

[0023] Figure 1 Exemplary infrared (IR) fundus images with various artifacts are provided, including stripe artifact 11, central reflection artifact 13, and eyelid 15 (e.g., considered as a shadow).

[0024] Figure 2 and Figure 3 Examples of tracking frames are shown (each frame includes an exemplary reference image and an exemplary live image), in which reference points and a set of landmarks (from the reference image) are tracked in the live IR image.

[0025] Figure 4 This describes the tracking system in the case where eyelid 31 and central reflection 33 are present in the live IR image 23.

[0026] Figure 5 Two additional examples of the invention are shown.

[0027] Figure 6 The statistical data of registration error and eye movement test results are shown for different acquisition modes and motion levels.

[0028] Figure 7 An exemplary anterior segment image is provided, which shows changes (in a timely manner) in pupil size, iris pattern, eyelid and eyelid movement, and lack of contrast in the eyelid region during the same acquisition.

[0029] Figure 8A , Figure 8B and Figure 9A , Figure 9B and Figure 9C This illustrates the tracking of reference points and a set of landmarks in the live image.

[0030] Figure 10 An example of a tracking algorithm according to an embodiment of the present invention is illustrated.

[0031] Figure 11 Provided Figure 10 The tracking test results of the embodiments.

[0032] Figure 12 This explains how to use ONH to determine the optimal tracking FOV (tracking window) position for a given OCT FOV (acquisition / scanning window).

[0033] Figure 13 The second embodiment of the present invention is described, which is used to determine the optimal tracking FOV location without using OHN or other predefined physiological landmarks.

[0034] Figure 14A , Figure 14B , Figure 14C and Figure 14D Additional examples of this method for identifying the optimal tracking FOV relative to the OCT FOV are provided.

[0035] Figure 15 Scenario 1 is illustrated, in which a reference image of the retina from the gaze of a previously visited patient is available.

[0036] Figure 16 Scenario 2 is illustrated, in which the retinal image quality algorithm detects a reference image during the initial alignment process (either by an operator or automatically).

[0037] Figure 17 Scenario 3 is illustrated, in which the quality algorithms from the previously accessed reference image and retinal image are unavailable.

[0038] Figure 18 An alternative solution for scenario 3 is explained.

[0039] Figure 19A Two examples are shown: small (top) and normal (bottom) pupil acquisition modes.

[0040] Figure 19B Statistics on registration error, eye movement, and number of keypoints are shown for a total of 29,529 images from 45 image sequences.

[0041] Figure 20A This describes the motion of the current image (white border) relative to the reference image (gray border), with eye motion parameters Δx, Δy, and rotation φ relative to the reference image.

[0042] Figure 20B Examples from three different patients are shown: one with good fixation, another with systematic eye movements, and the third with random eye movements.

[0043] Figure 21 A table was provided showing eye movement statistics for 15 patients.

[0044] Figure 22An example of a slit-scan ophthalmic system used for imaging the fundus is illustrated.

[0045] Figure 23 A general frequency-domain optical coherence tomography system for collecting 3D image data of an eye suitable for use with the present invention is described.

[0046] Figure 24 An exemplary OCT B scan image of a normal retina of the human eye is shown, and various typical retinal layers and boundaries are illustratively identified.

[0047] Figure 25 An example of a frontal vascular system image is shown.

[0048] Figure 26 An exemplary B-scan of the vascular system (OCTA) image is shown.

[0049] Figure 27 An example of a multilayer perceptron (MLP) neural network is illustrated.

[0050] Figure 28 This illustrates a simplified neural network consisting of an input layer, hidden layers, and an output layer.

[0051] Figure 29 This illustrates an example convolutional neural network architecture.

[0052] Figure 30 The example U-Net architecture is illustrated.

[0053] Figure 31 This describes an example computer system (or computing device or computer). Detailed Implementation

[0054] This invention provides an improved eye-tracking system, such as for fundus cameras, optical coherence tomography (OCT) systems, and OCT angiography systems. The invention is described herein using an infrared (IR) camera that tracks any eye in a series of live images, but it should be understood that the invention can use other imaging modalities (e.g., color images, fluorescence images, OCT scans, etc.).

[0055] This tracking system / method can begin by first identifying / detecting (e.g., salient) reference points (e.g., significant physical features that can be consistently (e.g., reliably and / or easily and / or quickly) identified in an image). For example, a reference point (or reference template) could correspond to the optic nerve head (ONH) (and its reference location) or another significant / consistent point / feature, such as a lesion. Reference points can be selected from the reference IR image using deep learning or other knowledge-based computer vision methods.

[0056] Alternatively or additionally, a set of candidate points can be identified in the image stream, and the most consistent candidate points within a set of images can be selected as reference anchor points for the series of live images. In this way, the anchor points / templates used in the series of live images can vary as the quality of the live image stream changes and as the different candidate points become more significant / consistent.

[0057] This tracking algorithm uses templates centered on a reference point extracted from a reference IR image to track a live IR image. Additional templates offset from the reference point center are extracted to increase the number of templates and corresponding landmarks in the IR image. These templates can be used to detect the same locations in different IR images as a set of landmarks, which can be used for registration between the reference image and the live IR image, thus enabling real-time tracking of a series of IR images. The advantage of generating a set of templates by offsetting the reference location is that it eliminates the need for vessel enhancement or complex image feature detection algorithms. However, if the tracking algorithm needs to be performed on high-resolution images for more accurate tracking, the real-time performance of the tracking algorithm will be affected by using these additional algorithms. Assuming a set of templates is extracted from the reference IR image, their corresponding locations in the live IR image (as a set of landmarks) can be determined in the live IR image in small boundary regions far from the reference point through template matching (e.g., normalized cross-correlation). Optionally, if no match is found with the initial set of templates, more templates can be searched. Once a corresponding match is found, all matches or subsets of matches can be used to compute the transformation (x and y shifts and rotations) between the IR reference image and the live IR image. Optionally, if the number of matches is no greater than a threshold (e.g., half of the identified landmarks), the current live image is discarded and no correction is made for the tracking error. Assuming sufficient matches are found, the transformation determines the amount of motion between the live IR image and the reference image. Theoretically, the transformation could be computed using two corresponding landmarks (a reference point and a landmark with high confidence) from the IR reference and the live image. However, more than two landmarks will be used for tracking to ensure more robust tracking.

[0058] Figure 2 , Figure 3 and Figure 4 Examples of tracking frames are shown (each frame includes an exemplary reference image and an exemplary live image), where reference points and a set of landmarks (from the exemplary reference image) are tracked in the live IR image. Figure 2 , Figure 3 and Figure 4In each of these, the top image (21A, 21B, and 21C, respectively) in each tracking frame is an exemplary reference IR image, and the bottom images (23A, 23B, and 23C, respectively) are exemplary live IR images. The dashed boxes are ONH templates, and the white boxes are the corresponding templates in the IR reference image and the live image. The template can be adaptively selected for each live IR image. For example, Figure 2 and Figure 3 The same reference images 21A / 21B and the same anchor point 25 are shown, but Figure 2 Additional boundary marker 27A and Figure 3 The additional markers 27B in the image are different. In this case, the markers are dynamically selected in each IR live image based on their detection confidence.

[0059] Figure 4 This describes the tracking system in the case where eyelid 31 and central reflection 33 are present in the live IR image 23. Note that in this case, tracking does not depend on the presence of blood vessels (as opposed to...). Figure 2 and Figure 3 (The opposite of the example) thus avoids confusing the eyelid with blood vessels.

[0060] Figure 5 Two additional examples of the invention are shown. The top row of images shows the invention applied to normal pupil acquisition, and the bottom row shows the invention applied to small pupil acquisition. In both cases, the landmarks are detected relative to the ONH location (e.g., a reference anchor point). In this example, tracking parameters (xy translation and rotation) are calculated through registration between the reference image (RI) and the moving image (MI). This method requires the ONH 41 location extracted from a feature-rich region of the reference image RI and a set of RI landmarks (e.g., auxiliary points) 43. For illustrative purposes, one of the landmarks 43 is shown within a restricted region or search window 42 and its relative distance 45 to the ONH 41. The ONH 41 in the reference image RI can be detected using a neural network system with a U-Net architecture. A general discussion of neural networks (including the U-Net architecture) is provided below. The ONH 41' in a moving (e.g., live) image MI can be detected using template matching with the ONH template 41 extracted from the reference image RI. Each reference marker template 43 and its relative distance 45 to ONH 41 (and optionally its relative orientation) are used to search for corresponding markers 43' in the motion image MI that have the same / similar distance 45' from ONH 41' (e.g., within a restricted region or window 42'). A set of markers corresponding to high confidence is used to calculate the tracking parameters.

[0061] In an exemplary implementation, infrared (IR) images (11.52 x 9.36 mm, pixel size 15 μm / pixel) were acquired using a CLARUS 500 (ZEISS, Dublin, CA) at a frame rate of 50 Hz, employing normal and small pupil acquisition modes with induced eye movement. Each eye was scanned using three different levels of motion: good gaze, systematic, and random eye movement. The registered images are displayed in a single image for visualizing the registration (see [link to documentation]). Figure 5 The average distance error between the registration movement and the reference landmark was calculated as the registration error. Statistics on the registration error and eye movement for each acquisition mode and each level of motion are reported. Approximately 500 images were collected from 15 eyes.

[0062] Figure 6 The results show statistical data on registration errors and eye movements for different acquisition modes and levels of motion using all eyes. The mean and standard deviation of registration errors are similar for normal and small pupil acquisition modes, indicating that the tracking algorithm has similar performance in both modes. The reported registration errors are important information for designing OCT scanning patterns. Using a computing system equipped with an Intel i7 CPU, 2.6GHz, and 32GB RAM, the average tracking time for a single image was measured to be 13 milliseconds. Therefore, this invention provides a real-time retinal tracking method using IR images with good tracking performance, which is an important component of OCT image acquisition systems.

[0063] Although the above examples are described as being applied to the posterior segment of the eye (e.g., the fundus), it should be understood that the invention can also be applied to other parts of the eye, such as the anterior segment. Real-time and efficient tracking of the anterior segment image is crucial in automated OCT angiography image acquisition. Anterior segment tracking is essential due to involuntary eye movements during image acquisition, especially in OCTA scans. Anterior segment LSO images can be used to track the movement of the anterior segment of the eye. It can be assumed that eye movement is a rigid motion with motion parameters (e.g., translation and rotation) that can be used to control the OCT beam.

[0064] Localized movements and lack of contrast in the anterior portion of the eye (such as changes in pupil size / shape, consistent eyelash and eyelid movements, and iris patterns that are squeezed or enlarged during tracking (due to changes in pupil size / shape)) can affect automated real-time processing, thereby reducing the success rate and reliability of tracking. Furthermore, the appearance of anatomical features in an image can change significantly over time, depending on the subject's gaze (e.g., the angle of gaze). Figure 7Exemplary anterior segment images are provided, showing variations in pupil size, iris pattern, eyelid and eyelash movements, and a lack of contrast in the eyelid region (in real-time) during the same acquisition. Therefore, a method is needed to robustly track the anterior segment of the eye in real-time using LSO or other imaging techniques.

[0065] Previous front-end tracking systems used a reference image with a set of landmarks extracted from the image. The tracking algorithm then tracked a series of real-world images by independently searching for landmarks in each real-world image, using the landmarks extracted from the reference image. That is, the matching landmarks between the reference and real-world images were determined independently. Assuming a rigid (or affine) transformation, independent landmark matching between two images becomes challenging due to local motion and lack of contrast. Complex landmark matching algorithms are typically required to compute the rigid transformation. If higher-resolution images are needed for more accurate tracking, the real-time tracking performance of this previous method is often compromised.

[0066] The above tracking embodiments (for example, see Figures 2 to 6 This approach provides efficient landmark matching detection between two images, but some implementations may have limitations. Some of the hypothetical reference and live images in the above embodiments contain obvious or unique anatomical features, such as ONHs, which are robustly detected due to the uniqueness of the anatomical features. In this method, landmarks are detected relative to reference (anchor) points (e.g., ONHs, lesions, or specific vascular patterns) in the live image. The distance (or relative distance) between the reference point in the reference image and the selected landmark remains constant in both images. Therefore, landmark detection in the live image becomes a simpler problem by searching small regions at known distances (and optionally, in direction) from the reference (anchor) point. The robustness of landmark detection is guaranteed because the distance between the reference point and the landmark location is constant.

[0067] In contrast, this embodiment offers several advantages over the embodiments described above. Similar to the embodiments described above, references (e.g., anchor points) are selected from candidate landmarks extracted from a reference image; however, the selected reference anchor points are not necessarily obvious or unique anatomical / physical features of the eye (e.g., ONH, pupil, iris boundary, or center). Although the anchor points may not be unique anatomical / physical features, the distance between the reference anchor points and the selected auxiliary landmarks remains constant in both the reference and live images. Therefore, landmark detection in the live image becomes a simpler problem by searching small regions at known distances from the reference anchor points. Robustness of landmark detection is guaranteed because the distance between the reference anchor points and the auxiliary landmark points is constant. A subset of best-matching landmarks can be selected to compute the rigid transformation through an exhaustive search of subsets. A similar method can also be applied to retinal tracking using IR images (as described in the embodiments above), where unique anatomical landmarks are invisible (or cannot be found) in the image / scan region or within the detector's field of view (e.g., periphery).

[0068] The difference between this embodiment and some embodiments described above is that the reference (anchor) point is selected from a set of landmark candidates extracted from a reference image. Reference points can be selected / picked based on traceable reference points in subsequent images (e.g., in an image stream) to ensure consistent and robust detection of that point. Essentially, temporal image information (e.g., changes in a series of images over time) is incorporated into the reference point selection method. For example, all images in a series of images (e.g., selected with gaze or variable intervals) or selected images can be examined to determine if the current landmark remains the best landmark to use as a reference anchor landmark. As different candidate landmarks become easier to trace (e.g., easier, faster, unique, and / or always detectable), they replace the previous reference point and become the new reference anchor point. All other landmark points can then be rereferenced relative to the new reference anchor point.

[0069] This embodiment is particularly useful for situations where the scanned (or imaged) area of ​​the eye (field of view) does not contain obvious or unique anatomical features (e.g., ONH) or where anatomical features are not necessarily used as reference points (e.g., the pupil changes during tracking due to its size / shape, such as over time). Therefore, in this embodiment, the uniqueness of the anatomical features selected as reference points is not required.

[0070] This embodiment can first detect reference (anchor) points from a set of candidate landmarks extracted from a reference image or a series of consecutive live images. For example, reference landmark candidates can be regions with good texture characteristics, such as the iris region facing the outer edge of the iris. An entropy filter can highlight regions with good texture characteristics, followed by additional image processing and analysis techniques to generate a mask containing landmark candidates to be selected as reference point candidates. Reference points located in regions with high contrast and texture can be selected as reference (anchor) points, which can be tracked in subsequent live images. Deep learning (e.g., neural network) methods / systems can be used to identify image regions with high contrast and good texture characteristics.

[0071] This embodiment uses templates centered on reference (anchor) points extracted from a reference image to track the live image. Additional templates extracted from the reference image are generated, centered on landmark candidates. These templates can be used to detect locations in the live image that correspond to a set of landmarks, which can be used for registration between the reference and live images, thereby tracking image sequences over time. Assuming a set of templates is extracted from the reference image, their corresponding locations in the live image (as a set of landmarks) can be determined in small regions of the live image from the reference point through template matching (e.g., normalized cross-correlation). Once all corresponding matches are found, a subset of the matches can be used to compute transformations (x and y translations and rotations) between the reference and live images.

[0072] The transformation determines the amount of motion between the live and reference images. The transformation can be calculated using two corresponding landmarks in the live and reference images. However, more than two landmarks can be used for tracking to ensure robustness.

[0073] A subset of matching landmarks can be determined through an exhaustive search. For example, in each iteration, two pairs of corresponding landmarks can be selected from the reference image and the present image. These two pairs can then be used to compute a rigid transformation. The error between each transformed reference image landmark (using the rigid transformation) and the present image landmark is determined. Landmarks associated with errors less than a predefined threshold can be selected as inliers. This process can be repeated for all (or most) feasible pairs (e.g., combinations). The transformation that creates the maximum number of inliers can then be chosen as the rigid transformation used for tracking.

[0074] Figure 8A , Figure 8B and Figure 9A , Figure 9B , Figure 9CThis illustrates the tracing of a set of landmarks in the reference points and the live imagery. The images in the left column are reference images, and the images in the right column are live images. The selected reference points are marked with a circled cross on each reference image. As can be seen, the selection of reference anchor points changes over time as different parts of the live image stream vary (e.g., shape or quality). Figure 8A , Figure 8B and Figure 9A , Figure 9B , Figure 9C The examples show variations in pupil size and shape, eyelid movement, and low-contrast effects.

[0075] In summary, motion artifacts pose a challenge to optical coherence tomography angiography (OCTA). While motion tracking solutions exist to correct these artifacts in retinal OCTA, motion tracking in the anterior segment (AS) of the eye remains unresolved. This is currently an obstacle to using AS-OCTA for diagnosing corneal, iris, and scleral diseases. This embodiment has been demonstrated for motion tracking in the anterior segment of the eye.

[0076] In a particular embodiment, a telecentric additional lens assembly with internal fixation is used to achieve CIRRUS with good patient alignment and fixation (fx). TM Frontal imaging was performed on the CIRRUS 6000 AngioPlex (ZEISS, Dublin, CA). A wide-field (20x14mm) line scan ophthalmoscopy (LSO) image set, totaling 6973 images (4798 central fixations and 2175 peripheral fixations), was acquired from 25 eyes of 15 subjects using additional lenses on the CIRRUS 6000. Motion within these image sets was then tracked using an algorithm employing rigid registration based on real-time landmarks between reference images and other (motion) images from the same set.

[0077] Figure 10 An example of a tracking algorithm according to an embodiment of the present invention is illustrated. In this example, anchor points and selected landmarks are found in a moving image and used to calculate translation and rotation values ​​for registration. Figure 10 The overlapping images at the bottom are used for visual verification. This embodiment first detects anchor points in regions of a reference image with high texture values. Then, the anchor point is located in the moving image by searching for a template (image region) centered on the anchor point location in the reference image. Next, the landmark from the reference image is found in the moving image by searching for a landmark template at the same distance from the anchor point in the reference image. Finally, translation and rotation are calculated using the landmark pair with the highest confidence value. The registration error is the average distance between corresponding landmarks in the two images. This value is calculated after visual confirmation of landmark matching and successful registration.

[0078] Figure 11The tracking test results are provided. The insets in the histograms of registration error and rotation angle illustrate their respective distribution parameters. The translation vector is plotted at the center with concentric rings every 500 μm. The inset shows the distribution parameters of the translation amplitude.

[0079] As mentioned above, real-time and efficient tracking of IR fundus images is crucial in automated retinal OCT image acquisition. Assuming the tracking and OCT acquisition fields of view (FOV) are located in the same retinal region, retinal tracking becomes more challenging when the patient's gaze is not straight or off-center. That is, the tracking and OCT (acquisition) FOVs are typically located in the same area of ​​the retina. Thus, motion tracking information can be used, for example, to correct OCT positioning during the OCT scan procedure. However, the location where the OCT scan is performed (e.g., the OCT acquisition FOV or OCT FOV) may not be on the retina with sufficient physical features / structures to robustly track motion. Current preferred methods identify / determine the optimal tracking FOV location before tracking and OCT acquisition.

[0080] Several challenges complicate efficient tracking. For example, IR images may not contain enough distributed retinal features (such as blood vessels) to track when the gaze deviates from the center. Another complexity is the greater curvature of the eye in the peripheral regions. Furthermore, eye movements introduce more nonlinear distortion in the current image relative to the reference image, which can lead to inaccurate tracking. Due to the nonlinear relationship between the two images, the transformation between the current image and the reference image may not be a rigid transformation.

[0081] One feasible solution is to place the tracking FOV at a location with well-defined retinal landmarks and features (e.g., around the ONH and large vessels) that can be robustly detected. One problem with this approach is that the large distance between the tracking FOV and the OCT acquisition FOV introduces rotation angle errors because the rotation anchor point is located within the tracking FOV rather than the OCT acquisition FOV.

[0082] To overcome the above challenges, the tracking FOV (e.g., in an IR image) can be made to at least partially overlap with the OCT acquisition FOV for eccentric fixation, or it can be placed as close as possible to the OCT acquisition FOV.

[0083] This paper proposes a method for optimal and dynamic localization of the tracking field of view (FOV) for patient gaze. In this method, a tracking algorithm (as described above, or other suitable tracking algorithms) is used to optimize the positional distribution of the tracking FOV by maximizing tracking performance using a set of metrics such as tracking error, landmarks (keypoints), and the number of landmarks. The tracking location (center of the FOV) that maximizes tracking performance is selected / assigned / identified as the desired location for a given patient gaze.

[0084] Essentially, this invention dynamically seeks the optimal tracking FOV (for a given patient gaze) that can effectively track OCT scans of eccentric gaze (e.g., gaze in the peripheral region of the eye). In this method, the tracking area can be placed at a location on the retina different from the OCT scan area. The optimal area on the retina relative to the OCT FOV is identified and used for retinal tracking. Therefore, a tracking FOV can be dynamically found for each eye. For example, the optimal tracking FOV can be determined based on tracking performance over a series of aligned images.

[0085] Figure 12 This illustrates the use of ONH to determine the location of the optimal tracking FOV (tracking window) for a given OCT FOV (acquisition / scanning window). In this example, the IR preview image (or a portion / window within the IR preview image) typically has a wide FOV (e.g., a 90-degree FOV) for patient alignment and can also be used to define the tracking FOV. These images, along with the IR tracking algorithm, can be used to determine the optimal tracking location relative to the OCT FOV. Dashed box 61 defines the OCT FOV, and dashed boxes 63A and 63B indicate the movement (repositioning) of non-optimal tracking FOVs until the optimal tracking FOV (solid black box 65) is identified. Non-optimal tracking FOVs 63A / 63B are moved toward the ONH at a distance from the center of the OCT FOV (e.g., indicated by the white solid line 67). Referencing the optimal tracking FOV (solid black box 65) in the IR preview image enables robust tracking of the remaining IR preview images within the same tracking FOV.

[0086] This document provides two embodiments (or implementations) of the present invention. The first embodiment uses a reference point. This implementation relies on a detectable reference point on the retina. For example, the reference point could be the center of the optic nerve head (ONH). The implementation can be summarized as follows:

[0087] 1) Collect a series of wide (e.g., 90-degree) FOV IR preview images (e.g. for patient alignment) or other suitable fundus images.

[0088] 2) Use / specify one of the collected images as a reference image.

[0089] 3) Detect the ONH center in the reference image.

[0090] 4) Crop the tracking FOV at the center of the OCT FOV in the reference image and use the cropped FOV as the tracking reference image (e.g., the current non-optimal tracking FOV 63A / 63B).

[0091] 5) Use the tracking reference image to track the remaining IR preview images in the set.

[0092] 6) Update the objective function. As is known in the art, the objective function in a mathematical optimization problem is a real-valued function whose value is minimized or maximized over a set of feasible alternatives. In this example, the objective function value is updated using, for example, the tracking outputs of all remaining IR preview images (e.g., tracking error, landmark distribution, number of landmarks, etc.).

[0093] 7) Update the tracking reference image by clipping the tracking FOV along a connecting line towards the ONH center (between the tracking and OCT FOV centers), e.g., line 67. An alternative to the connecting line could be a non-linear dynamic path from the tracking FOV center to the OCT FOV center. This non-linear dynamic path can be determined for each scan / eye.

[0094] 8) Repeat steps 5) to 8) until the objective function is minimized to achieve the maximum allowable distance between the OCT FOV and the optimal tracking FOV (e.g., constraint optimization).

[0095] Figure 13 A second embodiment of the invention is shown, used to determine the optimal tracking field of view (FOV) location without using OHN or other predefined physiological landmarks. Figure 12 All similar elements have similar reference numerals and are described above. The method searches for the optimal tracking FOV 65 around OCT FOV 61. The optimal location of the tracking FOV (solid black wireframe box) in the reference IR preview image enables robust tracking of the remaining IR preview image within the same tracking FOV. Figure 14A , Figure 14B , Figure 14C and Figure 14D Additional examples of this method for identifying the optimal tracking FOV relative to the OCT FOV are provided.

[0096] The second recommended solution / implementation does not use a reference point. This approach can be summarized as follows:

[0097] 1) Collect a series of wide FOV IR preview images (e.g., for patient alignment).

[0098] 2) Use one of the images as a reference image.

[0099] 3) Crop the tracking FOV at the center of the OCT FOV in the reference image, and use the cropped FOV as the tracking reference image.

[0100] 4) Remaining IR preview images in the tracking set relative to the tracking reference image.

[0101] 5) Update the objective function value using, for example, the tracking outputs of all remaining IR preview images (such as tracking error, landmark distribution, number of landmarks, etc.).

[0102] 6) Update the tracking reference image by cropping the tracking FOV to a region in the IR preview reference image that has rich anatomical features (e.g., blood vessels and lesions). Image saliency methods can be used to update the tracking FOV location.

[0103] 7) Repeat steps 4) to 7) until the objective function is minimized to the maximum allowable distance between the OCT and the tracking FOV (constraint optimization).

[0104] The retinal tracking method described above can also be used for automatic capture (e.g., OCT scans and / or fundus images). This will contrast with existing methods that use pupil tracking for automatic alignment and capture, but the inventors are unaware of any existing methods that use retinal tracking for alignment and automatic capture.

[0105] Automated patient alignment and image capture create a positive and efficient operator and patient experience. After initial operator calibration, the system can perform automatic tracking and OCT acquisition. Fundus images can be used to align with the scanned area on the retina. However, automated capture can be challenging because: eye movements during alignment; blinking and partial blinking; and alignment stability can lead to rapid instrument misalignment due to eye movements, focusing, operator errors, etc.

[0106] Retinal tracking algorithms can be used to lock onto fundus images and track incoming moving images. These algorithms require a reference image to calculate the geometric transformation between the reference and moving images. Tracking algorithms for automatic alignment and capture can be used in various scenarios.

[0107] For example, in the first scenario (Scenario 1), a reference image of the retina from a previous doctor's office visit is available. In this case, when the eye is stable (no movement or minimal movement), the reference image can be used for alignment and triggering automatic capture, and a series of retinal images can be robustly tracked. In the second scenario (Scenario 2), the retinal image quality algorithm may detect the reference image during initial alignment (by an operator or automatically). In this case, the reference image detected by the image quality algorithm can be used for alignment and triggering automatic capture in a similar manner to Scenario 1. The third scenario (Scenario 3) may be if a reference image from a previous visit and the retinal image quality algorithm are not available. In this case, the algorithm can track a sequence of images starting from the last image in the previous sequence as a reference image. The algorithm can repeat this process until a continuous and robust sequence of images is tracked, which can trigger automatic capture.

[0108] In this embodiment, an automatic alignment and automatic capture method is described for the three scenarios described above using a retinal tracking system. The basic idea is to use a tracking algorithm to evaluate the fundus image (e.g., to determine if the fundus image is a high-quality retinal image for a given gaze) and eye movement relative to the gaze position during alignment. Automatic capture is triggered if eye movement is minimal at the gaze position. Furthermore, the tracking output (e.g., xy translation and rotation relative to the gaze position) can also be used for automatic alignment in a motorized system by moving hardware components (e.g., chin rest or headrest, eyepiece, etc.).

[0109] As mentioned above, IR preview images (90-degree FOV) are typically used for patient alignment. These images, along with IR tracking algorithms such as those described above or other known IR tracking algorithms, can be used to determine whether the image can be continuously and robustly tracked using a reference image. Below are some examples applicable to the three different scenarios described above.

[0110] Figure 15 Scenario 1 is illustrated, where a retinal reference image from a patient's gaze during a previous visit is available. In this scenario, alignment and automatic capture are straightforward issues because the reference image is known for a given gaze. Figure 15 The diagram illustrates the use of a reference image to track each moving image. The dotted wireframe frame is an untrackable image, while the dashed-line frame is a trackable image. The quality of the tracking determines whether the image is in the correct gaze and has good quality. The quality of the tracking can be measured using tracking outputs such as tracking error, landmark distribution, number of landmarks, xy translation, and rotation of the moving image relative to the reference image (as described above). This tracking output can also be used for automatic alignment in motorized systems by moving hardware components such as chin rests or headrests, and eyepieces.

[0111] Automatic capture can be triggered if a predetermined number N of consecutive motion images are robustly tracked (e.g., with a predetermined confidence or quality metric). This indicates that the patient's eye movements are minimal and the gaze is correct. The tracking output can also be used to guide the operator or patient (graphically or using sound / language / text) to achieve better alignment.

[0112] Figure 16Scenario 2 is described, in which a retinal image quality algorithm detects a reference image during initial alignment (either manually or automatically). In this scenario, during alignment, a suitable IR image quality algorithm detects a reference image from a sequence of moving images. The operator performs initial alignment to bring the retina into the desired field of view and fixation. The IR image quality algorithm then determines the quality of a series of moving images. A reference image is then selected from a set of candidate reference images. The best reference image is selected based on its image quality score. Once the reference image is selected, automatic capture or automatic alignment can be triggered, as described above with reference to Scenario 1.

[0113] In scenario 3, the reference image and retinal image quality algorithm from previously accessed images are unavailable. In this scenario, the algorithm tracks a series of images starting from the last image in the previous sequence, serving as the reference image (solid white frame). Figure 17 The algorithm repeats this process until a continuous and robust sequence of images (dashed frame) is tracked, which can trigger automatic capture. In this method, the operator can perform initial alignment to bring the retina into the desired field of view and fixation.

[0114] The number of images in a sequence depends on tracking performance. For example, if tracking is not possible, a new sequence can be started with a new reference image from the last image in the previous sequence.

[0115] Figure 18 An alternative solution for Scenario 3 is described. This method selects a reference image from a continuous sequence of images that are tracked robustly and continuously. Once the reference image is selected, automatic capture or automatic alignment can be triggered, just like the method in Scenario 1.

[0116] The aforementioned imaging tracking application can be used to extract various statistical data to identify various characteristic problems affecting tracking. For example, the image sequence used for tracking can be analyzed to determine whether the image has characteristics of systematic motion, random motion, or good fixation. The ophthalmic system of the present invention can then be used to inform the system operator or patient of potential problems affecting tracking and provide suggested solutions.

[0117] Various types of artifacts in OCT can affect the diagnosis of ophthalmic diseases. Eccentricity artifacts and motion artifacts are considered important. Eccentricity artifacts are caused by fixation errors, resulting in displacement of the analytical grid on topographic maps for specific disease types. Eccentricity artifacts primarily occur in subjects with poor concentration, poor vision, or fixation deviation. Despite being instructed to fixate, involuntary eye movements still occur, with varying intensity in different directions during alignment and acquisition.

[0118] Motion artifacts are caused by eye saccades, changes in head position, or respiratory movements. Motion artifacts can be overcome with eye-tracking systems. However, eye-tracking systems often cannot handle saccades, inattentiveness, or poor vision. In these cases, the scan cannot be fully completed to the end.

[0119] To improve patient fixation and reduce distraction (especially for longer scan times), particularly for patients with poor attention or vision, eye movement analysis during alignment and acquisition can be a useful tool. It can inform both the operator and patient that more careful attention requires better fixation or eye movement control. For example, visual cues can be provided to the operator, and auditory cues to the patient, potentially leading to more successful scans. The operator can adjust hardware components based on the motion analysis output. The patient can be guided to the fixation target until the scan is complete.

[0120] This embodiment describes a method for eye-tracking analysis. The basic idea is to use retinal tracking output for real-time or post-acquisition analysis to generate a set of messages, including audio messages, that can inform the operator and patient of their gaze status and eye movements during alignment and acquisition. Providing motion analysis results post-acquisition can help the operator understand the reasons for poor scan quality so that the operator can take appropriate measures that may lead to a successful scan.

[0121] Eye-tracking analysis is helpful in the following aspects:

[0122] 1) Self-alignment: The patient can receive instructions from the device to align themselves.

[0123] 2) Automatic acquisition: During the acquisition process, the patient can be notified of changes in fixation or large movements. Fixation on the same location is important for small scanning fields of view.

[0124] 3) When the gaze target is unavailable, messages (e.g., in the form of sound) can keep the patient focused.

[0125] 4) Eye-tracking analysis results can be used in post-processing algorithms to address residual motion.

[0126] The aforementioned real-time retinal tracking method using infrared reflectance (IR) images for eccentric gaze was tested in a proof-of-concept application. As mentioned above, OCT acquisition systems rely on robust and real-time retinal tracking methods to capture reliable OCT images for visualization and further analysis. Tracking the retina via eccentric gaze can be challenging due to the lack of sufficiently rich anatomical features in the images. The proposed robust and real-time retinal tracking algorithm finds at least one anatomical feature with high contrast as a reference point (RP) to improve tracking performance.

[0127] In this example, as described above, the xy translation and rotation of the registration between the reference image and the moving image based on real-time keypoints (KPs) are calculated as tracking parameters. This tracking method relies on a unique RP and a set of reference image KPs extracted from the reference image. The location of the RP in the reference image is robustly detected using a fast image saliency method. Any suitable saliency method known in the art can be used. Examples of saliency methods can be found in: (1) X. Hou and L. Zhang, “Saliency Detection: A Spectral Residual Approach”, CVPR, 2007; (2) C. Guo, Q. Ma, and L. Zhang, “Spatio-temporal saliency detection using phase spectrum of quaternion fourier transform”, CVPR, 2008; and (3) B. Schauerte, B. Kühn, K. Kroschel, R. Stiefelhagen, “Multimodal Saliency-based Attention for Object-based Scene Analysis”, IROS, 2011.

[0128] Figure 19A Two examples are shown: small (top) and normal (bottom) pupil acquisition modes. RP (+) is detected using a fast image saliency algorithm. Keypoints (white circles) are detected relative to the RP location. RP locations in the moving image are detected by template matching using RP templates extracted from the reference image. Each reference KP template and its relative distance to the RP are used to search for a corresponding motion KP with the same distance to the RP location in the moving image, as indicated by the dashed arrows.

[0129] In this example implementation, the tracking parameters are calculated from a subset corresponding to KP with high confidence. Prototype software was used to collect IR image sequences (11.52 x 9.36 mm, pixel size 15 μm / pixel, and 50 Hz frame rate) from a CLARUS 500 (ZEISS, Dublin, CA). The registered images are displayed as single images to visualize the registration (e.g., Figure 19A (The rightmost image in the dataset). Calculate the average distance error between the registration motion and the reference KP for each moving image as the registration error.

[0130] The report included statistics on registration error, number of kPs, and eye movements. Figure 19BStatistics on registration error, eye movement, and number of keypoints are presented for a total of 29,529 images from 45 image sequences. Forty-five sequences were collected from one or both eyes of ten subjects / patients, with an average of 650 images per sequence. Patients' gazes were off-center. The average registration error of 15.3 ± 2.7 μm indicates that accurate tracking is feasible in the OCT domain with an A-scan spacing greater than 15 μm. The average execution time for tracking, measured using an Intel i7-8850H CPU, 2.6 GHz, and 32 GB RAM, was 15 ms. Therefore, this embodiment demonstrates the robustness of the proposed tracking algorithm, which is based on a real-time retinal tracking method using IR fundus images. This tool could be an essential component of any OCT image acquisition system.

[0131] Most eye-tracking-based analyses aim to identify and analyze an individual's visual attention patterns while performing specific tasks, such as reading, searching, scanning images, and driving. The anterior segment of the eye (e.g., the pupil and iris) is used for eye-tracking analysis. This method uses retinal tracking outputs (eye-tracking parameters) for each frame of a line-scan ophthalmoscopy (LSO) or infrared reflectance (IR) fundus image. Eye-tracking parameters (x, y, e.g., translation and rotation) recorded over a period of time can be used for statistical analysis, which may include statistical moment analysis of the eye-tracking parameters. Time-series analyses (e.g., Kalman filtering and particle filtering) can be used to predict future eye movements. This system can also use statistical and time-series analyses to generate messages to notify the operator and patient.

[0132] In this invention, eye-tracking analysis can be used during and / or after acquisition. Retinal tracking algorithms using LSO or IR images can be used to calculate eye-tracking parameters, such as x and y translation and rotation. The motion parameters are calculated relative to a reference image, which is captured through the initial gaze or using any of the methods described above. Figure 20A This describes the motion of the current image (white border) relative to the reference image (gray border), with eye-tracking parameters Δx, Δy, and rotation φ relative to the reference image. The current image is registered to the reference image, and then the two images are averaged.

[0133] For each fundus image, eye movement parameters are recorded, which can then be used for statistical analysis over a period of time. Examples of statistical analysis include statistical moment analysis of eye movement parameters. Time series analysis can be used for future eye movement prediction. Prediction algorithms include Kalman filtering and particle filtering. Statistical and time series analysis can be used to generate informative messages to notify operators and patients to take action.

[0134] In embodiments with motion analysis during alignment and acquisition, time-series analysis for eye movement prediction (next position and velocity) can alert the patient if he / she is deviating from the initial gaze position.

[0135] In embodiments where motion analysis is performed after acquisition, statistical analysis can be applied after the current acquisition has ended (regardless of whether the acquisition was successful or failed). An example of statistical analysis includes the total gaze deviation (mean of xy motion) from the initial gaze position and the distribution of eye movements (standard deviation) as a measure of the severity of eye movements during acquisition.

[0136] Figure 20B Examples from three different patients are shown: one with good fixation, another with systematic eye movements, and the third with random eye movements. Eye movement calculations can be applied to IR images relative to a reference image with an initial fixation.

[0137] Figure 21 A table is provided showing eye movement statistics for 15 patients. The mean represents the overall gaze deviation from the initial gaze position. The standard deviation represents a measure of eye movement during acquisition. Scans containing systematic or random eye movements showed significantly larger mean and standard deviations compared to scans with good gaze, which can be used as an indicator of poor gaze. For this study, significantly larger mean or standard deviations were defined as 116 and 90 micrometers, respectively. Eye and gaze point analysis can highlight its use as operator or patient feedback by providing informative messages to reduce motion during OCT image acquisition, which is important for any subsequent data processing.

[0138] The following provides a description of various hardware and architectures applicable to this invention.

[0139] Fundus imaging system

[0140] Two types of imaging systems used for fundus imaging are flood illumination imaging systems (or flood illumination imagers) and scanning illumination imaging systems (or scanning imagers). A flood illumination imager simultaneously floods the entire field of view (FOV) of interest of the sample, for example, by using a flash, and captures a full-frame image of the sample (e.g., the fundus) with a full-frame camera (e.g., a camera with a sufficiently large two-dimensional (2D) light sensor array to capture the desired FOV as a whole). For example, a flood illumination fundus imager illuminates the fundus of the eye with light and captures a full-frame image of the fundus in a single image capture sequence from the camera. A scanning imager provides a scanning beam that scans across an object (e.g., the eye), and as the scanning beam scans across the object, it images at different scanning locations, producing a series of image fragments that can be reconstructed, for example, a montage, to create a synthetic image of the desired FOV. The scanning beam can be a point, a line, or a two-dimensional region, such as a slit or a wide line. Examples of fundus imagers are provided in U.S. Patents 8,967,806 and 8,998,411.

[0141] Figure 22An example of a slit-scanning ophthalmic system SLO-1 for imaging the fundus F, which is the inner surface of the eye E opposite to the lens (or crystalline lens) CL, and may include the retina, optic disc, macula, fovea, and posterior pole. In this example, the imaging system is in a so-called “scan-to-de-scan” configuration, where the scan line beam SB scans the fundus F through the optical components of the eye E (including the corneal Crn, iris Irs, pupil Ppl, and lens CL). In the case of a floodlight fundus imager, a scanner is not required, and light is applied to the entire desired field of view (FOV) at a time. Other scanning configurations are known in the art, and a specific scanning configuration is not critical to the invention. As depicted, the imaging system includes one or more light sources LtSrc, preferably a multicolor LED system or a laser system, wherein the optical spread has been appropriately adjusted. An optional slit Slt (adjustable or static) is located in front of the light source LtSrc and can be used to adjust the width of the scan line beam SB. Furthermore, the slit Slt can remain stationary during imaging or can be adjusted to different widths to allow for different levels of confocality and different applications, whether for a specific scan or during a scan used to suppress reflections. An optional objective lens ObjL can be placed in front of the slit Slt. The objective lens ObjL can be any existing lens, including but not limited to refractive, diffractive, reflective, or hybrid lenses / systems. Light from the slit Slt passes through the pupil separator SM and is directed to the scanner LnScn. It is desirable to bring the scanning plane and the pupil plane as close as possible to reduce vignetting in the system. Optional optics DL can be included to control the optical distance between the images of the two components. The pupil separator SM can transmit the illumination beam from the light source LtSrc to the scanner LnScn and reflect the detection beam from the scanner LnScn (e.g., reflected light returning from the eye E) towards the camera Cmr. The task of the pupil separator SM is to separate the illumination and detection beams and help suppress system reflections. The scanner LnScn can be a rotating galvanometer scanner or other types of scanners (e.g., piezoelectric or voice coil, microelectromechanical systems (MEMS) scanners, electro-optic deflectors, and / or rotating polygon scanners). Depending on whether pupil separation is performed before or after the scanner LnScn, the scan can be divided into two steps, with one scanner in the illumination path and a separate scanner in the detection path. A specific pupil separation arrangement is described in detail in U.S. Patent No. 9,456,746, which is incorporated herein by reference in its entirety.

[0142] From the scanner LnScn, an illumination beam passes through one or more optics, in this case a scanning lens SL and an ophthalmic lens or eyepiece OL, which allows the pupil of the eye E to image onto the system's image pupil. Typically, the scanning lens SL receives the scanning illumination beam from the scanner LnScn at any of a plurality of scanning angles (incident angles) and produces a scanning line beam SB with a substantially flat surface focal plane (e.g., a collimated optical path). The ophthalmic lens OL can then focus the scanning line beam SB onto the object to be imaged. In this example, the ophthalmic lens OL focuses the scanning line beam SB onto the fundus F (or retina) of the eye E to image the fundus. In this way, the scanning line beam SB produces a transverse scanning line across the fundus F. One possible configuration of these optics is a Keplerian telescope, in which the distance between the two lenses is selected to produce an approximately telecentric intermediate fundus image (4-f configuration). The ophthalmic lens OL can be a single lens, an achromatic lens, or an arrangement of different lenses. As those skilled in the art will know, all lenses can be refractive, diffractive, reflective, or a combination of these. The focal lengths of the ophthalmic lens (OL), scanning lens (SL), and the size and / or form of the pupillary separator (SM) and scanner (LnScn) can vary depending on the desired field of view (FOV). Therefore, an arrangement can be envisioned in which multiple components can be switched in and out of the beam path according to the field of view, for example, by using flip-up optics, motorized wheels, or detachable optics. Since changes in the field of view result in different beam sizes across the pupil, changes in FOV can also be combined to alter pupillary separation. For example, a field of view of 45° to 60° is typical or standard for fundus cameras. Higher fields of view, such as 60°–120° or larger, wide field of view FOVs may also be feasible. Wide field of view FOVs may be required for combinations of wide-line fundus imaging (BLFI) with other imaging modalities, such as optical coherence tomography (OCT). The upper limit of the field of view can be determined by the achievable working distance combined with the physiological conditions surrounding the human eye. Since the typical human retina has a field of view (FOV) of 140° horizontally and 80° to 100° vertically, an asymmetrical field of view may be required to achieve the highest possible FOV on the system.

[0143] The scanning line beam SB passes through the pupil Ppl of the eye E and is directed towards the retina or fundus surface F. The scanner LnScn1 adjusts the position of the light on the retina or fundus F so that a lateral range of areas on the eye E is illuminated. Reflected or scattered light (or emitted light in the case of fluorescence imaging) is guided back along a similar path to the illumination to define the collection beam CB on the detection path to the camera Cmr.

[0144] In the "scan-de-scan" configuration of this exemplary slit-scan ophthalmic system SLO-1, the light returning from the eye E is "de-scanned" by the scanner LnScn on its way to the pupillator SM. That is, the scanner LnScn scans the illumination beam from the pupillator SM to define a scanning illumination beam SB passing through the eye E, but since the scanner LnScn also receives the returning light from the eye E at the same scanning position, the scanner LnScn has the effect of de-scanning the returning light (e.g., canceling the scanning action) to define a non-scanning (e.g., stable or stationary) collected beam from the scanner LnScn to the pupillator SM, which folds the collected beam toward the camera Cmr. At the pupillator SM, the reflected light (or the emitted light in the case of fluorescence imaging) is separated from the illumination light onto a detection path pointing toward the camera Cmr, which can be a digital camera with a light sensor for capturing an image. An imaging (e.g., objective) lens ImgL can be positioned in the detection path to image the fundus onto the camera Cmr. Similar to the case for the objective lens ObjL, the imaging lens ImgL can be any type of lens known in the art (e.g., refractive, diffractive, reflective, or hybrid lens). Additional operational details, particularly methods for reducing artifacts in images, are described in PCT Publication No. WO2016 / 124644, the contents of which are incorporated herein by reference in their entirety. The camera Cmr captures the received images, and for example, it creates an image file, which can be generated by one or more (electronic) processors or computing devices (e.g., ...). Figure 31 The computer system further processes the data. Therefore, the collected beam (returning from all scan positions of the scan line beam SB) is collected by the camera Cmr, and the full-frame image Img can be constructed from the synthesis of the individually captured collected beams, as by montage. However, other scanning configurations are also conceivable, including configurations where the illumination beam scans on the eye E and the collected beam scans on the camera's light sensor array. PCT Publication WO2012 / 059236 and U.S. Patent Publication No. 2015 / 0131050 (incorporated herein by reference) describe several embodiments of a slit scanning ophthalmoscope, including various designs in which the returned light sweeps across the camera's light sensor array and the returned light does not sweep across the camera's light sensor array.

[0145] In this example, the camera Cmr is connected to a processor (e.g., a processing module) Proc and a display (e.g., a display module, computer screen, electronic screen, etc.) Dspl. Both can be part of the imaging system itself, or they can be part of a separate, dedicated processing and / or display unit, such as a computer system, where data is transferred from the camera Cmr to the computer system via cable or a computer network including wireless networks. The display and processor can be a single, all-in-one unit. The display can be a conventional electronic display / screen or a touchscreen type and can include a user interface for displaying and receiving information from and from the instrument operator or user. The user can interact with the display using any type of user input device known in the art, including but not limited to a mouse, knob, button, pointer, and touchscreen.

[0146] During imaging, it may be necessary for the patient to maintain a fixed gaze. This is achieved by providing a fixation target that guides the patient's gaze. The fixation target can be inside or outside the instrument, depending on which area of ​​the eye is being imaged. Figure 22 An embodiment of an internal gaze target is illustrated. In addition to the primary light source LtSrc used for imaging, a second optional light source FxLtSrc, such as one or more LEDs, can be positioned to image a light pattern onto the retina using a lens FxL, a scanning element FxScn, and a reflector / mirror FxM. The gaze scanner FxScn can move the position of the light pattern, and the reflector FxM guides the light pattern from the gaze scanner FxScn to the fundus F of the eye E. Preferably, the gaze scanner FxScn is positioned so that it lies in the pupillary plane of the system, so that the light pattern on the retina / fundus can be moved according to the desired gaze position.

[0147] Slit-lamp ophthalmoscope systems can operate in different imaging modes depending on the light source and wavelength-selective filtering elements used. When imaging the eye using a series of colored LEDs (red, blue, and green), true-color reflectance imaging (similar to the imaging observed by clinicians when examining the eye with a handheld or slit-lamp ophthalmoscope) can be achieved. Images of each color can be progressively built up as each LED is turned on at each scanning position, or images of each color can be captured individually. The three-color images can be combined to display a true-color image or displayed individually to highlight different features of the retina. The red channel best highlights the choroid, the green channel highlights the retina, and the blue channel highlights the anterior segment of the retina. Furthermore, light of specific frequencies (e.g., a single colored LED or laser) can be used to excite different fluorophores in the eye (e.g., autofluorescence), and the generated fluorescence can be detected by filtering out the excitation wavelength.

[0148] Fundus imaging systems can also provide infrared reflective images, for example, by using an infrared laser (or other infrared light source). The advantage of infrared (IR) mode is that the eye is not sensitive to IR wavelengths. This allows users to continuously capture images without disturbing the eye (e.g., in preview / alignment mode) to assist the user during instrument alignment. Furthermore, IR wavelengths increase the ability to penetrate tissue and provide improved visualization of choroidal structures. Additionally, fluorescein angiography (FA) and indocyanine green (ICG) angiography can be performed by collecting images after a fluorescent dye has been injected into the subject's bloodstream. For example, in FA (and / or ICG), a series of time-lapse images can be captured after a photoreactive dye (e.g., a fluorescent dye) has been injected into the subject's bloodstream. It is important to note that caution must be exercised, as fluorescent dyes can cause life-threatening allergic reactions in some individuals. High-contrast, grayscale images are captured by exciting the dye using a selected specific light frequency. As the dye flows through the eye, different parts of the eye emit bright light (e.g., fluorescence), allowing the progression of the dye to be discerned, thus identifying blood flow through the eye.

[0149] Optical coherence tomography system

[0150] Typically, optical coherence tomography (OCT) uses low-coherence light to produce two-dimensional (2D) and three-dimensional (3D) internal views of biological tissues. OCT is capable of in vivo imaging of retinal structures. OCT angiography (OCTA) produces blood flow information, such as blood flow from vessels within the retina. Examples of OCT systems are provided in U.S. Patent Nos. 6,741,359 and 9,706,915, and examples of OCTA systems can be found in U.S. Patent Nos. 9,700,206 and 9,759,544, all of which are incorporated herein by reference in their entirety. Exemplary OCT / OCTA systems are provided herein.

[0151] Figure 23A general-purpose frequency-domain optical coherence tomography (FD-OCT) system for collecting 3D image data of an eye suitable for use with this invention is described. The FD-OCT system OCT_1 includes a light source LtSrc1. Typical light sources include, but are not limited to, broadband light sources with short-time coherence lengths or swept-frequency laser sources. The beam from the light source LtSrc1 is typically routed via fiber Fbr1 to illuminate a sample, such as an eye E; a typical sample is tissue within the human eye. The light source LrSrc1 can be, for example, a broadband light source with a short-time coherence length in the case of spectral-domain OCT (SD-OCT) or a wavelength-tunable laser source in the case of swept-frequency source OCT (SS-OCT). A scanning beam can be used, typically between the output of fiber Fbr1 and the sample E, such that the beam (dashed line Bm) scans laterally across the area of ​​the sample to be imaged. The beam from the scanner Scnr1 can be passed through a scanning lens SL and an ophthalmic lens OL and focused onto the sample E to be imaged. The scanning lens SL can receive the beam from the scanner Scnr1 at multiple incident angles and produce substantially collimated light, which the ophthalmic lens OL can then focus onto the sample. This example illustrates a scanning beam that needs to be scanned in two lateral directions (e.g., the x and y directions in the Cartesian plane) to scan the desired field of view (FOV). An example of this is a point-field OCT, which uses a point-field beam to scan the sample. Thus, the scanner Scnr1 is illustratively shown as comprising two sub-scanners: a first sub-scanner Xscn for scanning the point-field beam on the sample in a first direction (e.g., the horizontal x direction); and a second sub-scanner Yscn for scanning the point-field beam on the sample across a second direction (e.g., the vertical y direction). If the scanning beam is a line-field beam (e.g., a line-field OCT), which can sample the entire line portion of the sample at once, then perhaps only one scanner is needed to scan the line-field beam on the sample across the desired FOV. If the scanning beam is a full-field beam (e.g., a full-field OCT), then perhaps no scanner is needed, and the full-field beam can be applied to the entire desired FOV at once.

[0152] Regardless of the beam used, the light scattered from the sample (e.g., the sample light) is collected. In this example, the scattered light returning from the sample is collected into the same fiber Fbr1 used to guide the light for illumination. The reference light from the same light source LtSrc1 travels via a separate path, in this case involving fiber Fbr2 and a back reflector RR1 with adjustable optical delay. Those skilled in the art will recognize that a transmission reference path can also be used, and the adjustable delay can be placed in either the sample arm or the reference arm of the interferometer. The collected sample light is combined with the reference light, for example in a fiber coupler Cplr1, to form an optical interference in an OCT photodetector Dtctr1 (e.g., a photodetector array, a digital camera, etc.). Although a single fiber port is shown leading to detector Dtctr1, those skilled in the art will recognize that various designs of the interferometer can be used for balanced or unbalanced detection of the interference signal. The output of detector Dtctr1 is provided to a processor (e.g., an internal or external computing device) Cmp1, which converts the observed interference into depth information of the sample. Depth information can be stored in memory associated with processor Cmp1 and / or displayed on a display (e.g., computer / electronic display / screen) Scn1. Processing and storage functions can be located within the OCT instrument, or the functions can be offloaded (e.g., executed thereon) to an external processor (e.g., an external computing device), to which the collected data can be transferred. Figure 31 An example of a computing device (or computer system) is shown. This unit may be dedicated to data processing or to performing other common tasks not specific to the OCT device. The processor (computing device) Cmp1 may include, for example, a field-programmable gate array (FPGA), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a graphics processing unit (GPU), a system-on-a-chip (SoC), a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), or a combination thereof, which may perform some or all of the processing steps in a serial and / or parallel manner via one or more host processors and / or one or more external computing devices.

[0153] The sample and reference arms in an interferometer can consist of bulk optics, fiber optics, or hybrid bulk optics systems, and can have different architectures, such as Michelson, Mach-Zehnder, or common-path-based designs known to the art. The beams used herein should be interpreted as any carefully oriented optical path. Instead of a mechanical scanning beam, an optical field can illuminate a one-dimensional or two-dimensional region of the retina to generate OCT data (e.g., see U.S. Patent No. 9,332,902; D. Hillmann et al., “Holoscopy–Holographic Optical Coherence Tomography,” Optics Express, 36(13):2390 2011; Y. Nakamura et al., “High-Speed ​​Three Dimensional Human Retinal Imaging by Line Field Spectral Domain Optical Coherence Tomography,” Optics Express, 15(12):7103 2007; Blazkiewicz et al., “Signal-To-Noise Ratio Study of Full-Field Fourier-Domain Optical Coherence Tomography,” Applied Optics, 44(36):7722 (2005)). In time-domain systems, the reference arm needs to have an adjustable optical delay to generate interference. Balanced detection systems are commonly used in TD-OCT and SS-OCT systems, while spectrometers are used for the detection port of SD-OCT systems. The invention described herein can be applied to any type of OCT system. Various aspects of the invention can be applied to any type of OCT system or other types of ophthalmic diagnostic systems and / or multiple ophthalmic diagnostic systems, including but not limited to fundus imaging systems, field testing devices, and scanning laser polarimeters.

[0154] In Fourier domain optical coherence tomography (FD-OCT), each measurement is a real-valued spectral interferogram (Sj(k)). The real-valued spectral data typically undergoes several post-processing steps, including background subtraction and dispersion correction. The Fourier transform of the processed interferogram results in a complex-valued OCT signal output Aj(z) = |Aj|eiφ. The absolute value of this complex-valued OCT signal, |Aj|, ​​reveals the scattering intensity profile for different path lengths; therefore, scattering is a function of depth (z-direction) in the sample. Similarly, the phase φj can also be extracted from the complex-valued OCT signal. The scattering profile as a function of depth is called the axial scan (A-scan). A set of A-scans measured at adjacent locations in the sample produces a cross-sectional image of the sample (tomogram or B-scan). The collection of B-scans collected at different lateral locations on the sample constitutes a data volume or cube. For a given amount of data, the term fast axis refers to the scan direction along a single B-scan, while slow axis refers to the axis along which multiple B-scans are collected. The term "cluster scan" refers to a single data unit or block generated by repeated acquisitions at the same (or substantially the same) location (or region) for analyzing motion contrast, which can be used to identify blood flow. A cluster scan can consist of multiple A-scans or B-scans acquired at approximately the same location on the sample at relatively short time intervals. Because the scans in a cluster scan belong to the same region, the static structure remains relatively unchanged from scan to scan within the cluster scan, while motion contrast between scans that meet predefined criteria can be identified as blood flow.

[0155] Various methods for generating B-scans are known in the art, including, but not limited to: along the horizontal or x-direction, along the vertical or y-direction, along the diagonals of x and y, or in a circular or spiral pattern. A B-scan can be xz-dimensional, but can be any cross-sectional image including the z-dimensional dimension. Figure 24 This image shows a sample OCT B-scan image of a normal human retina. An OCT B-scan of the retina provides a view of the retinal tissue structure. For illustrative purposes, Figure 24 Various canonical retinal layers and their boundaries were identified. The identified retinal boundary layers include (from top to bottom): Internal Limiting Membrane (ILM) layer 1, Retinal Nerve Fiber Layer (RNFL or NFL) layer 2, Ganglion Cell Layer (GCL) layer 3, Internal Reticular Layer (IPL) layer 4, Inner Nuclear Layer (INL) layer 5, Outer Reticular Layer (OPL) layer 6, Outer Nuclear Layer (ONL) layer 7, Knot between the Outer Segment (OS) and Inner Segment (IS) of the light receptor (indicated by reference symbol layer 8), External or External Membrane (ELM or OLM) layer 9, Retinal Pigment Epithelium (RPE) layer 10, and Bruch's Membrane (BM) layer 11.

[0156] In OCT angiography or functional OCT, analytical algorithms can be applied to OCT data collected at the same or substantially the same sample location on the sample at different times (e.g., cluster scans) to analyze motion or flow (see, for example, the publications of U.S. Patent Nos. 2005 / 0171438, 2012 / 0307014, 2010 / 0027857, 2012 / 0277579 and U.S. Patent No. 6,549,801, all of which are incorporated herein by reference in their entirety). OCT systems can use any of a variety of OCT angiography processing algorithms (e.g., motion contrast algorithms) to identify blood flow. For example, motion contrast algorithms can be applied to intensity information derived from image data (intensity-based algorithms), phase information from image data (phase-based algorithms), or complex image data (complex-based algorithms). A frontal image is a 2D projection of 3D OCT data (e.g., by averaging the intensity of each individual A-scan so that each A-scan defines pixels in the 2D projection). Similarly, a frontal vascular system image is an image displaying motion contrast signals, where the data dimension corresponding to depth (e.g., along the z-direction of an A-scan) is displayed as a single representative value (e.g., a pixel image in a 2D projection), typically by summing or integrating all or isolated portions of the data (see, for example, U.S. Patent No. 7,301,644, the entire contents of which are incorporated herein by reference). An OCT system providing angiographic imaging capabilities may be referred to as an OCT angiography (OCTA) system.

[0157] Figure 25 An example of a frontal vascular system image is shown. After processing the data using any motion contrast technique known in the art to enhance motion contrast, pixel ranges corresponding to a given tissue depth from the surface of the internal limiting membrane (ILM) of the retina can be summed to generate a frontal (e.g., orthographic view) image of the vascular system. Figure 26 An exemplary B-scan of the vascular system (OCTA) is shown. As illustrated, structural information may not be clearly defined because blood flow may pass through multiple retinal layers, making them less distinct than in a structural OCT B-scan. Figure 24As shown. Nevertheless, OCTA provides a non-invasive technique for microvascular imaging of the retina and choroid, which can be crucial for the diagnosis and / or monitoring of various pathologies. For example, OCTA can be used to identify diabetic retinopathy by recognizing microaneurysms, neovascular complexes, and quantifying avascular and non-perfused areas of the fovea. Furthermore, OCTA has shown excellent concordance with fluorescein angiography (FA), a more routine but more covert technique that requires dye injection to observe vascular flow in the retina. Additionally, in dry age-related macular degeneration, OCTA has been used to monitor the pervasive reduction in choroidal capillary flow. Similarly, in wet age-related macular degeneration, OCTA can provide qualitative and quantitative analysis of the choroidal neovascular membrane. OCTA has also been used to investigate vascular occlusion, such as assessing unperfused areas and the integrity of superficial and deep neural plexuses.

[0158] Neural Networks

[0159] As described above, this invention can utilize neural network (NN) machine learning (ML) models. For completeness, a general discussion of neural networks is provided herein. This invention can use any of the following neural network architectures, individually or in combination. A neural network or neural network is a network of interconnected neurons (nodes), where each neuron represents a node in the network. Groups of neurons can be arranged hierarchically, with the output of one layer fed forward to the next layer in a multilayer perceptron (MLP) arrangement. An MLP can be understood as a feedforward neural network model that maps a set of input data to a set of output data.

[0160] Figure 27 An example of a multilayer perceptron (MLP) neural network is illustrated. Its structure may include multiple hidden (e.g., inner) layers HL1 to HLn, which map an input layer InL (receiving a set of inputs (or vector inputs) in_1 to in_3) to an output layer OutL, which produces a set of outputs (or vector outputs), such as out_1 and out_2. Each layer can have any given number of nodes, which are exemplarily shown as circles within each layer in this document. In this example, the first hidden layer HL1 has two nodes, while hidden layers HL2, HL3, and HLn each have three nodes. Generally, the deeper the MLP (e.g., the more hidden layers in the MLP), the greater its learning capacity. The input layer InL receives vector input (illustrated as a three-dimensional vector consisting of in_1, in_2, and in_3) and can apply the received vector input to the first hidden layer HL1 in the sequence of hidden layers. The output layer OutL receives the output from the last hidden layer (e.g., HLn) in the multi-layer model, processes its input, and produces a vector output (exemplarily shown as a two-dimensional vector consisting of out_1 and out_2).

[0161] Typically, each neuron (or node) produces a single output, which is fed forward to neurons in the immediately preceding layer. However, each neuron in a hidden layer can receive multiple inputs, either from the input layer or from the output of a neuron in the immediately preceding hidden layer. Typically, each node can apply a function to its inputs to generate an output for that node. Nodes in hidden layers (such as learning layers) can apply the same function to their respective inputs to produce their respective outputs. However, some nodes, such as those in the input layer InL, receive only one input and may be passive, meaning they simply relay the value of their single input to their output; for example, they provide a copy of their input to their output, as indicated by the dotted arrow within the node in the input layer InL.

[0162] For illustrative purposes, Figure 28 A simplified neural network consisting of an input layer InL', a hidden layer HL1', and an output layer OutL' is shown. The input layer InL' is shown as having two input nodes i1 and i2, which receive inputs Input_1 and Input_2 respectively (e.g., the input nodes of layer InL' receive a two-dimensional input vector). The input layer InL' feeds forward to the hidden layer HL1', which has two nodes h1 and h2, and this hidden layer in turn feeds forward to the output layer OutL', which has two nodes o1 and o2. The interconnections or links between neurons (shown as solid arrows) have weights w1 to w8. Typically, in addition to the input layers, nodes (neurons) can receive the output of the node immediately preceding them as input. Each node can compute its output by multiplying each of its inputs by the corresponding interconnection weights of each input, summing the products of its inputs, adding (or multiplying by) a constant defined by another weight or bias that may be associated with that particular node (e.g., node weights w9, w10, w11, w12 correspond to nodes h1, h2, o1, o2, respectively), and then applying a nonlinear or logarithmic function to the result. The nonlinear function can be called an activation function or a transfer function. Several activation functions are known in the art, and the choice of a particular activation function is not critical to this discussion. However, it is important to note that the operation of an ML model or the behavior of a neural network depends on the weight values, which can be learned so that the neural network provides the desired output for a given input.

[0163] During the training or learning phase, the neural network learns (e.g., is trained to determine) appropriate weight values ​​to achieve the desired output for a given input. Before training the neural network, an initial (e.g., random and optional non-zero) value, such as a random number seed, can be assigned individually to each weight. Various methods for assigning initial weights are known in the art. The weights are then trained (optimized) so that, for a given training vector input, the neural network produces an output close to the desired (predetermined) training vector output. For example, the weights can be progressively adjusted over thousands of iterations using a technique called backpropagation. In each iteration of backpropagation, the training input (e.g., a vector input or training input image / sample) is fed forward through the neural network to determine its actual output (e.g., a vector output). The error of each output neuron or output node is then calculated based on the actual neuron output and the target training output of that neuron (e.g., the training output image / sample corresponding to the current training input image / sample). The output is then backpropagated through the neural network (in the direction from the output layer back to the input layer), updating the weights based on the degree of influence of each weight on the overall error, thereby bringing the output of the neural network closer to the desired training output. This cycle is then repeated until the actual output of the neural network is within an acceptable error range of the desired training output for a given training input. It's understandable that each training input may require multiple backpropagation iterations to reach the desired error range. Typically, an epoch refers to one backpropagation iteration across all training samples (e.g., one forward propagation and one backpropagation), so training a neural network may require many epochs. Generally, the larger the training set, the better the performance of the trained ML model, so various data augmentation methods can be used to increase the size of the training set. For example, when the training set includes corresponding pairs of training input and training output images, the training images can be divided into multiple corresponding image segments (or patches). Corresponding patches from the training input and training output images can be paired to limit multiple training patch pairs from one input / output image pair, which expands the training set. However, training on a large training set places high demands on computational resources, such as memory and data processing resources. The computational requirements can be reduced by dividing the large training set into multiple mini-batches, where the mini-batch size limits the number of training samples in one forward / backward pass. In this case, one epoch can include multiple mini-batches. Another problem is that the NN may overfit the training set, thus reducing its ability to generalize from a specific input to different inputs. Overfitting can be mitigated by creating an ensemble of neural networks or by randomly dropping nodes from the neural network during training, which effectively removes the dropped nodes. Various dropout conditioning methods, such as inverse dropout, are known in the art.

[0164] Please note that the operations of a trained neural network (NN) are not direct algorithms for operational / analysis steps. In fact, when a trained NN receives input, that input is not analyzed in the conventional sense. Instead, regardless of the subject or nature of the input (e.g., a vector constrained to a live image / scan or a vector constrained to some other entity, such as a demographic description or activity record), the input will be subjected to the same pre-defined architecture of the trained neural network (e.g., the same node / layer arrangement, trained weights and biases, pre-defined convolution / deconvolution operations, activation functions, pooling operations, etc.), and it may be unclear how the architecture of the trained network produces its output. Furthermore, the values ​​of trained weights and biases are not deterministic and depend on many factors, such as the amount of time the neural network is used for training (e.g., the number of epochs in training), the random initial values ​​of the weights before training begins, the computer architecture of the machine training the NN, the selection of training samples, the distribution of training samples across multiple mini-batches, the choice of activation function, the choice of error function to correct the weights, and even if training is interrupted on one machine (e.g., with the first computer architecture) and completed on another machine (e.g., with different computer architectures). Crucially, the reasons why a trained ML model achieves certain outputs are not yet clear, and extensive research is underway to determine the factors upon which the outputs of ML models are based. Therefore, the processing of real-time data by neural networks cannot be simplified to simple step-by-step algorithms. Instead, its operation depends on its training architecture, the training sample set, the training sequence, and various conditions during the training of the ML model.

[0165] In summary, building a neural network (NN) machine learning model can include a learning (or training) phase and a classification (or operation) phase. During the learning phase, the neural network can be trained for a specific purpose and can be provided with a set of training examples, including training (sample) inputs and training (sample) outputs, and optionally a set of validation examples to test the progress of the training. During this learning process, various weights associated with the nodes and node interconnections in the neural network are incrementally adjusted to reduce the error between the actual output of the neural network and the desired training output. In this way, multilayer feedforward neural networks (as discussed above) can be made capable of approximating any measurable function to any desired accuracy. The result of the learning phase is a machine learning (ML) model that has been learned (e.g., trained). In the operation phase, a set of test inputs (or real-time inputs) can be submitted to the learned (trained) ML model, which can apply its learned knowledge to produce output predictions based on the test inputs.

[0166] and Figure 27 and Figure 28Like regular neural networks, convolutional neural networks (CNNs) consist of neurons with learnable weights and biases. Each neuron receives input, performs an operation (e.g., a dot product), and optionally undergoes non-linear processing. However, a CNN might receive raw image pixels at one end (e.g., the input) and provide a classification (or category) score at the other end (e.g., the output). Since CNNs expect images as input, they are optimized for volume (e.g., the pixel height and width of the image, and the image depth, e.g., color depth, such as RGB depth defined by the three colors red, green, and blue). For example, a CNN layer might be optimized for neurons arranged in 3D. Neurons in a CNN layer might also be connected to small regions of the previous layer, rather than all neurons in a fully connected CNN. The final output layer of a CNN can reduce the entire image to a single vector (classification) arranged along the depth dimension.

[0167] Figure 29 An example convolutional neural network architecture is provided. A convolutional neural network can be construed as a sequence of two or more layers (e.g., layer 1 to layer N), where each layer can include a (image) convolution step, a (result) weighted sum step, and a nonlinear function step. Convolution can be performed on the input data by applying filters (or kernels), for example, generating feature maps over a moving window of the input data. Each layer and its components can have different predefined filters (from a filter bank), weights (or weighting parameters), and / or function parameters. In this example, the input data is an image with a given pixel height and width, which can be the raw pixel values ​​of the image. In this example, the input image is shown as a depth image with three color channels RGB (red, green, and blue). Optionally, the input image can undergo various preprocessing steps, and the preprocessed results can be used to replace or supplement the original input image. Some examples of image preprocessing might include: retinal angiography segmentation, color space transformation, adaptive histogram equalization, connected component generation, etc. Within a layer, a dot product can be computed between a given weight and a small region connecting the weights in the input volume. Many ways to configure a CNN are known in the art, but as an example, layers can be configured to apply element-wise activation functions, such as a maximum (0, x) threshold at zero. Pooling functions (e.g., along the xy direction) can be performed to downsample the volume. Fully connected layers can be used to determine the classification output and produce a one-dimensional output vector that has been found useful for image recognition and classification. However, for image segmentation, a CNN needs to classify each pixel. Since each CNN layer tends to downsample the input image, another stage is needed to upsample the image back to its original resolution. This can be achieved by applying a transposed convolution (or deconvolution) stage TC, which typically does not use any predefined interpolation methods but instead has learnable parameters.

[0168] Convolutional neural networks have been successfully applied to many computer vision problems. As mentioned above, training a CNN typically requires a large training dataset. The U-Net architecture, based on CNNs, can usually be trained on a smaller training dataset than a regular CNN.

[0169] Figure 30 An example U-Net architecture is illustrated. This exemplary U-Net includes an input module (or input layer or stage) that receives an input U-in of any given size (e.g., an input image or image patch). For illustrative purposes, the image size of any stage or layer is indicated in a box representing the image; for example, the input module contains the number "128×128" to indicate that the input image U-in consists of 128×128 pixels. The input image can be a fundus image, an OCT / OCTA frontal image, a B-scan image, etc. However, it should be understood that the input can be of any size or dimension. For example, the input image can be an RGB color image, a monochrome image, a volumetric image, etc. The input image passes through a series of processing layers, each illustrated with exemplary sizes, but these sizes are for illustrative purposes only and will depend on, for example, the image size, the convolutional filters, and / or the pooling stages. This architecture includes a contraction path (illustrated here as consisting of four encoding modules), followed by an expansion path (illustrated here as consisting of four decoding modules), and copy and pruning links between the corresponding modules / stages (e.g., CC1 to CC4), where the output of an encoding module in the contraction path is copied and concatenated to the upconverted input of the corresponding decoding module in the expansion path (e.g., appended to its back side). This results in a typical U-shaped feature, hence the name of the architecture. Optionally, for computational reasons, a "bottleneck" module / stage (BN) can be positioned between the contraction and expansion paths. The bottleneck BN can consist of two convolutional layers (with batch normalization and optional dropout).

[0170] A shrinking path is analogous to an encoder, typically capturing contextual (or feature) information using feature maps. In this example, each encoding module in the shrinking path may include two or more convolutional layers, illustratively indicated by the asterisk “*”, and may be followed by a max-pooling layer (e.g., a downsampling layer). For example, the input image U-in is illustratively shown as undergoing two convolutional layers, each with 32 feature maps. It can be understood that each convolutional kernel produces one feature map (e.g., the output of a convolutional operation with a given kernel is an image commonly referred to as a “feature map”). For example, the input U-in undergoes a first convolution that applies 32 convolutional kernels (not shown) to produce an output consisting of 32 corresponding feature maps. However, as is known in the art, the number of feature maps produced by a convolutional operation can be adjusted (up or down). For example, the number of feature maps can be reduced by averaging feature map groups, discarding some feature maps, or other known feature map reduction methods. In this example, the first convolution is followed by a second convolution, the output of which is constrained to 32 feature maps. Another approach to conceiving of feature maps might be to view the output of the convolutional layer as a 3D image, where the 2D dimension of the 3D image is given by the listed XY plane pixel dimensions (e.g., 128 × 128 pixels), and its depth is given by the number of feature maps (e.g., 32 planar image depths). Following this analogy, the output of the second convolution (e.g., the output of the first encoding module in the shrinking path) could be described as a 128 × 128 × 32 image. The output of the second convolution is then pooled, which reduces the 2D dimension of each feature map (e.g., the X and Y dimensions can each be halved). The pooling operation can be embodied in the downsampling operation, as indicated by the down arrow. Several pooling methods, such as max pooling, are known in the art, and the specific pooling method is not critical to this invention. The number of feature maps may double with each pooling operation, starting with 32 feature maps in the first encoding module (or block), 64 in the second encoding module, and so on. Thus, the shrinking path forms a convolutional network consisting of multiple encoding modules (or stages or blocks). As a typical convolutional network, each encoding module provides at least one convolutional stage, followed by an activation function (e.g., a rectified linear unit (ReLU) or a sigmoid layer), not shown, and a max-pooling operation. Typically, the activation function introduces non-linearity into the layer (e.g., to help avoid overfitting), receives the layer's results, and determines whether to "activate" the output (e.g., to determine if the value of a given node meets a predefined criterion for forwarding the output to the next layer / node). In summary, shrinking paths generally reduce spatial information while increasing feature information.

[0171] The expansion path, similar to a decoder, can provide localization and spatial information to the results of the contraction path, among other things, although downsampling and any max pooling are performed during the contraction phase. The expansion path comprises multiple decoding modules, each concatenating its current upconverted input with the output of the corresponding encoding module. In this way, features and spatial information are combined in the expansion path through a series of upconvolutions (e.g., upsampling, transposed convolutions, or deconvolutions) and concatenations with high-resolution features from the contraction path (e.g., via CC1 to CC4). Therefore, the output of the deconvolutional layers is concatenated with the corresponding (optionally cropped) feature maps from the contraction path, followed by two convolutional layers and an activation function (optionally batch normalization).

[0172] The output of the last expansion block in the expansion path can be fed into another processing / training block or layer, such as a classifier block, which can be trained together with the U-Net architecture. Alternatively or additionally, the output of the last upsampling block (at the end of the expansion path) can be submitted to another convolutional operation (e.g., output convolution) before producing its output U-out, as indicated by the dotted arrow. The kernel size of the output convolution can be chosen to reduce the dimension of the last upsampling block to the desired size. For example, a neural network may have multiple features per pixel before reaching the output convolution, which could provide a 1×1 convolutional operation to combine these multiple features at the pixel-by-pixel level into a single output value for each pixel.

[0173] Computing devices / systems

[0174] Figure 31 Example computer systems (or computing devices or computer apparatuses) are illustrated. In some embodiments, one or more computer systems may provide the functionality described or illustrated herein and / or perform one or more steps of one or more methods described or illustrated herein. Computer systems may take any suitable physical form. For example, a computer system may be an embedded computer system, a system-on-a-chip (SOC), a single-board computer system (SBC) (e.g., a computer module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, a computer system grid, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the computer system may reside in a cloud, which may include one or more cloud components in one or more networks.

[0175] In some embodiments, the computer system may include a processor Cpnt1, a memory Cpnt2, a storage device Cpnt3, an input / output (I / O) interface Cpnt4, a communication interface Cpnt5, and a bus Cpnt6. The computer system may also optionally include a display Cpnt7, such as a computer monitor or screen.

[0176] Processor Cpnt1 includes hardware for executing instructions, such as the hardware that constitutes a computer program. For example, processor Cpnt1 may be a central processing unit (CPU) or a general-purpose graphics processing unit (GPGPU). Processor Cpnt1 may retrieve (or fetch) instructions from internal registers, internal caches, memory Cpnt2, or storage device Cpnt3, decode and execute instructions, and write one or more results to internal registers, internal caches, memory Cpnt2, or storage device Cpnt3. In a particular embodiment, processor Cpnt1 may include one or more internal caches for data, instructions, or addresses. Processor Cpnt1 may include one or more instruction caches and one or more data caches, such as for storing data tables. Instructions in the instruction cache may be copies of instructions in memory Cpnt2 or storage device Cpnt3, and the instruction cache may accelerate the retrieval of these instructions by processor Cpnt1. Processor Cpnt1 may include any suitable number of internal registers and may include one or more arithmetic logic units (ALUs). Processor Cpnt1 may be a multi-core processor; or may include one or more processors Cpnt1. Although this disclosure describes and illustrates a particular processor, this disclosure considers any suitable processor.

[0177] Memory Cpnt2 may include main memory for storing instructions for processor Cpnt1 to execute during processing or to save temporary data. For example, a computer system may load instructions or data (e.g., a data table) from storage device Cpnt3 or from another source (e.g., another computer system) into memory Cpnt2. Processor Cpnt1 may load instructions and data from memory Cpnt2 into one or more internal registers or internal caches. To execute instructions, processor Cpnt1 may retrieve and decode instructions from internal registers or internal caches. During or after instruction execution, processor Cpnt1 may write one or more results (which may be intermediate or final results) to internal registers, internal caches, memory Cpnt2, or storage device Cpnt3. Bus Cpnt6 may include one or more memory buses (each of which may include an address bus and a data bus) and may couple processor Cpnt1 to memory Cpnt2 and / or storage device Cpnt3. Optionally, one or more memory management units (MMUs) facilitate data transfer between processor Cpnt1 and memory Cpnt2. The memory Cpnt2 (which may be fast volatile memory) may include random access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM). The storage device Cpnt3 may include long-term or high-capacity storage for data or instructions. The storage device Cpnt3 may be internal or external to the computer system and includes one or more of the following: disk drives (e.g., hard disk drives, HDDs or solid-state drives, SSDs), flash memory, ROM, EPROM, optical disks, magneto-optical disks, magnetic tape, Universal Serial Bus (USB) accessible drives, and other types of non-volatile memory.

[0178] The I / O interface Cpnt4 can be software, hardware, or a combination of both, and includes one or more interfaces (e.g., serial or parallel communication ports) for communicating with I / O devices, enabling communication with people (e.g., users). For example, I / O devices may include keyboards, keypads, microphones, monitors, mice, printers, scanners, speakers, cameras, styluses, tablets, touchscreens, trackballs, video cameras, other suitable I / O devices, or combinations thereof.

[0179] The communication interface Cpnt5 provides a network interface for communicating with other systems or networks. The communication interface Cpnt5 may include a Bluetooth interface or other types of packet-based communication. For example, the communication interface Cpnt5 may include a network interface controller (NIC) and / or a wireless NIC or a wireless adapter for communicating with a wireless network. The communication interface Cpnt5 can provide communication with Wi-Fi networks, ad hoc networks, personal area networks (PANs), wireless PANs (e.g., Bluetooth WPANs), local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), cellular telephone networks (e.g., Global System for Mobile Communications (GSM) networks), the Internet, or a combination of two or more of these.

[0180] The Cpnt6 bus can provide communication links between the aforementioned components of a computing system. For example, the Cpnt6 bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand bus, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Fast (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or a combination of two or more of these.

[0181] Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0182] Herein, one or more computer-readable non-transitory storage media may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. Where appropriate, computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.

[0183] Although the invention has been described in conjunction with several specific embodiments, it will be apparent to those skilled in the art that many further alternatives, modifications, and variations will be apparent from the foregoing description. Therefore, the invention described herein is intended to include all such alternatives, modifications, applications, and variations that may fall within the spirit and scope of the appended claims.

Claims

1. An eye-tracking method, comprising: Capture multiple images of the eye, including a reference image and one or more live images; A reference anchor point is defined in the reference image; One or more auxiliary points are defined in the reference image; Obtain the distance and direction between the reference anchor point and the one or more auxiliary points in the reference image; In the selected live image: a) Identify the initial matching point that matches the reference anchor point; b) Based on the distance and direction of the selected auxiliary point relative to the matched reference anchor point, search for matching points of the selected auxiliary point within the region; Based on the matching points of the initial matching point and the selected auxiliary point, the tracking error between the reference image and the selected live image is corrected.

2. The method according to claim 1, wherein: The reference image defines a plurality of the auxiliary points; Searching for matching points of the selected auxiliary points is part of searching for matching auxiliary points in the selected live image; In response to the fact that the number of auxiliary points matched in the selected live image is not greater than a predetermined minimum value, no tracking error correction is performed on the selected live image.

3. The method according to claim 2, wherein, The pre-defined minimum value is greater than half of the multiple auxiliary points.

4. The method according to any one of claims 1 to 3, wherein, The reference image and the live image are infrared images.

5. The method according to any one of claims 1 to 3, wherein, The captured images are images of the retina of the eye.

6. The method according to any one of claims 1 to 3, further comprising: Identify significant physical features in the reference image; The reference anchor point is defined based on the significant physical features.

7. The method according to claim 6, wherein: The reference anchor point is part of a reference anchor template, which consists of multiple identifiers that collectively define the significant physical feature. as well as Identifying the initial matching point is part of identifying the initial matching template that matches the reference anchor template.

8. The method according to claim 1, wherein: Identify significant physical features in the reference image; The reference anchor point is defined based on the significant physical features; The reference anchor point is part of a reference anchor template, which consists of multiple identifiers that collectively define the significant physical feature; and Identifying the initial matching point is part of identifying an initial matching template that matches the reference anchor template; The one or more auxiliary points are part of a corresponding one or more auxiliary templates, the auxiliary templates being composed of multiple identifiers that collectively define the corresponding auxiliary physical features in the reference image; and The search for matching points of the selected auxiliary point is based on the offset position of the selected auxiliary template relative to the reference anchor template, searching within the region for a portion of the matching corresponding to the selected auxiliary template.

9. The method according to claim 8, wherein, The significant physical features in the reference image and the one or more live images are identified using a neural network.

10. The method according to claim 9, wherein, The significant physical feature is the predefined retinal structure.

11. The method according to claim 10, wherein, The prominent physical features are the optic disc, optic nerve head, lesions, or specific vascular patterns.

12. The method according to claim 6, wherein, The prominent physical features are the optic nerve head, pupil, iris boundary, or center of the eye.

13. The method according to any one of claims 1 to 3, further comprising: Identify multiple candidate anchor points in the reference image; Search for matching points among the multiple candidate anchor points in the selected live image; The best-matching candidate anchor point in the selected live image is designated as the reference anchor point; The candidate anchor point for the best match is either the candidate with the highest confidence in the selected live image or the candidate that finds a match the fastest.

14. The method of claim 13, further comprising: Search for matching points for the multiple candidate anchor points in the multiple live images; The best-matching candidate anchor point among multiple selected live images is designated as the reference anchor point.

15. The method according to claim 14, wherein, The best-matching candidate anchor point is the candidate anchor point that has the highest number of successful matches in a series of consecutively selected live images.

16. The method according to claim 15, wherein, The series consists of a predetermined number of live images.

17. The method according to claim 13, wherein, The candidate anchor points are identified based on their salience in the reference image.

18. The method according to claim 13, wherein, Candidate anchor points that were not designated as the reference anchor point were designated as the auxiliary points.

19. The method of claim 1, further comprising: A plurality of the reference anchor points are defined in the reference image; In the selected live images: a) Identify multiple initial matching points that match multiple reference anchor points, and transform the live image into the reference image based on the identified multiple reference anchor points as coarse registration; b) Based on the position of the selected auxiliary point relative to the plurality of reference anchor points, search for matching points of the selected auxiliary point within the region.

20. The method of claim 1, further comprising: Use an OCT system to define the OCT acquisition field of view on the eye; in: The reference anchor point is confined within a tracking field of view that can move within the reference image; The tracking field of view moves around the reference image to the optimal position determined as the tracking image, while also at least partially overlapping with the OCT acquisition field of view.

21. The method according to claim 20, wherein, The optimal position is based on the output of a tracking algorithm, which includes one or more of the following: tracking error, landmark distribution, and number of landmarks.

22. An eye-tracking method, comprising: Capture multiple images of the retina of the eye, the multiple images including a reference image and one or more live images; Identify significant physical features in the reference image; The reference anchor template is defined based on the aforementioned significant physical characteristics; One or more auxiliary templates are defined based on other physical features in the reference image; Obtain the distance and orientation between the reference anchor template and the one or more auxiliary templates in the reference image; Store the position of the auxiliary template relative to the reference anchor template; In each live image: a) Identify an initial matching region that matches the reference anchor template, the initial matching region defining a corresponding real anchor template, the position of the real anchor template matching the position of the reference anchor template; b) Search for matching regions of the one or more auxiliary templates, each found matching region defining another corresponding template in the live image, wherein the search for each auxiliary template is limited to a bound region, the position of the bound region relative to the live anchor template being based on the distance and direction of the auxiliary template relative to the matching reference anchor template; as well as Based on two or more pairs of corresponding templates of the reference image and the selected live image, the tracking error between the reference image and the selected live image is corrected.

23. The method according to claim 22, wherein, The significant physical features in the reference image and each live image are identified using a neural network.

Citation Information

Patent Citations

  • High speed spectral domain functional optical coherence tomography and optical doppler tomography for in vivo blood flow dynamics and tissue structure

    US20050171438A1

  • Method and apparatus for ultrahigh sensitive optical microangiography

    US20120307014A1

  • Systems and methods for broad line fundus imaging

    US20150131050A1

  • Phase-resolved optical coherence tomography and optical doppler tomography for imaging fluid flow in tissue with fast scanning speed and high velocity sensitivity

    US6549801B1

  • Optical coherence tomography optical scanner

    US6741359B2