Real-time ir fundus image tracking in the presence of artifacts using reference landmarks

By using a reference point for landmark matching in IR fundus images, the system addresses the challenge of tracking ocular motion in the presence of artifacts, achieving robust and efficient eye movement tracking.

JP2026034463APending Publication Date: 2026-02-27CARL ZEISS MEDITEC INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025202748
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-04-29
Filing Date
2025-11-25
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Current eye tracking systems face challenges in accurately tracking ocular motion due to the presence of artifacts in infrared (IR) fundus images, which affects real-time processing and reliability, especially when high-resolution images are required.

Method used

The system identifies a reference point or template in the IR image, such as the optic nerve head (ONH), and matches additional landmarks relative to this point, maintaining a constant distance for robust landmark detection, allowing real-time tracking without advanced image processing.

Benefits of technology

This method enables efficient, real-time eye movement tracking even with high-resolution images, reducing the impact of image artifacts and ensuring accurate tracking performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026034463000001_ABST
    Figure 2026034463000001_ABST
Patent Text Reader

Abstract

To provide a more efficient system / method for eye motion tracking.SOLUTION: A system and method for eye motion tracking. An anchor point and a plurality of auxiliary points are selected from the reference image. Individual live images in the sequence of images are then searched for a match between the anchor point and the ancillary point. First, an anchor point is found, and then the search for individual auxiliary points is limited to a search window defined by the known distance and / or orientation of the auxiliary point to be searched relative to the anchor point.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to motion tracking, and more particularly to tracking ocular motion of the anterior and posterior segments of the eye. [Background technology]

[0002] Fundus imaging, such as can be obtained by using a fundus camera, generally provides a frontal, planar view of the fundus as seen through the pupil of the eye. Fundus imaging may use different frequencies of light, such as white light, red light, blue light, green light, infrared (IR), etc., to image tissue, or may use selected frequencies to excite fluorescent molecules within specific tissues (e.g., autofluorescence) or to excite fluorescent dyes injected into the patient (e.g., fluorescence angiography). A more detailed description of different fundus imaging techniques is provided below.

[0003] OCT is a noninvasive imaging technique that uses light waves to generate cross-sectional images of retinal tissue. For example, OCT allows for the visualization of distinct tissue layers of the retina. Generally, OCT systems are interferometric imaging systems that determine the scattering profile of a sample along the OCT beam by detecting the interference of light reflected from the sample with a reference beam, forming a three-dimensional (3D) representation of the sample. Each scattering profile in the depth direction (e.g., z-axis or axial direction) can be individually reconstructed into an axial scan or A-scan. Cross-sectional two-dimensional (2D) images (B-scans) and extended 3D volumes (C-scans or cube scans) can be constructed from multiple A-scans acquired as the OCT beam is scanned / translated through a set of transverse (e.g., x-axis and y-axis) locations on the sample. OCT also allows for the construction of planar, en face (e.g., en face) 2D images of selected portions of a tissue volume (e.g., a target tissue slab (subvolume) or target tissue layer(s) of the retina). OCTA is an extension of OCT and can identify (e.g., render in an image format) the presence or absence of blood flow in tissue layers. OCTA can identify blood flow by identifying differences (e.g., contrast differences) over time in multiple OCT images of the same retinal region and designating differences that meet predetermined criteria as blood flow. A more detailed description of OCT and OCTA is provided below.

[0004] Real-time and efficient tracking of fundus images (e.g., infrared (IR) fundus images) is important in automated retinal OCT image acquisition. Retinal tracking is particularly important due to involuntary eye movements during image acquisition, especially between OCT and OCTA scans.

[0005] IR images can be used to track retinal movement. However, poor IR image quality and the presence of various artifacts can affect automated, real-time processing, thereby reducing success rates and reliability. IR image quality can change significantly over time depending on fixation, focus, vignetting, eyelash, stripe, and central reflex artifacts. Therefore, a method is needed that can robustly track the retina using IR images in real time. Figure 1 provides an example IR fundus image with various artifacts, including stripe artifacts 11, central reflex artifacts 13, and eyelashes 15 (e.g., seen as dark shadows).

[0006] Current tracking systems use a reference image with a set of landmarks extracted from the image. Then, the tracking algorithm tracks the live image by searching for landmarks in each live image using the landmarks extracted from the reference image. Landmark matching between the reference image and the live image is determined independently. Therefore, the presence of artifacts in the image, such as stripe and center reflection artifacts, makes matching a challenging problem. Typically, advanced image processing algorithms are required to enhance the image before landmark detection. Using these additional algorithms impairs the real-time performance of the tracking algorithm, especially when the tracking algorithm needs to run on high-resolution images for more accurate tracking.

[0007] In summary, prior art tracking systems use a reference fundus image with a set of landmarks extracted from the reference image. Then, a tracking algorithm uses the landmarks extracted from the reference image to track a series of live images by independently searching for each landmark in each live image. Landmark matches between the reference image and the live image are determined independently. Therefore, matching landmarks is a challenging problem due to the presence of artifacts in the images (such as stripe and central reflection artifacts). Advanced image processing algorithms are required to enhance the IR image before landmark detection. The addition of these advanced algorithms can hinder their use in real-time applications, especially when the tracking algorithm is required to run on high-resolution (e.g., large) images for more accurate tracking. Summary of the Invention [Problem to be solved by the invention]

[0008] It is an object of the present invention to provide a more efficient system / method for eye movement tracking. Another object of the present invention is to provide real-time eye movement tracking using high resolution images. [Means for solving the problem]

[0009] The above objectives are achieved in a method / system for eye tracking. Unlike prior art tracking systems, the present system does not independently search for matching landmarks. Rather, the present invention identifies a reference (anchor) point / template (e.g., a landmark) and matches additional landmarks relative to the location of the reference point. Multiple landmarks can then be detected in a live IR image relative to the reference (anchor) point or template obtained from the reference image.

[0010] A fiducial point may be selected that is a salient anatomical / physical feature that is easily identified and that can be predicted to be present in subsequent live images (e.g., when imaging the posterior segment of the eye, the optic nerve head (ONH), a lesion, or a particular vascular pattern). Alternatively, for example, when no salient, unchanging anatomical feature exists (such as when imaging the anterior segment), a fiducial anchor point may be selected from a pool / group of candidate fiducial points based on the current state of the series of images. As the quality of the series of live images changes or different salient features become apparent, the fiducial anchor point is modified / changed accordingly. Thus, in this alternative embodiment, the fiducial anchor point may change over time depending on the images being captured / collected.

[0011] It should be understood that the reference point or template may include one or more distinctive features (pixel identifiers) that together define (e.g., identify) a particular landmark (e.g., ONH, a lesion, or a particular vascular pattern) used as a reference physical landmark. The distance between the reference point and the selected landmark in the reference IR image and the live IR image is maintained constant (or their relative distance is maintained constant) in both images. Thus, landmark detection in the live IR image becomes a simpler problem by searching a small region (e.g., a border region or a predefined / fixed-size window) at a predetermined / specific distance from the reference point. Due to the constant distance between the reference point and the landmark location, the robustness of landmark detection is improved / facilitated. Once an initial landmark is matched, the search for additional landmarks may be further limited to a specific direction / orientation / angle (e.g., in addition to a specific distance) relative to the already matched landmark.

[0012] The speed of our method can be improved, especially when dealing with high-resolution images, because no advanced image processing algorithms are required to enhance the IR image before landmark detection, ensuring the real-time performance of the tracking algorithm on high-resolution images for more accurate tracking.

[0013] In summary, the present invention may begin by detecting salient points (e.g., ONH locations or other points, such as lesions) as fiducial / anchor points in a selected reference IR image (or other imaging modality) using, for example, deep learning or knowledge-based computer vision methods. To increase the number of templates and corresponding landmarks in the IR image, additional templates / points offset from the fiducial point center are extracted from the reference IR image. Optionally, multiple fiducial / anchor points can be used for general purposes. For example, multiple images of the eye, including a reference image and one or more live images, may be captured. Multiple fiducial anchor points are then defined in the reference image, and one or more auxiliary points may also be defined in the reference image. Multiple initial matching points that match (all or part of) the multiple fiducial anchor points are then identified in the selected live image, and the selected live image is transformed into the reference image as a coarse registration based on (e.g., using) the identified multiple fiducial anchor points. After this coarse alignment, the selected auxiliary point can be searched for matches within an area (e.g., bounded by a search area / FOV / window) based on the position of the selected auxiliary point relative to multiple matched reference anchor points. Tracking errors between the reference image and the selected live image can then be corrected based on these matched points. This technique can be useful when there is significant geometric transformation between the reference image and the live image during tracking, and for more complex tracking systems. For example, if there is a large rotation (or affine / projective relationship) between the two images, two or more anchor points can be used to first coarsely align the two images to more accurately search for additional landmarks. The tracking algorithm then tracks the live IR image (or other corresponding imaging modality) using a template centered on the reference point and additional templates extracted from the reference IR image.Given a set of templates extracted from the reference IR image, their corresponding locations (as a set of landmarks) in the live IR image can be determined by template matching in a small region away from the reference points in the live IR image. All or some of the matches can be used to calculate the transformation (x and y shifts and rotations) between the IR reference image and the live IR image. In this way, landmarks (e.g., matching templates or matching points) are detected relative to the reference landmarks. This leads to real-time operation (e.g., processing is limited to a small region of the image) and robust tracking (e.g., the distance between the reference landmarks and the landmarks is known, eliminating false positives and providing an additional check to verify the validity of match candidates).

[0014] The advantage of this method compared to the prior art is that landmarks are searched for (and tracked dependently) and detected in specific regions / windows in a live image whose positions and / or sizes are defined relative to a reference landmark. Therefore, real-time tracking can be performed with minimal pre-processing of high-resolution IR images, and no advanced image processing techniques are required during tracking. Furthermore, the present invention is less sensitive to the presence of image artifacts (such as stripe artifacts, central reflection artifacts, and eyelashes) due to the fact that tracking is defined relative to a reference point.

[0015] Furthermore, the present invention may be extended to move a fundus image (e.g., IR image) tracking region / window with a given FOV that at least partially overlaps with the OCT scan FOV. To maintain robust tracking, the IR FOV may be moved (while maintaining overlap with the OCT FOV) until a position is reached that includes a maximum (or sufficient) number of easily / robustly identifiable anatomical features (e.g., ONH, lesions, or specific vascular patterns). This tracking information can then be used to provide motion compensation to the OCT system with respect to the OCT scan.

[0016] Thus, the reference image can be used to align and trigger automatic capture (e.g., from an OCT system) when the eye is stable (no or minimal movement), and a series of retinal images are tracked robustly.

[0017] The present invention provides various metrics for quantifying the quality of tracking images and possible causes of poor tracking images. Accordingly, the present invention can extract various statistics from historical data to identify various characteristic problems affecting tracking. For example, the present invention can analyze a series of images used for tracking to determine whether the images have characteristics of generalized motion, random motion, or good fixation. An ophthalmic system using the present invention can then inform the system operator or the patient of problems that may affect tracking and provide suggested solutions.

[0018] Other objects and achievements of the present invention, together with a fuller understanding of the invention, will become apparent and understood by reference to the following description and claims taken in conjunction with the accompanying drawings.

[0019] To facilitate the understanding of the present invention, several publications are cited or referenced herein. All publications cited or referenced herein are incorporated by reference in their entirety.

[0020] The embodiments disclosed herein are merely examples, and the scope of the present disclosure is not limited thereto. Features of any embodiment described in one claim category, e.g., a system, may also be claimed in other claim categories, e.g., a method. Dependencies or back-references in the appended claims are selected for formality reasons only. However, any subject matter available from a careful back-reference to a previous claim may also be claimed, thereby disclosing any combination of claims and their features and may be claimed regardless of the dependencies selected in the appended claims. [Brief explanation of the drawings]

[0021] In the drawings, like reference numbers / letters refer to like elements. [Figure 1] FIG. 1 provides an exemplary infrared (IR) fundus image with various artifacts, including a stripe artifact 11, a central reflection artifact 13, and eyelashes 15 (e.g., seen as dark shadows). [Figure 2] FIG. 1 illustrates example tracking frames (each including an exemplary reference image and an exemplary live image) in which a reference point (from the reference image) and a set of landmarks are tracked in the live IR image. [Figure 3] FIG. 1 illustrates example tracking frames (each including an exemplary reference image and an exemplary live image) in which a reference point (from the reference image) and a set of landmarks are tracked in the live IR image. [Figure 4] FIG. 1 shows the tracking system in the presence of eyelashes 31 and a central reflection 33 in a live IR image 23. [Figure 5] 1A-1C illustrate two additional examples of the present invention. [Figure 6] FIG. 10 shows statistics of test results obtained for registration error and eye movement for different acquisition modes and movement levels. [Figure 7] 10A-10C provide exemplary anterior segment images with variations (over time) in pupil size, iris pattern, eyelid and eyelash movement, and lack of contrast in the eyelid region within the same acquisition. [Figure 8A-8B] FIG. 1 illustrates tracking of a reference point and a set of landmarks in a live image. [Figures 9A-9C] FIG. 1 illustrates tracking of a reference point and a set of landmarks in a live image. [Figure 10] FIG. 2 illustrates an example of a tracking algorithm according to an embodiment of the present invention. [Figure 11] FIG. 11 provides follow-up test results for the embodiment of FIG. 10. [Figure 12] FIG. 1 illustrates the use of an ONH to determine the optimal tracking FOV (tracking window) location for a given OCT FOV (acquisition / scan window). [Figure 13] FIG. 10 illustrates a second embodiment of the present invention for determining optimal tracking FOV position without using OHNs or other predefined physiological landmarks. [Figures 14A-14D] 10A-10C provide additional examples of the present method for identifying an optimal tracking FOV relative to an OCT FOV. [Figure 15] FIG. 1 illustrates scenario 1, where a retinal reference image from a previous visit is available for the patient's fixation. [Figure 16] FIG. 10 illustrates scenario 2 where the retinal image quality algorithm detects the reference image during initial alignment (operator or automatic). [Figure 17] FIG. 10 illustrates scenario 3 where a reference image from a previous visit and a retinal image quality algorithm are not available. [Figure 18] FIG. 10 illustrates an alternative solution for scenario 3. [Figure 19A] FIG. 10 shows two examples for small (top) and normal (bottom) pupil acquisition modes. [Figure 19B] FIG. 1 shows statistics on registration error, eye movements, and number of keypoints for a total of 29,529 images from a series of 45 images. [Figure 20A] FIG. 10 illustrates the motion of the current image (white border) relative to the reference image (gray border) with eye motion parameters Δx, Δy, and rotation φ relative to the reference image. [Figure 20B] Figure 1 shows examples from three different patients: one with good fixation, another with regular eye movements, and a third with random eye movements. [Figure 21] FIG. 1 provides a table showing eye movement statistics for 15 patients. [Figure 22] FIG. 1 illustrates an example of a slit-scan ophthalmic system for imaging the fundus. [Figure 23] FIG. 1 illustrates a generalized frequency-domain optical coherence tomography system used to collect 3D image data of the eye suitable for use in the present invention. [Figure 24] FIG. 1 shows an exemplary OCT B-scan image of a normal retina of a human eye, illustratively identifying various normal retinal layers and boundaries. [Figure 25] FIG. 10 shows an example of an en face vascular image. [Figure 26] FIG. 1 shows an exemplary B-scan of a vasculature (OCTA) image. [Figure 27] FIG. 1 illustrates an example of a multi-layer perceptron (MLP) neural network. [Figure 28] FIG. 1 illustrates a simplified neural network consisting of an input layer, a hidden layer, and an output layer. [Figure 29] FIG. 1 illustrates an exemplary convolutional neural network architecture. [Figure 30] FIG. 1 illustrates an exemplary U-Net architecture. [Figure 31] FIG. 1 illustrates an exemplary computer system (or computing device or computer). DETAILED DESCRIPTION OF THE INVENTION

[0022] The present invention provides an improved eye-tracking system, such as for use with fundus cameras, optical coherence tomography (OCT) systems, and OCT angiography systems. Although the invention is described herein using an infrared (IR) camera tracking any eye in a series of live images, it should be understood that the invention may be practiced using other imaging modalities (e.g., color images, fluorescence images, OCT scans, etc.).

[0023] The tracking system / method may begin by first identifying / detecting a reference point (e.g., a salient physical feature) that is invariantly (e.g., reliably and / or easily and / or quickly) identified within the image. For example, the reference point (or reference template) may correspond to the optic nerve head (ONH) (and its reference location), or another salient / invariant point / feature, such as a lesion. The reference point may be selected from a reference IR image using deep learning or other knowledge-based computer vision methods.

[0024] Alternatively or additionally, a series of candidate points may be identified in the stream of images, and the most invariant candidate point in the set of images may be selected as the reference anchor point for the series of live images. In this way, the anchor points / templates used in the series of live images may change as the quality of the live image stream changes and different candidate points become more salient / invariant.

[0025] The tracking algorithm tracks the live IR image using templates centered on a reference point extracted from the reference IR image. To increase the number of templates and corresponding landmarks in the IR image, additional templates offset from the reference point center are extracted. These templates can be used to detect the same locations in different IR images as a set of landmarks that can be used for alignment between the reference and live IR images, leading to tracking of a series of IR images over time. The advantage of generating a set of templates by offsetting the reference location is that vascular enhancement or advanced image feature detection algorithms are not required. The use of these additional algorithms would impair the real-time performance of the tracking algorithm, especially if the tracking algorithm needs to be run on high-resolution images for more accurate tracking. Given that a set of templates is extracted from the reference IR image, their corresponding locations (as a set of landmarks) in the live IR image are determined by template matching (e.g., normalized cross-correlation) in a small boundary region away from the reference point in the live IR image. Optionally, if there is no match with the initial set of templates, more templates may be searched. Once corresponding matches are found, all or some of the matches can be used to calculate a transformation (x and y shifts and rotation) between the IR reference image and the live IR image. Optionally, if the number of matches is not greater than a threshold (e.g., half of the identified landmarks), the current live image is discarded and not corrected for tracking error. Assuming sufficient matches are found, the transformation determines the amount of motion between the live IR image and the reference image. Theoretically, the transformation can be calculated using two corresponding landmarks (a reference point and a landmark with a high confidence) in the IR reference and live images. However, to ensure more robust tracking, more than two landmarks are used for tracking.

[0026] 2, 3, and 4 show example tracking frames (each including an exemplary reference image and an exemplary live image) in which a reference point (from the exemplary reference image) and a set of landmarks are tracked in the live IR image. In each of FIGS. 2, 3, and 4, the top image (21A, 21B, and 21C, respectively) in each tracking frame is an exemplary reference IR image, and the bottom image (23A, 23B, and 23C, respectively) is an exemplary live IR image. The dotted boxes are ONH templates, and the white boxes are corresponding templates in both the IR reference image and the live image. Templates can be adaptively selected for each live IR image. For example, FIGS. 2 and 3 show the same reference images 21A / 21B and the same anchor point 25, but the additional landmark 27A in FIG. 2 is different from the additional landmark 27B in FIG. 3. In this case, landmarks are dynamically selected in each live IR image based on their detection confidence.

[0027] Figure 4 shows the tracking system in the presence of eyelashes 31 and a central reflection 33 in the live IR image 23. Note that in this case, tracking does not depend on the presence of blood vessels (in contrast to the examples of Figures 2 and 3), thus avoiding confusing eyelashes with blood vessels.

[0028] FIG. 5 shows two additional examples of the present invention. The top row of images shows the present invention applied to a normal pupil acquisition, and the bottom row of images shows the present invention applied to a small pupil acquisition. In both cases, landmarks are detected relative to the ONH position (e.g., fiducial anchor point). In this example, tracking parameters (xy translation and rotation) are calculated by registration between a reference image (RI) and a moving image (MI). The method requires the location of the ONH 41 and a set of RI landmarks (e.g., auxiliary points) 43 extracted from feature-rich regions of the reference image RI. For illustrative purposes, one of the landmarks 43 is shown within a bounding region or search window 42, and its relative distance 45 to the ONH 41 is indicated. The ONH 41 in the reference image RI can be detected using a neural network system with a U-Net architecture. A schematic description of a neural network including a U-Net architecture is provided below. The ONH 41' in the dynamic (e.g., live) image MI can be detected by template matching using the ONH template 41 extracted from the reference image RI. Each reference landmark template 43 and its relative distance 45 (and optionally relative orientation) to the ONH 41 is used to search for corresponding landmarks 43' (e.g., within a boundary region or window 42') that have the same / similar distance 45' from the ONH 41' in the dynamic image MI. Some landmark correspondences with high confidence are used to calculate tracking parameters.

[0029] In an exemplary embodiment, infrared (IR) images (11.52 × 9.36 mm with a pixel size of 15 μm / pixel) were collected at a frame rate of 50 Hz using a Clarus 500 (Zeiss, Dublin, CA) using normal and small-pupil acquisition modes with guided eye movement. Each eye was scanned using three different movement levels: good fixation, regular eye movement, and random eye movement. The registered images were displayed in a single image to visualize the registration (see Figure 5). The average distance error between the registered movement landmarks and the reference landmarks was calculated as the registration error. Registration error and eye movement statistics were reported for each acquisition mode and each movement level. Approximately 500 images were collected from 15 eyes.

[0030] Figure 6 shows the obtained statistics for alignment error and eye movement for different acquisition modes and movement levels using all eyes. The mean and standard deviation of alignment error for normal and small-pupil acquisition modes are similar, indicating that the tracking algorithm has similar performance for both modes. The reported alignment error is important information to help design OCT scan patterns. The tracking time for a single image was measured to be an average of 13 ms using a computing system with an Intel i7 2.6 GHz CPU and 32 GB of RAM. Thus, the present invention provides a real-time retinal tracking method using IR images, an important part of the OCT image acquisition system, with demonstrated good tracking performance.

[0031] Although the above examples are described as being applied to the posterior segment of the eye (e.g., the fundus), it is understood that the present invention can also be applied to other parts of the eye (e.g., the anterior segment of the eye). Real-time and efficient tracking of anterior segment images is important in automated OCT angiography image acquisition. Anterior segment tracking is important due to involuntary eye movement during image acquisition, particularly in OCTA scans. Anterior segment LSO images can be used to track the movement of the anterior segment of the eye. Eye movement can be assumed to be rigid body motion, with motion parameters such as translation and rotation that can be used to steer the OCT beam.

[0032] Local motion and lack of contrast in the anterior segment of the eye (e.g., changes in pupil size / shape, constant eyelash and eyelid movement, constricted or dilated iris patterns during tracking due to changes in pupil size / shape, etc.) can affect automated real-time processing, thereby reducing the success and reliability of tracking. Additionally, the appearance of anatomical features in an image can vary significantly over time depending on the subject's gaze (e.g., gaze angle). Figure 7 provides an example anterior segment image with changes in pupil size, iris pattern, eyelid and eyelash movement (over time), and lack of contrast in the eyelid region within the same acquisition. Therefore, a method is needed that can robustly track the anterior segment of the eye using LSO or other imaging modalities in real time.

[0033] Conventional anterior segment tracking systems use a reference image with a set of landmarks extracted from the image. Then, a tracking algorithm uses the landmarks extracted from the reference image to track a series of live images by independently searching for landmarks in each live image. That is, matching landmarks between the reference image and the live image are determined independently. Independent matching of landmarks between two images assuming a rigid (or affine) transformation is a challenging problem due to local motion and lack of contrast. Advanced landmark matching algorithms have typically been required to calculate the rigid transformation. When high-resolution images are required for more accurate tracking, the real-time tracking performance of such conventional approaches is typically compromised.

[0034] While the tracking embodiments described above (see, e.g., FIGS. 2-6 ) provide efficient landmark matching detection between two images, some embodiments may have limitations. Some of the above-described embodiment(s) assume that the reference image and the live image contain obvious or unique anatomical features, such as the ONH, that are robustly detectable due to the uniqueness of the anatomical features. In this approach, landmarks are detected relative to a reference (anchor) point in the live image (e.g., the ONH, a lesion, or a specific vascular pattern). The distance (or relative distance) between the reference point and the selected landmark in the reference image and the live image is maintained constant in both images. Therefore, landmark detection in the live image becomes a simpler problem by searching a small region at a known distance (and optionally orientation) from the reference (anchor) point. The robustness of landmark detection is guaranteed by the constant distance between the reference point and the landmark location.

[0035] In contrast, the present embodiment has several advantages over the above-described embodiment(s). Similar to the above-described embodiment(s), a reference (e.g., anchor) point is selected from candidate landmarks extracted from the reference image. However, the selected reference anchor point does not necessarily have to be an obvious or inherent anatomical / physical feature of the eye (e.g., ONH, pupil, iris border, or iris center). Although the anchor point may not be an inherent anatomical / physical feature, the distance between the reference anchor point and the selected auxiliary landmark in the reference image and the live image is maintained constant in both images. Therefore, landmark detection in the live image becomes a simpler problem by searching a small region at a known distance from the reference anchor point. The robustness of landmark detection is guaranteed by the constant distance between the reference anchor point and the landmark auxiliary point. Some best-matching landmarks can be selected by some exhaustive search to calculate a rigid transformation. A similar approach may also be applied to retinal tracking using IR images (embodiment described above), where no unique anatomical landmarks are visible (or found) within the image / scan field or within the field of view (e.g., periphery) of the detector.

[0036] A difference in this embodiment compared to some of the above-described embodiments is that the reference (anchor) point is selected from a group of landmark candidates extracted from the reference image. The reference point may be selected / picked based on, for example, its trackability in subsequent images (e.g., in a stream of images) to ensure consistent and robust detection of the point. Essentially, temporal image information (e.g., changes in a series of images over time) is incorporated into the reference point selection method. For example, all images in a series of images or selected images (e.g., images selected at fixed or variable intervals) may be examined to determine whether the current landmark is still the best landmark to use as the reference anchor landmark. As a different landmark candidate becomes more trackable (e.g., more easily, more quickly, more uniquely, and / or more consistently detectable), it replaces the previous reference point and becomes the new reference anchor point. All other landmark points may then be re-referenced relative to the new reference anchor point.

[0037] This embodiment is particularly useful in situations where the scan (or image) area (field of view) of the eye does not contain an obvious or unique anatomical feature (e.g., ONH), or where an anatomical feature is not necessarily useful for selection as a reference point (e.g., a pupil whose size / shape changes during tracking, e.g., over time). Thus, this embodiment does not require the uniqueness of the anatomical feature selected as a reference point.

[0038] This embodiment may first detect reference (anchor) points from a set of candidate landmarks extracted from a reference image or a series of consecutive live images. For example, reference landmark candidates may reside in regions with significant texture characteristics, such as the iris region toward the outer edge of the iris. An entropy filter may highlight the regions with significant texture characteristics, followed by additional image processing and analysis techniques to generate a mask containing the landmark candidates to be selected as reference point candidates. Reference points located in areas with high contrast and texture that are trackable in subsequent live images may be selected as reference (anchor) points. Deep learning (e.g., neural network) methods / systems may be used to identify image regions with high contrast and significant texture characteristics.

[0039] This embodiment tracks the live image using templates centered on fiducial (anchor) points extracted from the reference image. Additional templates are generated centered on candidate landmarks extracted from the reference image. These templates can be used to find the same locations in the live image as a set of landmarks that can be used for registration between the reference image and the live image, leading to tracking of a series of images over time. Given a set of templates extracted from the reference image, their corresponding locations (as a set of landmarks) in the live image are determined by template matching (e.g., normalized cross-correlation) in a small region away from the fiducial points in the live image. Once all corresponding matches are found, some matches are used to calculate the transformation (x and y shifts and rotations) between the reference image and the live image.

[0040] The transformation determines the amount of motion between the live image and the reference image. The transformation can be calculated using two corresponding landmarks in the reference image and the live image. However, to ensure the robustness of the tracking, more than two landmarks can be used for tracking.

[0041] Some matching landmarks may be determined by an exhaustive search. For example, at each iteration, two pairs of corresponding landmarks may be selected from the reference image and the live image. A rigid transformation may then be calculated using the two pairs. The error between each transformed reference image landmark (using the rigid transformation) and the live image landmark is determined. Landmarks associated with an error smaller than a predetermined threshold may be selected as inliers. This procedure may be repeated for all (or most) traversable two-pairs (e.g., combinations of two-pairs). The transformation that produces the largest number of inliers may then be selected as the rigid transformation for tracking.

[0042] 8 and 9 illustrate the tracking of a reference point and a set of landmarks in a live image. The images in the left column are reference images, and the images in the right column are live images. The selected reference point is marked with a circled cross for each reference image. As shown, the selection of the reference anchor point changes over time as various portions of the live image stream change (e.g., in shape or quality). Pupil size and shape, eyelid movement, and low-contrast changes are visible in the examples of FIGS. 8 and 9.

[0043] In summary, motion artifacts pose a challenge in optical coherence tomography angiography (OCTA). While motion tracking solutions exist to correct these artifacts in retinal OCTA, the motion tracking problem remains unsolved for the anterior segment (AS) of the eye. This currently poses an obstacle to the use of AS-OCTA for the diagnosis of diseases of the cornea, iris, and sclera. The present embodiment is demonstrated for motion tracking of the anterior segment of the eye.

[0044] In a specific embodiment, a telecentric add-on lens assembly with internal fixation was used to enable imaging of the anterior segment with a CIRRUS™ 6000 AngioPlex (ZEISS, Dublin, CA) at good patient alignment and fixation (fx). Using this add-on lens on the CIRRUS™ 6000, wide-field (20 × 14 mm) line scanning ophthalmoscope (LSO) image sets were acquired from 25 eyes of 15 subjects, resulting in a total of 6973 images (4798 at central fixation and 2175 at peripheral fixation). Motion in these image sets was then tracked by an algorithm using real-time landmark-based rigid registration between the reference image and other (dynamic) images from the same set.

[0045] FIG. 10 illustrates an example of a tracking algorithm according to an embodiment of the present invention. In this example, anchor points and selected landmarks are found in the dynamic image and used to calculate translation and rotation values ​​for registration. The bottom overlay image in FIG. 10 is shown for visual verification. In this embodiment, anchor points are first detected in areas of the reference image with high texture values. Next, this anchor point is identified in the dynamic image by searching for a template (image region) centered on the reference image anchor point location. Next, landmarks from the reference image are found in the dynamic image by searching for landmark templates whose distance to the anchor point is the same as in the reference image. Finally, the landmark pair with the highest confidence value is used to calculate the translation and rotation. The registration error is the average distance between corresponding landmarks in both images. This value is calculated after visually confirming the landmark match and confirming successful registration.

[0046] Figure 11 provides the tracking test results. The insets of the registration error histogram and rotation angle histogram show the individual distribution parameters. The translation vector is plotted at the center, with concentric rings every 500 μm. The inset shows the distribution parameters for the translation magnitude.

[0047] As mentioned above, real-time, efficient tracking of IR fundus images is important in automated retinal OCT image acquisition. Retinal tracking becomes more challenging when the patient's gaze is not straight or off-center, assuming that the tracking and OCT acquisition fields of view (FOVs) are located over the same retinal region. That is, the tracking and OCT (acquisition) FOVs are typically located over the same region on the retina. In this way, motion tracking information can be used, for example, to correct OCT positioning during OCT scanning. However, the location where the OCT scan is being performed (e.g., the OCT acquisition FOV or OCT FOV) may not be at a location on the retina with sufficiently unique physical features / structures to track motion in a robust manner. A currently preferred approach is to identify / determine the optimal tracking FOV location prior to tracking and OCT acquisition.

[0048] There are several challenges that complicate efficient tracking. For example, IR images may not contain enough dispersed retinal features (such as blood vessels) to track when fixation is off-center. Another complication is that the curvature of the eye is greater in the peripheral regions of the eye. Furthermore, eye movement can generate more nonlinear distortion in the current image relative to the reference image, which can lead to inaccurate tracking. The transformation between the current image relative to the reference image may not be a rigid transformation due to the nonlinear relationship between the two images.

[0049] One possible solution is to place the tracking FOV where there are clear retinal landmarks and features that can be robustly detected (e.g., around the ONH and large blood vessels). The problem with this approach is that the large distance between the tracking FOV and the OCT acquisition FOV introduces rotation angle errors due to the location of the rotation anchor point being within the tracking FOV but not within the OCT acquisition FOV.

[0050] To overcome the above challenges, the tracking FOV (e.g., in the IR image) can be made to at least partially overlap with the OCT acquisition FOV for off-center fixation, or can be positioned as close as possible to the OCT acquisition FOV.

[0051] A method for optimal and dynamic positioning of a tracking FOV with respect to a patient's fixation is presented herein. In this approach, a tracking algorithm (such as that described above or another suitable tracking algorithm) is used to optimize the position of the tracking FOV by maximizing tracking performance using a set of metrics such as tracking error, landmark (keypoint) distribution, and number of landmarks. The tracking position (center of the FOV) that maximizes tracking performance is selected / designated / identified as the desired position for a given patient's fixation.

[0052] Essentially, the present invention dynamically finds an optimal tracking FOV (for a given patient fixation) that enables good tracking for OCT scans of off-center fixations (e.g., fixations in the peripheral region of the eye). In this approach, the tracking region can be located at a different retinal location than the OCT scan region. The optimal region on the retina for the OCT FOV is identified and used for retinal tracking. Thus, the tracking FOV can be dynamically found for each eye. For example, the optimal tracking FOV can be determined based on tracking performance for a series of alignment images.

[0053] FIG. 12 illustrates the use of an ONH to determine the location of the optimal tracking FOV (tracking window) for a given OCT FOV (acquisition / scan window). In this example, an IR preview image (or a section / window within the IR preview image), typically with a wide FOV (e.g., a 90-degree FOV), can be used for patient alignment and to define the tracking FOV. These images, along with an IR tracking algorithm, can be used to determine the optimal tracking position for the OCT FOV. The dotted box 61 defines the OCT FOV, and the dashed boxes 63A and 63B indicate moved (repositioned) non-optimal tracking FOVs that are moved until an optimal tracking FOV (solid black box 65) is identified. The non-optimal tracking FOVs 63A / 63B are moved toward the ONH within a distance to the OCT FOV center (e.g., indicated by the solid white line 67). The optimal tracking FOV (solid black box 65) in the reference IR preview image enables robust tracking of the remaining IR preview images within the same tracking FOV.

[0054] Two embodiments (or implementations) of the present invention are provided herein. The first embodiment uses a reference point. This embodiment relies on a detectable reference point on the retina. The reference point can be, for example, the center of the optic nerve head (ONH). The embodiment can be summarized as follows: 1) Collect a series of wide (e.g., 90 degree) FOV IR preview images (such as those used in patient alignment), or other suitable fundus images. 2) Use / designate one of the collected images as the reference image. 3) Detect the ONH center in the reference image. 4) Crop the tracking FOV at the center of the OCT FOV in the reference image and use the cropped FOV as the tracking reference image (e.g., the current non-optimal tracking FOV 63A / 63B). 5) The tracking reference image is used to track the remaining IR preview images in the set. 6) Update the objective function. As is known in the art, the objective function in a mathematical optimization problem is a real-valued function whose value is minimized or maximized over a set of feasible alternatives. In this case, the objective function value is updated using tracking outputs such as tracking error, landmark distribution, and number of landmarks for all remaining IR preview images. 7) Update the tracking reference image by cropping the tracking FOV toward the ONH center along a connecting line (e.g., line 67) (between the tracking FOV center and the OCT FOV center). An alternative to the connecting line can be a nonlinear dynamic path from the tracking FOV center to the OCT FOV center. The nonlinear dynamic path can be determined for each scan / eye. 8) Repeat steps 5) to 8) until the objective function is minimized for the maximum allowable distance between the OCT FOV and the optimal tracking FOV (e.g., constrained optimization).

[0055] FIG. 13 illustrates a second embodiment of the present invention for determining an optimal tracking FOV position without using OHNs or other predefined physiological landmarks. All components similar to those in FIG. 12 have similar reference numbers and are described above. This approach searches for an optimal tracking FOV 65 around the OCT FOV 61. The optimal location of the tracking FOV (black solid box) within the reference IR preview image enables robust tracking of the remaining IR preview images within the same tracking FOV. FIGS. 14A, 14B, 14C, and 14D provide additional examples of this method for identifying an optimal tracking FOV relative to the OCT FOV.

[0056] The second proposed solution / embodiment does not use a reference point. This approach can be summarized as follows: 1) Acquire a series of wide FOV IR preview images (e.g., those used in patient alignment). 2) Use one of the images as a reference image. 3) Crop the tracking FOV at the center of the OCT FOV in the reference image and use the cropped FOV as the tracking reference image. 4) Track the remaining IR preview images in the set against the tracking reference image. 5) Update the objective function value using the tracking outputs such as tracking error, landmark distribution, and number of landmarks for all remaining IR preview images. 6) Update the tracking reference image by cropping the tracking FOV towards regions of the IR preview reference image that have many anatomical features such as blood vessels and lesions. Image saliency techniques can be used to update the tracking FOV position. 7) Repeat steps 4) to 7) until the objective function is minimized for the maximum allowable distance between the OCT and tracking FOV (constrained optimization).

[0057] The retinal tracking methods described above may also be used for automatic capture (e.g., of OCT scans and / or fundus images). This is in contrast to prior art methods that use pupil tracking for automatic alignment and capture, although the present invention's use of retinal tracking for alignment and automatic capture is not known in prior art approaches.

[0058] Automated patient alignment and image capture creates a positive and effective operator and patient experience. After initial alignment by the operator, the system can perform automatic tracking and OCT acquisition. Fundus images can be used to register the scan area on the retina. However, automatic capture can be challenging due to eye movement during alignment, blinking and partial blinking, and alignment stability, which can rapidly cause the device to become misaligned due to, for example, eye movement, focusing, or operator error.

[0059] A retinal tracking algorithm can be used to lock onto the fundus image and track the incoming dynamic image. The retinal tracking algorithm requires a reference image to calculate the geometric transformation between the reference image and the dynamic image. The tracking algorithm for automatic alignment and automatic capture can be used in different scenarios.

[0060] For example, the first scenario (Scenario 1) is when a reference image of the retina from a previous clinic visit is available for the patient's fixation. In this case, the reference image can be used to register and trigger automatic capture when the eye is stable (no or minimal movement), and a series of retinal images are robustly tracked. The second scenario (Scenario 2) is when the retinal image quality algorithm detects the reference image during initial alignment (by an operator or automatically). In this case, the reference image detected by the image quality algorithm can be used similarly to Scenario 1 to register and trigger automatic capture. The third scenario (Scenario 3) is when a reference image from a previous clinic visit and the retinal image quality algorithm are not available. In this case, the algorithm can track a series of images as reference images, starting with the last image in the previous series. The algorithm can repeat this process until a consecutive series of images is tracked continuously and robustly, thereby triggering automatic capture.

[0061] This embodiment describes an automatic alignment and automatic capture technique that uses the methods described in the above three scenarios for using a retinal tracking system. The basic idea is to use a tracking algorithm to evaluate the fundus image (e.g., determine whether the fundus image is a good quality retinal image for a given fixation) and evaluate eye movement relative to the fixation position during alignment. If eye movement is minimal at the fixation position, automatic capture is triggered. In addition, the tracking output (e.g., xy translation and rotation relative to the fixation position) can also be used for automatic alignment in a motorized system by moving hardware components such as a chin rest or headrest, eyepieces, etc.

[0062] As mentioned above, IR preview images (90 degree FOV) are typically used for patient alignment. These images can be used in conjunction with an IR tracking algorithm such as the one described above or other known IR tracking algorithms to determine whether the image can be tracked continuously and robustly using the reference images. Below are some embodiments suitable for use in the three different scenarios described above.

[0063] FIG. 15 illustrates Scenario 1, where a reference image of the retina from a previous visit is available for the patient's fixation. In this scenario, alignment and automatic capture are straightforward because the reference image is known for a given fixation. FIG. 15 illustrates tracking each dynamic image using the reference image. The dotted box represents a non-trackable image, while the dashed box represents a trackable image. Tracking quality determines whether the image is at the correct fixation and of good quality. Tracking quality can be measured using tracking outputs such as tracking error, landmark distribution, number of landmarks, x-y translation, and rotation of the dynamic image relative to the reference image (as described above). Tracking outputs can also be used for automatic alignment in motorized systems by moving hardware components such as a chin rest or head rest, eyepieces, etc.

[0064] Automatic capture can be triggered when a predefined number N of consecutive dynamic images can be robustly tracked (e.g., with a predefined confidence or quality metric). This indicates minimal patient eye movement and accurate fixation. The tracking output can also be used to guide the operator or patient (graphically or by using sound / verbal / text) for better alignment.

[0065] FIG. 16 illustrates Scenario 2, in which a retinal image quality algorithm detects a reference image during initial alignment (by an operator or automatically). In this scenario, a reference image is detected from a series of dynamic images during alignment using an appropriate IR image quality algorithm. The operator performs an initial alignment to bring the retina into the desired field of view and fixation. The IR image quality algorithm then determines the quality of the series of dynamic images. A reference image is then selected from a set of candidate reference images. The best reference image is selected based on the image quality score. Once a reference image is selected, automatic capture or alignment can be triggered, as described above with reference to Scenario 1.

[0066] In Scenario 3, no reference images from a previous visit and no retinal image quality algorithm are available. In this scenario, the algorithm tracks a series of images starting from the last image in the previous series as the reference image (solid white box). Figure 17 shows that the algorithm repeats this process until a series of consecutive images are continuously and robustly tracked (dashed box), which can trigger automatic capture. This approach allows the operator to perform initial alignment to place the retina at the desired field of view and fixation.

[0067] The number of images in the series depends on the tracking performance, for example, if tracking is not possible, the new series can start with a new reference image from the last image in the previous series.

[0068] Figure 18 shows an alternative solution for Scenario 3. This approach allows for the selection of a reference image from a continuous, robustly tracked sequence of images. Once the reference image is selected, automatic capture or alignment can be triggered, similar to the method in Scenario 1.

[0069] The image tracking application described above can be used to extract various statistics to identify various characteristic problems affecting tracking. For example, a series of images used for tracking can be analyzed to determine whether the images have characteristics of systemic movement, random movement, or good fixation. An ophthalmic system using the present invention can then notify the system operator or the patient of problems that may affect tracking and provide suggested solutions.

[0070] Various types of artifacts in OCT can affect the diagnosis of eye diseases. Off-center artifacts and motion artifacts are important artifacts. Off-center artifacts are due to fixation errors and cause a displacement of the analysis grid on the topography map for certain disease types. Off-center artifacts most often occur in subjects with poor attention, poor visual acuity, or eccentric fixation. Even when patients are asked to fixate, involuntary eye movements with different intensities in different directions still occur during alignment and acquisition.

[0071] Motion artifacts are caused by ocular saccades, changes in head position, or respiratory movements. Motion artifacts can be overcome by eye-tracking systems. However, eye-tracking systems generally cannot handle saccadic movements or patients with poor attention or poor vision. In these cases, the scan cannot be fully completed.

[0072] To improve patient fixation and reduce distraction (especially for longer scan times), especially in patients with poor attention or poor vision, eye movement analysis during alignment and acquisition can be a useful tool to notify the operator and patient that more careful attention is needed for better fixation or eye movement control. For example, visual notifications for the operator and audio notifications for the patient may be provided, which may lead to more successful scans. The operator can adjust hardware components according to the movement analysis output. The patient can be guided to the fixation target until the scan is completed.

[0073] In this embodiment, a method of eye movement analysis is described. The basic idea is to use the retinal tracking output for real-time or post-acquisition analysis to generate a set of messages, including audio messages, that can inform the operator and patient during alignment and acquisition about the state of fixation and eye movement. Providing the movement analysis results after acquisition can help the operator understand the reasons for poor scan quality, so that the operator can take appropriate action that may lead to a successful scan.

[0074] Eye movement analysis can be useful in the following cases: 1) Self-alignment: The patient can receive instructions from the device to align themselves. 2) Automatic acquisition: During acquisition, the patient can be notified of any change in fixation or large movements. Fixation in the same position is important in small scan fields. 3) If a fixation target is not available, a message (e.g., in the form of a voice) can keep the patient in fixation. 4) The eye movement analysis results can be used in post-processing algorithms to resolve residual movement.

[0075] The above-described real-time retinal tracking method for off-center fixation using infrared reflectance (IR) images was tested in a proof-of-concept application. As explained above, the OCT acquisition system relies on a robust, real-time retinal tracking method to capture reliable OCT images for visualization and further analysis. Tracking the retina at off-center fixation can be challenging due to the lack of many relevant anatomical features in the images. The robust, real-time retinal tracking algorithm proposed in this application detects at least one anatomical feature with high contrast as a reference point (RP) to improve tracking performance.

[0076] In this example, as described above, real-time keypoint (KP)-based registration between the reference image and the dynamic image calculates xy translation and rotation as tracking parameters. The tracking method relies on a unique RP extracted from the reference image and a set of reference image KPs. The location of the RP in the reference image is robustly detected using a fast image saliency method. Any suitable saliency method known in the art can be used. Examples of saliency include: (1) X. Hou and L. Zhang, "Saliency Detection: CVPR dual Approach," CVPR, 2007; (2) C. Guo, Q. Ma, and L. Zhang, "Spatio-temporal saliency detection using phase spectrum of quaternion fourier transform," CVPR, 2008; and (3) B. Schauerte, B. Kuhn, K. Kroschel, and R. Stiefelhagen, "Multimodal Saliency-based Attention for Object-based Scene Analysis." Analysis,” IROS, 2011.

[0077] Figure 19A shows two examples for small (upper) and normal (lower) pupil acquisition modes. RP(+) is detected using a fast image saliency algorithm. Keypoints (white circles) are detected based on the RP position. The RP position in the dynamic image is detected by template matching using RP templates extracted from the reference image. Each reference KP template and its relative distance to the RP are used to search for the corresponding motion KP with the same distance from the RP position in the dynamic image, as indicated by the dashed arrows.

[0078] In this exemplary embodiment, tracking parameters were calculated from a subset of KP correspondences with high confidence. Using prototype software, a series of IR images (11.52 × 9.36 mm, pixel size 15 μm / pixel, and frame rate 50 Hz) was collected from a CLARUS 500 (Zeiss, Dublin, California). The registered images were displayed in a single image to visualize the registration (e.g., the rightmost image in Figure 19A). For each dynamic image, the average distance error between the registered motion KP and the reference KP was calculated as the registration error.

[0079] Statistics on registration error, number of KPs, and eye movement were reported. Figure 19B shows the statistics on registration error, eye movement, and number of keypoints for a total of 29,529 images from 45 image series. Forty-five image series, each with an average of 650 images, were collected from one or both eyes of 10 subjects / patients. Patient fixation was off-center. The average registration error of 15.3 ± 2.7 μm indicates that accurate tracking is possible in the OCT domain with A-scan intervals greater than 15 μm. The tracking execution time was measured at an average of 15 ms using an Intel i7-8850H CPU-H 2.6 GHz CPU and 32 GB of RAM. Thus, this embodiment demonstrates the robustness of the present tracking algorithm, which is based on a real-time retinal tracking method using IR fundus images. This tool can be an important part of any OCT image acquisition system.

[0080] Eye-tracking-based analyses often aim to identify and analyze an individual's visual attention patterns as they perform specific tasks, such as reading, searching, scanning images, or driving. For eye movement analysis, the anterior segment of the eye (e.g., pupil and iris) is used. This method uses retinal tracking outputs (eye movement parameters) for each frame of a line-scan ophthalmoscope (LSO) or infrared-reflectance (IR) fundus image. The eye movement parameters (x, y, e.g., translation and rotation) recorded over a period of time can be used for statistical analysis, which may include statistical moment analysis of the eye movement parameters. Future eye movements can be predicted using time series analysis, such as Kalman filtering and particle filters. The system can also generate messages to inform the operator and patient using statistical and time series analysis.

[0081] In the present invention, eye movement analysis can be used during and / or after acquisition. Retinal tracking algorithms using LSO or IR images can be used to calculate eye movement parameters such as x and y translation and rotation. The movement parameters are calculated at initial fixation or relative to a reference image captured using one of the methods described above. Figure 20A shows the movement of the current image (white border) relative to the reference image (gray border) due to eye movement parameters Δx, Δy, and rotation φ relative to the reference image. The current image was aligned with the reference image, and then the two images were averaged.

[0082] For each fundus image, eye movement parameters are recorded, which can be used for statistical analysis over a period of time. Examples of statistical analysis include statistical moment analysis of eye movement parameters. Time series analysis can be used to predict future eye movement. Prediction algorithms include Kalman filtering and particle filters. Statistical analysis and time series analysis can be used to generate informative messages to inform operators and patients regarding their actions.

[0083] In embodiments with motion analysis during alignment and acquisition, time series analysis for eye movement prediction (next position and velocity) can alert the patient if they are drifting away from their initial fixation position.

[0084] In embodiments with post-acquisition motion analysis, statistical analysis can be applied after the current acquisition is completed (whether the acquisition was successful or unsuccessful). One example statistical analysis includes the overall fixation offset (average xy movement) from the initial fixation position and the distribution (standard deviation) of eye movements as a measure of eye movement severity during acquisition. Figure 20B shows examples from three different patients: one with good fixation, another with regular eye movements, and a third with random eye movements. Eye movement calculations can be applied to the IR image relative to a reference image with an initial fixation.

[0085] Figure 21 provides a table showing eye movement statistics for 15 patients. The mean indicates the overall fixation offset from the initial fixation position. The standard deviation indicates a measure of eye movement within an acquisition. Scans containing regular or random eye movements exhibited significantly larger mean values ​​and standard deviations compared to scans with good fixation, which can be used as an indicator of poor fixation. A significantly larger mean value or standard deviation could be defined as 116 microns and 90 microns, respectively, for this study. Eye and fixation analysis can highlight its use as feedback to the operator or patient by providing helpful messages for reducing movement in OCT image acquisition, which is important for any subsequent data processing.

[0086] The following describes various hardware and architectures suitable for the present invention. Fundus Imaging System Two categories of imaging systems used to image the fundus are flood-illumination imaging systems (or flood-illumination imagers) and scanning-illumination imaging systems (or scanning imagers). Flood-illumination imagers simultaneously flood the entire field of view (FOV) of interest of the object with light, such as with a flash lamp, and capture a full-frame image of the object (e.g., the fundus) using a full-frame camera (e.g., a camera having a two-dimensional (2D) photosensor array sized collectively to capture the desired FOV). For example, a flood-illumination fundus imager floods the fundus of the eye with light and captures a full-frame image of the fundus in a single image capture sequence of the camera. Scanning imagers provide a scanning beam that is scanned across the object, e.g., the eye, and the scanning beam is imaged at different scanning locations as the scanning beam is scanned across the object, creating a series of image segments that can be reconfigured, e.g., combined, to create a composite image of the desired FOV. The scanning beam can be a point, a line, or a two-dimensional region such as a slit or a wide line. Examples of fundus imaging devices are described in US Pat. Nos. 8,967,806 and 8,998,411.

[0087] FIG. 22 illustrates an example of a slit-scanning ophthalmic system SLO-1 for imaging the fundus F, which is the inner surface of the eye E opposite the ocular lens (or crystalline lens) CL and may include the retina, optic disc, macula, fovea, and posterior pole. In this example, the imaging system is in a so-called “scan-descan” configuration, in which a scanning line beam SB traverses the optical components of the eye E (including the cornea Crn, iris Irs, pupil Ppl, and lens) to scan across the entire fundus F. In the case of a projected light fundus imaging device, no scanner is required, and light is irradiated across the entire desired field of view (FOV) at once. Other scanning configurations are known in the art, and the specific scanning configuration is not critical to the present invention. As shown, the imaging system includes one or more light sources LtSrc, preferably a polychromatic LED system or a laser system with a suitably adjusted etendue. An optional slit Slt (adjustable or stationary) may be positioned in front of the light source LtSrc and used to adjust the width of the scanning line beam SB. Additionally, the slit Slt can remain stationary during imaging or can be adjusted to different widths to allow for different confocal levels and different applications for a particular scan or during scanning used for reflection suppression. An optional objective lens ObjL can be placed before the slit Slt. The objective lens ObjL can be any one of several state-of-the-art lenses, including, but not limited to, refractive, diffractive, reflective, or hybrid lenses / systems. Light from the slit Slt passes through a pupil-splitting mirror SM and is directed to the scanner LnScn. It is desirable to keep the scan plane and pupil plane as close as possible to reduce vignetting in the system. An optional optical system DL can be included to manipulate the optical distance between the images of the two components. The pupil-splitting mirror SM can pass the illumination beam from the light source LtSrc to the scanner LnScn and reflect the detection beam from the scanner LnScn (e.g., reflected light returning from the eye E) toward the camera Cmr. The task of the pupil-splitting mirror SM is to split the illumination beam from the scanner LnScn and assist in suppressing system reflections.Scanner LnScn can be a rotating galvo scanner or other type of scanner (e.g., piezoelectric or voice coil, microelectromechanical system (MEMS) scanner, electro-optic deflector, and / or rotating polygon scanner). Depending on whether pupil splitting occurs before or after scanner LnScn, the scan can be split into two steps with one scanner in the illumination path and a separate scanner in the detection path. Specific pupil splitting arrangements are described in detail in U.S. Pat. No. 9,456,746, the entire contents of which are incorporated herein by reference.

[0088] From the scanner LnScn, the illumination beam passes through one or more optical systems—in this case, a scan lens SL and an ophthalmic or ocular lens OL—that enable the pupil of the eye E to be imaged onto the system's image pupil. Generally, the scan lens SL receives the scanning illumination beam from the scanner LnScn at any of a number of scan angles (angles of incidence) and generates a scan line beam SB having a substantially planar focal plane (e.g., a collimated optical path). The ophthalmic lens OL can focus the scan line beam SB onto an imaging object. In this example, the ophthalmic lens OL can focus the scan beam SB onto the fundus F (or retina) of the eye E to image the fundus. In this way, the scan line beam SB creates a transverse scan line that moves across the fundus F. One possible configuration of these optical systems is a Keplerian telescope, in which the distance between the two lenses is selected to generate a nearly telecentric intermediate fundus image (4-f configuration). The ophthalmic lens OL can be a single lens, an achromatic lens, or an arrangement of different lenses. All lenses can be refractive, diffractive, reflective, or hybrid, as known to those skilled in the art. The size and / or shape of the ophthalmic lens OL, the scanning lens SL, the pupil-splitting mirror SM, and the scanner LnScn can vary depending on the desired field of view (FOV). Therefore, an arrangement can be envisioned in which multiple components can be switched in and out of the beam path, for example, by using a flip of the optical system, a motorized wheel, or removable optical elements, depending on the field of view. Because a change in field of view results in a different beam size on the pupil, the pupil splitting can also be changed to accommodate a change in FOV. For example, a field of view of 45° to 60° is a typical or standard FOV for a fundus camera. Higher fields of view, such as wide-field FOVs of 60° to 120° or more, may also be feasible. A wide-field FOV may be desirable for combining a wide-line fundus imager (BLFI) with another imaging modality, such as optical coherence tomography (OCT). The upper limit of the field of view may be determined by the accessible working distance combined with the physiological conditions around the human eye. Since a typical human retina has an FOV of 140° horizontally and 80°-100° vertically, it may be desirable to have an asymmetric field of view with as high an FVO as possible on the system.

[0089] The scanning beam SB passes through the pupil Ppl of the eye E and is directed onto the retina or fundus, i.e., surface F. The scanner LnScn1 adjusts the position of the light on the retina or fundus F so that a range of lateral positions of the eye E is illuminated. The reflected or scattered light (or emitted light in the case of fluorescence imaging) is directed along a path similar to the illumination, defining a focused beam CB on a detection path to the camera Cmr.

[0090] In the “scan-descan” configuration of the exemplary slit-scanning ophthalmic system SLO-1 of the present invention, light returning from eye E is “descanned” by scanner LnScn on its way to pupil-splitting mirror SM. That is, scanner LnScn scans illumination beam SB from pupil-splitting mirror SM to define a scanning illumination beam SB across eye E, but because scanner LnScn also receives returning light from eye E at the same scanning position, it effectively descans the returning light (e.g., cancels the scanning motion) to define a non-scanning (e.g., stationary) focused beam from scanner LnScn to pupil-splitting mirror SM, which folds the focused beam toward camera Cmr. At pupil-splitting mirror SM, reflected light (or emitted light, in the case of fluorescence imaging) is separated from the illumination light onto a detection path that is directed to camera Cmr, which may be a digital camera with a photosensor for capturing an image. An imaging (e.g., objective) lens ImgL may be positioned in the detection path so that the fundus is imaged onto camera Cmr. As with the objective lens ObjL, the imaging lens ImgL can be any type of lens known in the art (e.g., refractive, diffractive, reflective, or hybrid lens). Additional operational details, particularly methods for reducing artifacts in images, are described in International Publication WO 2016 / 124644, the entire contents of which are incorporated herein by reference. The camera Cmr captures the received images and, for example, creates an image file, which can be further processed by one or more (electronic) processors or computing devices (e.g., the computer system of FIG. 20). Thus, focused beams (returned from all scanning positions of the scanning line beam SB) are collected by the camera Cmr, and a full-frame image Img can be constructed, such as by montage, from a combination of the individually captured focused beams. However, other scanning configurations are also envisioned, including those in which the illumination beam is scanned across the eye E and the focused beam is scanned across the camera's photosensor array.WO 2012 / 059236 and U.S. Patent Application Publication No. 2015 / 0131050, which are incorporated herein by reference, describe several embodiments of scanning slit ophthalmoscopes, including various designs, such as designs in which the returning light is swept across the camera's photosensor array, and designs in which the returning light is not swept across the camera's photosensor array.

[0091] In this example, the camera Cmr is connected to a processor (e.g., processing module) Proc and a display (e.g., display module, computer screen, electronic screen, etc.) Dspl, both of which may be part of the imaging system itself or may be part of separate, dedicated processing and / or display units, such as a computer system, with data passed from the camera Cmr to the computer system via a cable or computer network, including a wireless network. The display and processor may be an integrated unit. The display may be a traditional electronic display / screen or may be a touchscreen and may include a user interface for displaying information to and receiving information from an equipment operator or user. A user may interact with the display using any type of user input device known in the art, including, but not limited to, a mouse, knob, button, pointer, and touchscreen.

[0092] It may be desirable for the patient's gaze to remain fixed while imaging is being performed. One way to achieve gaze fixation is to provide a fixation target to which the patient can be instructed to gaze. The fixation target can be internal or external to the device, depending on which region of the eye is being imaged. One embodiment of an internal fixation target is shown in FIG. 11. In addition to the primary light source LtSrc used for imaging, an optional second light source FxLtSrc, such as one or more LEDs, can be positioned to image a light pattern onto the retina using a lens FxL, scanning elements FxScn, and a reflector / mirror FxM. The fixation scanner FxScn can move the position of the light pattern, and the reflector FxM directs the light pattern from the fixation scanner FxScn to the fundus F of the eye E. Preferably, the fixation scanner FxScn is positioned at the pupil plane of the system so that the light pattern on the retina / fundus can be moved according to the desired fixation position.

[0093] Slit-scanning ophthalmoscope systems can operate in different imaging modes depending on the light source and wavelength-selective filtering elements used. True-color reflectance imaging (similar to that observed by clinicians when examining the eye using a handheld or slit-lamp ophthalmoscope) can be achieved when imaging the eye using a series of colored LEDs (red, blue, and green). Images for each color can be built up stepwise with each LED turned on at each scanning position, or each color image can be captured completely separately. The three color images can be combined to display a true-color image, or displayed individually to highlight different features of the retina. The red channel best highlights the choroid, the green channel highlights the retina, and the blue channel highlights the pre-retinal layers. Additionally, specific frequencies of light (e.g., individual colored LEDs or lasers) can be used to excite different fluorophores (e.g., autofluorescence) within the eye, and the resulting fluorescence can be detected by filtering out the excitation wavelengths.

[0094] Fundus imaging systems can also provide infrared reflectance images, such as by using an infrared laser (or other infrared light source). Infrared (IR) mode is advantageous in that the eye is not sensitive to IR wavelengths. This infrared (IR) mode may allow the user to continuously capture images without obstructing the eye (e.g., in preview / alignment mode) to assist the user during device alignment. IR wavelengths also have high penetration through tissue and may provide improved visualization of choroidal structures. Additionally, fluorescein angiography (FA) and indocyanine green (ICG) angiography imaging can be achieved by collecting images after a fluorescent dye is injected into the subject's bloodstream. For example, with FA (and / or ICG), a series of time-lapse images may be captured after injecting a photoreactive dye (e.g., a fluorescent dye) into the subject's bloodstream. Note that fluorescent dyes can cause life-threatening allergic reactions in some individuals, so caution is advised. High-contrast grayscale images are captured using specific light frequencies selected to excite the dye. As the dye flows through the eye, various parts of the eye glow brightly (e.g., fluoresce), allowing the viewer to see how the dye, and therefore blood, is progressing through the eye.

[0095] Optical coherence tomography system Generally, optical coherence tomography (OCT) uses low-coherence light to generate two-dimensional (2D) and three-dimensional (3D) internal views of biological tissues. OCT enables in vivo imaging of retinal structures. OCT angiography (OCTA) generates flow information, such as vascular flow, from within the retina. Examples of OCT systems are provided in U.S. Patent Nos. 6,741,359 and 9,706,915, and examples of OCTA systems are provided in U.S. Patent Nos. 9,700,206 and 9,759,544, all of which are incorporated herein by reference in their entireties. An exemplary OCT / OCTA system is provided herein.

[0096] FIG. 23 illustrates a generalized frequency-domain optical coherence tomography (FD-OCT) system for collecting 3D image data of the eye suitable for use with the present invention. The FD-OCT system OCT_1 includes a light source LtSrc1. Typical light sources include, but are not limited to, a broadband light source with a short temporal coherence length or a swept laser source. A beam of light from the light source LtSrc1 is typically guided by an optical fiber Fbr1 to illuminate a sample, such as an eye E, a typical sample being human intraocular tissue. The light source LrSrc1 may be, for example, a broadband light source with a short temporal coherence length in the case of spectral-domain OCT (SD-OCT) or a tunable laser source in the case of swept-source OCT (SS-OCT). The light may typically be scanned using a scanner Scnr1 between the output of the optical fiber Fbr1 and the sample E, such that the beam of light (dashed line Bm) is scanned laterally across the region of the sample to be imaged. The light beam from scanner Scnr1 passes through a scan lens SL and an ophthalmic lens OL and can be focused onto the sample E to be imaged. The scan lens SL can receive the light beam from scanner Scnr1 at multiple angles of incidence to generate substantially collimated light, which the ophthalmic lens OL can then focus onto the sample. This example shows a scanning beam that needs to be scanned in two lateral directions (e.g., the x and y directions on a Cartesian plane) to scan a desired field of view (FOV). This example is point-field OCT, which uses a point-field beam to scan across the sample. Thus, scanner Scnr1 is illustratively shown to include two sub-scanners: a first sub-scanner Xscn for scanning the point-field beam across the sample in a first direction (e.g., the horizontal x direction) and a second sub-scanner Yscn for scanning the point-field beam across the sample in an intersecting second direction (e.g., the vertical y direction). If the scanning beam is a line field beam (e.g., line field OCT) and can sample an entire line portion of the sample at a time, only one scanner may be required to scan the line field beam across the sample to span the desired FOV.If the scanning beam is a full-field beam (eg, full-field OCT), a scanner may not be required and the full-field light beam may be illuminated across the entire desired FOV at once.

[0097] Regardless of the type of beam used, light scattered from the sample (e.g., sample light) is collected. In this example, scattered light returning from the sample is collected into the same optical fiber Fbr1 used to route light for illumination. A reference beam derived from the same light source LtSrc1 travels along a separate path, which in this case includes optical fiber Fbr2 and a retroreflector RR1 with an adjustable optical delay. As will be appreciated by those skilled in the art, a transmissive reference path can also be used, with an adjustable delay located in either the sample or the reference arm of the interferometer. The collected sample light is combined with the reference beam, for example, in a fiber coupler Cplr1, to form optical interference within an OCT photodetector Dtctr1 (e.g., a photodetector array, digital camera, etc.). While one fiber port is shown reaching detector Dtctr1, those skilled in the art will appreciate that various interferometer designs can be used for balanced or unbalanced detection of the interference signal. The output from detector Dtctr1 is fed to a processor (e.g., an internal or external computing device) Cmp1, which converts the observed interference into sample depth information. The depth information may be stored in a memory associated with processor Cmp1 and / or displayed on a display (e.g., computer / electronic display / screen) Scn1. The processing and storage functions may be localized within the OCT device, or the functions may be offloaded to (e.g., executed on) an external processor (e.g., an external computer system) to which the collected data is transferred. An example of a computing device (or computer system) is shown in Figure 31. This unit may be dedicated to data processing or may perform other tasks that are quite general and not dedicated to the OCT device.The processor (computing device) Cmp1 may include, for example, a field programmable gate array (FPGA), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a system on a chip (SoC), a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), or combinations thereof, which may perform some or all of the processing steps in a serial and / or parallel manner with one or more host processors and / or one or more external computing devices.

[0098] The sample and reference arms in an interferometer can be constructed of bulk optics, fiber optics, or hybrid bulk optics systems and can have different architectures, such as Michelson, Mach-Zehnder, or common-path designs, as known to those skilled in the art. Light beam, as used herein, should be interpreted as any carefully directed optical path. Instead of mechanically scanning the beam, a light field can illuminate a one-dimensional or two-dimensional area of ​​the retina to generate OCT data (e.g., U.S. Pat. No. 9,332,902; D. Hillmann et al., "Holoscopy-holographic optical coherence tomography," Optics Letters, Vol. 36(13), p. 2290, 2011; Y. Nakamura et al., "High-Speed ​​three dimensional human retinal imaging by line field spectral domain optical coherence tomography," Optics Express, 2011). Express, 15(12), p. 7103, 2007; Blazkiewicz et al., "Signal-to-noise ratio study of full-field Fourier-domain optical coherence tomography," Applied Optics, 44(36), p. 7722 (2005). In time-domain systems, the reference arm must have an adjustable optical delay to create interference. Balanced detection systems are typically used in TD-OCT and SS-OCT systems, while a spectrometer is used at the detection port for SD-OCT systems. The invention described herein can be applied to either type of OCT system.Various aspects of the present invention may be applied to any type of OCT system or to other types of ophthalmic diagnostic systems and / or to multiple ophthalmic diagnostic systems, including, but not limited to, fundus imaging systems, visual field testing devices, and scanning laser polarimeters.

[0099] In Fourier-domain optical coherence tomography (FD-OCT), each measurement is a real-valued spectrally controlled interferogram (Sj(k)). The real-valued spectral data typically undergoes several post-processing steps, including background subtraction, dispersion correction, etc. A Fourier transform of the processed interferogram yields a complex OCT signal output Aj(z) = |Aj|eiφ. The absolute value of this complex OCT signal, |Aj|, ​​reveals the scattering intensity at different path lengths and, therefore, the scattering profile with respect to depth (z-direction) within the sample. Similarly, the phase φj can also be extracted from the complex OCT signal. The scattering profile with respect to depth is called an axial scan (A-scan). A collection of A-scans measured at adjacent locations within the sample produces a cross-sectional image (tomogram or B-scan) of the sample. A collection of B-scans collected at different lateral locations on the sample constitutes a data volume or cube. For a particular data volume, the fast axis refers to the scan direction along one B-scan, and the slow axis refers to the axis along which multiple B-scans are collected. The term "cluster scan" may refer to a unit or block of data generated by repeated acquisition at the same (or substantially the same) location (or region) to analyze motion contrast, which may be used to identify blood flow. A cluster scan can consist of multiple A-scans or B-scans collected at approximately the same location on the sample at a relatively short time interval. Because the scans in a cluster scan are of the same region, stationary structures remain relatively unchanged between scans in the cluster scan, while motion contrast between scans that meet predetermined criteria may be identified as blood flow.

[0100] Various methods for generating B-scans are known in the art, including, but not limited to, along the horizontal or x-direction, along the vertical or y-direction, along the x and y diagonals, or in a circular or spiral pattern. A B-scan may be in the xz dimension, but may also be any cross-sectional image including the z-dimension. An exemplary OCT B-scan image of a normal retina of a human eye is shown in FIG. 24. An OCT B-scan of the retina provides a view of the structure of the retinal tissue. For illustrative purposes, FIG. 24 identifies the various normal retinal layers and layer boundaries. The identified retinal boundary layers include (from top to bottom) the inner limiting membrane (ILM) layer 1, the retinal nerve fiber layer (RNFL or NFL) layer 2, the ganglion cell layer (GCL) layer 3, the inner plexiform layer (IPL) layer 4, the inner nuclear layer (INL) layer 5, the outer plexiform layer (OPL) layer 6, the outer nuclear layer (ONL) layer 7, the junction between the outer segments (OS) and inner segments (IS) of photoreceptors (indicated by reference numeral 8), the external limiting membrane (ELM or OLM) layer 9, the retinal pigment epithelium (RPE) layer 10, and the Bruch's membrane (BM) layer 11.

[0101] In OCT angiography or functional OCT, analysis algorithms may be applied to OCT data collected at different times (e.g., cluster scans) at the same or nearly the same sample location on the sample to analyze motion or flow (see, e.g., U.S. Patent Application Publication Nos. 2005 / 0171438, 2012 / 0307014, 2010 / 0027857, 2012 / 0277579, and U.S. Patent No. 6,549,801, all of which are incorporated by reference in their entireties). OCT systems may use any one of a number of OCT angiography processing algorithms (e.g., motion contrast algorithms) to identify blood flow. For example, motion contrast algorithms can be applied to intensity information derived from the image data (intensity-based algorithms), phase information from the image data (phase-based algorithms), or complex image data (complex-based algorithms). An en face image is a 2D projection of the 3D OCT data (e.g., by averaging the intensity of each individual A-scan, whereby each A-scan defines a pixel in the 2D projection). Similarly, an en face vascular image is an image that displays motion contrast signals in which the data dimension corresponding to depth (e.g., the z-direction along the A-scan) is displayed as a single representative value (e.g., a pixel in the 2D projection image), typically by summing or integrating all or isolated portions of the data (see, e.g., U.S. Pat. No. 7,301,644, incorporated herein by reference in its entirety). OCT systems that provide angiography capabilities may be referred to as OCT angiography (OCTA) systems.

[0102] FIG. 25 shows an example of an en face vasculature image. After processing the data and highlighting motion contrast using any of the motion contrast methods known in the art, an en face (e.g., front view) image of the vasculature may be generated by summing pixel ranges corresponding to a tissue depth from the surface of the retinal internal limiting membrane (ILM). FIG. 26 shows an exemplary B-scan of a vasculature (OCTA) image. As shown, structural information may be less clear because blood flow traverses multiple retinal layers, obscuring them more than in a structural OCT B-scan such as that shown in FIG. 24. Nevertheless, OCTA provides a noninvasive technique for imaging the retinal and choroidal microvasculature, which may be important for diagnosing and / or monitoring various pathologies. For example, OCTA may be used to identify diabetic retinopathy by identifying microaneurysms, neovascular complexes, and quantifying the foveal avascular zone and nonperfused areas. Furthermore, OCTA has been shown to show good agreement with fluorescein angiography (FA), a more traditional but less invasive technique that requires the injection of dye to observe vascular flow in the retina. Furthermore, in dry age-related macular degeneration (AMD), OCTA has been used to monitor the overall decrease in choriocapillaris flow. Similarly, in exudative AMD, OCTA can provide qualitative and quantitative analysis of choroidal neovascular membranes. OCTA has also been used to study vascular obstruction, for example, to assess nonperfused areas and the integrity of the superficial and deep plexuses.

[0103] Neural Networks As mentioned above, the present invention may use neural network (NN) machine learning (ML) models. For completeness, neural networks are generally described herein. The invention may use any of the following neural network architectures, alone or in combination: A neural network, or neural net, is a network of interconnected neurons (via nodes), with each neuron representing a node in the network. Collections of neurons may be arranged in layers, with the output of one layer being fed forward to the next layer in a multi-layer perceptron (MLP) arrangement. An MLP may be understood as a feed-forward neural network that maps a set of input data to a set of output data.

[0104] FIG. 25 illustrates an example of a multilayer perceptron (MLP) neural network. The structure may include multiple hidden (e.g., inner) layers HL1 through HLn, which map an input layer InL (which receives a set of inputs (or vector inputs) in_1 through in_3) to an output layer OutL, which generates a set of outputs (or vector outputs), e.g., out_1 and out_2. Each layer may have any number of nodes, which are shown here as circles within each layer for illustrative purposes. In this example, the first hidden layer HL1 has two nodes, and hidden layers HL2, HL3, and HLn each have three nodes. Generally, the deeper the MLP (e.g., the greater the number of hidden layers in the MLP), the greater its learning capacity. The input layer InL may receive vector inputs (shown for illustrative purposes as a three-dimensional vector consisting of in_1, in_2, and in_3) and feed the received vector inputs to the first hidden layer HL1 in the sequence of hidden layers. The output layer OutL receives the output from the last hidden layer in the multi-layer model, say HLn, and produces a vector output result (shown for illustration purposes as a two-dimensional vector consisting of out_1 and out_2).

[0105] Typically, each neuron (i.e., node) generates one output, which is fed forward to neurons in the immediately succeeding layer. However, each neuron in a hidden layer may receive multiple inputs, either from the input layer or from the outputs of neurons in the immediately preceding hidden layer. In general, each node may apply a function to its inputs to generate the output for that node. Nodes in a hidden layer (e.g., the training layer) may apply the same function to each of their inputs to generate their respective outputs. However, some nodes, e.g., nodes in the input layer InL, may receive only one input and be passive, meaning that they simply relay the value of that one input to their output, e.g., they provide a copy of that input to their output; this is indicated by the dashed arrows in the nodes of the input layer InL for illustrative purposes.

[0106] For illustrative purposes, Figure 28 shows a simplified neural network consisting of an input layer InL', a hidden layer HL1', and an output layer OutL'. The input layer InL' is shown to have two input nodes i1 and i2, which receive inputs Input_1 and Input_2, respectively (e.g., the input nodes of layer InL' receive a two-dimensional input vector). The input layer InL' feeds forward into one hidden layer HL1' with two nodes h1 and h2, which in turn feeds forward into an output layer OutL' with two nodes o1 and o2. The interconnections, or links, between neurons (shown with solid arrows for illustrative purposes) have weights w1 through w8. Typically, except for the input layer, a node (neuron) may receive as input the output of a node in the previous layer. Each node may calculate its output by multiplying each of its inputs by each input's corresponding interconnection weight, summing the products of the inputs, adding (or multiplying) a constant defined by other weights or biases that may be associated with that particular node (e.g., node weights w9, w10, w11, and w12 corresponding to nodes h1, h2, o1, and o2, respectively), and then applying a nonlinear or logarithmic function to the result. The nonlinear function may be referred to as an activation function or a transfer function. Multiple activation functions are known in the art, and the selection of a particular activation function is not important to this description. However, it should be noted that the operation of an ML model, the behavior of a neural net, depends on the values ​​of the weights, which the neural network may be trained to provide a desired output for a given input.

[0107] During a training, or learning, phase, a neural network learns (e.g., is trained to identify) appropriate weight values ​​to achieve a desired output for a given input. Before a neural network is trained, each weight may be individually assigned an initial (e.g., random, optionally non-zero) value, such as a random number seed. Various methods for assigning initial weights are known in the art. The weights are then trained (optimized) so that, for a given training vector input, the neural network produces an output that approximates a desired (predetermined) training vector output. For example, the weights may be gradually adjusted over thousands of iterative cycles by a method called backpropagation. In each backpropagation cycle, a training input (e.g., a vector input or training input image / sample) is passed forward through the neural network to provide its actual output (e.g., a vector output). The error of each output neuron, or output node, is then calculated based on the actual neuron's output and the supervised training output for that neuron (e.g., a training output image / sample corresponding to the current training input image / sample). It then propagates backward through the neural network (from the output layer back to the input layer), updating the weights based on how much influence each weight has on the overall error, thereby moving the neural network's output closer to the desired training output. This cycle is then repeated until the neural network's actual output is within an acceptable error range of the desired training output for that training input. As will be appreciated, each training input may require many backpropagation iterations to achieve the desired error range. Typically, an epoch refers to one backpropagation iteration (e.g., one forward pass and one backward pass) of all training samples, and training a neural network may require many epochs. Generally, the larger the training set, the better the performance of the trained ML model; therefore, various data augmentation methods may be used to increase the size of the training set.For example, if the training set includes pairs of corresponding training input images and training output images, the training images may be divided into multiple corresponding image segments (or patches). Corresponding patches from the training input images and training output images may be paired to define multiple training patch pairs from one input / output image pair, thereby expanding the training set. However, training a large training set increases the demands on computer resources, such as memory and data processing resources. The computational demands may be reduced by dividing the large training set into multiple mini-batches, the size of which determines the number of training samples in one forward / backward pass. In this case, one epoch may contain multiple mini-batches. Another problem is the possibility that the neural network may overfit the training set, reducing its ability to generalize from a specific input to different inputs. The overfitting problem may be mitigated by creating an ensemble of neural networks or by randomly dropping out nodes in the neural network during training, which effectively removes the dropped leads from the neural network. Various dropout adjustment methods, such as inverse dropout, are known in the art.

[0108] It should be noted that the operation of a trained NN machine model is not a simple algorithm of computation / analysis steps. Indeed, when a trained NN machine model receives an input, the input is not analyzed in the traditional sense. Rather, regardless of the purpose or nature of the input (e.g., vectors defining a live image / scan or vectors defining any other entity such as a demographic description or activity record), the input is subjected to the same architectural construction of the trained neural network (e.g., the same node / layer arrangement, trained weights and bias values, predetermined convolution / deconvolution operations, activation functions, pooling operations, etc.), and it may not be obvious how the architectural construction of the trained network generates its output. Furthermore, the values ​​of the trained weights and biases are not deterministic and depend on many factors, such as the amount of time given to the neural network for training (e.g., the number of epochs in training), the random starting values ​​of the weights before training begins, the computer architecture of the machine on which the NN is trained, the selection of training samples, the distribution of training samples among multiple mini-batches, the selection of activation functions, the selection of error functions that modify the weights, and even whether training is interrupted on one machine (e.g., with a first computer architecture) and completed on another machine (e.g., with a different computer architecture). The point is that the reasons why a trained ML model arrives at a particular output are not clear, and much research is currently being conducted to identify the factors on which ML models base their output. Therefore, the processing of neural networks on live data cannot be reduced to a simple algorithmic step. Rather, the operation depends on the training architecture, training sample set, training sequence, and various circumstances in the training of the ML model.

[0109] In general, constructing a neural network machine learning model may include a learning (or training) stage and a classification (or computation) stage. In the learning stage, a neural network may be trained for a specific purpose and provided with a set of training examples, including training (sample) inputs and training (sample) outputs, and optionally a set of validation examples for testing the training progress. During this learning process, various weights associated with nodes and node interconnections within the neural network are gradually adjusted to reduce the error between the neural network's actual output and the desired training output. In this way, a multi-layer feedforward neural network (such as those described above) may be able to approximate any measurable function to any desired accuracy. The result of the learning stage is a learned (e.g., trained) (neural network) machine learning (ML). In the computation stage, a set of test inputs (or live inputs) may be provided to the learned (trained) ML model, which may apply what it has learned to generate output predictions based on the test inputs.

[0110] Similar to the conventional neural networks of Figures 27 and 28, convolutional neural networks (CNNs) are also composed of neurons with learnable weights and biases. Each neuron receives an input and performs an operation (e.g., a dot product), optionally followed by a nonlinear transformation. However, CNNs receive raw image pixels at one end (e.g., the input) and provide a classification (or class) score at the other end (e.g., the output). Because CNNs expect images as input, they are optimized to handle volumes (e.g., image pixel height and width, and image depth, e.g., color depth, such as RGB depth defined by three colors: red, green, and blue). For example, CNN layers may be optimized for neurons arranged in three dimensions. Neurons in a CNN layer may connect to a small region of the previous layer rather than all of the neurons in a fully connected NN. The final output layer of a CNN may reduce the full image to a single vector (classification) arranged along the depth dimension.

[0111] FIG. 29 provides an exemplary convolutional neural network architecture. A convolutional neural network may be defined as a sequence of two or more layers (e.g., Layer 1 through Layer N), where a layer may include an (image) convolution step, a (resulting) weighted sum step, and a nonlinear function step. The convolution may be performed on the input data by, for example, applying a filter (or kernel) on a moving window over the input data to generate a feature map. Each layer and layer component may have different predetermined filters (from a filter bank), weights (or weighting parameters), and / or function parameters. In this example, the input data may be an image of a certain pixel height and width, or the raw pixel values ​​of this image. In this example, the input image is depicted as having a depth of three color channels, RGB (red, green, blue). Optionally, various preprocessing steps may be performed on the input image, and the results of the preprocessing steps may be input instead of or in addition to the raw image data. Some examples of image processing may include retinal vessel map segmentation, color space conversion, adaptive histogram equalization, connected component generation, etc. Within a layer, a dot product may be calculated between certain weights and their connected small regions within the input volume. While many methods for constructing CNNs are known in the art, by way of example, layers may be configured to apply element-wise activation functions, such as a max(0,x) threshold at zero. Pooling functions may be performed (e.g., along the x and y directions) to downsample the volume. Fully connected layers may be used to identify classification outputs and generate one-dimensional output vectors, which have proven useful in image recognition and classification. However, for image segmentation, CNNs must classify each pixel. Because each CNN layer tends to reduce the resolution of the input image, another stage is required to upsample the image to its original resolution. This may be achieved by applying a transposed convolution (or deconvolution) stage TC, which typically does not use any predetermined interpolation method but instead has learnable parameters.

[0112] Convolutional neural networks have been successfully applied to many problems in computer vision. As mentioned above, training a CNN generally requires a large training dataset. The U-Net architecture is based on a CNN and can generally be trained with a smaller training dataset than traditional CNNs.

[0113] FIG. 30 illustrates an exemplary U-Net architecture. This exemplary U-Net includes an input module (or input layer or stage), which receives an input U-in (e.g., an input image or image patch) of any size. For convenience, the image size at any stage or layer is indicated within a box representing the image; for example, in the input module, the number "128x128" is enclosed, indicating that the input image U-in is composed of 128x128 pixels. The input image may be a fundus image, an OCT / OCTA en face image, a B-scan image, etc. However, it should be understood that the input may be of any size or dimensionality. For example, the input image may be an RGB color image, a monochrome image, a volumetric image, etc. The input image passes through a series of processing layers, each of which is illustrated with exemplary sizes, but these sizes are for illustrative purposes only and will depend, for example, on the size of the image, the convolutional filters, and / or the pooling stage. This architecture consists of a convergent path (exemplary herein including four encoding modules) followed by an augmented path (exemplary herein including four decoding modules), with copy-and-crop links (e.g., CC1-CC4) between corresponding modules / stages that copy the output of one encoding module in the convergent path and connect (e.g., append) it to the upconverted input of the corresponding decoding module in the augmented path. This results in a characteristic U-shape, from which the architecture is named. Optionally, for computational considerations, a "bottleneck" module / stage (BN) can be placed between the convergent path and the augmented path. The bottleneck BN may consist of two convolutional layers (with batch normalization and optional dropout).

[0114] The convergent path is similar to an encoder and typically uses feature maps to capture context (or feature) information. In this example, each encoding module in the convergent path includes two or more convolutional layers, indicated by an asterisk symbol "*," which may be followed by a max-pooling layer (e.g., a downsampling layer). For example, the input image U-in is shown passing through two convolutional layers, each with 32 feature maps. It can be understood that each convolutional kernel produces a feature map (e.g., the output from a convolution operation with a given kernel is an image commonly referred to as a "feature map"). For example, the input U-in passes through a first convolution that applies 32 convolutional kernels (not shown), producing an output consisting of 32 individual feature maps. However, as is known in the art, the number of feature maps produced by a convolution operation can be adjusted (upward or downward). For example, the number of feature maps can be reduced by averaging groups of feature maps, deleting some feature maps, or other known methods of reducing feature maps. In this example, this first convolution is followed by a second convolution whose output is limited to 32 feature maps. Another way to envision the feature maps is to consider the output of the convolution layer as a 3D image whose 2D dimensions are given by the stated XY plane pixel dimensions (e.g., 128x128 pixels) and whose depth is given by the number of feature maps (e.g., the depth of the 32 plane images). Following this illustration, the output of the second convolution (e.g., the output of the first encoding module in the convergence path) can be described as a 128x128x32 image. The output from the second convolution is then subjected to a pooling operation, which reduces the 2D dimensions of each feature map (e.g., the X and Y dimensions can each be reduced by half). The pooling operation can be embodied within a downsampling process, as indicated by the downward arrows. Several pooling methods, such as max pooling, are known in the art, and the particular pooling method is not critical to the present invention.The number of feature maps doubles with each pooling, such as 32 feature maps in the first encoding module (or block), 64 feature maps in the second encoding module, and so on. Thus, the convergence path forms a convolutional network composed of multiple encoding modules (or stages or blocks). As is typical for convolutional networks, each encoding module may provide at least one convolution stage followed by an activation function (e.g., a rectified linear unit (ReLU) or sigmoid layer) (not shown) and a max-pooling operation. Generally, the activation function introduces nonlinearity into the layer (e.g., to avoid overfitting), receives the layer's results, and determines whether to "activate" the output (e.g., whether the value at a particular node meets a predetermined criterion to forward the output to the next layer / node). In summary, the convergence path generally reduces spatial information and increases feature information.

[0115] The extension path is similar to the decoder, notably providing localization and spatial information to the results of the convergence path, despite the downsampling and any max-pooling performed in the contraction stage. The extension path includes multiple decoding modules, each of which combines its current upconverted input with the output of a corresponding encoding module. Thus, features and spatial information are combined in the extension path through a series of upconvolutions (e.g., upsampling or transposed convolutions, i.e., deconvolutions) and combinations (e.g., via CC1-CC4) with high-resolution features from the convergence path. Thus, the output of the deconvolution layer is combined with the corresponding (optionally cropped) feature map from the convergence path, followed by two convolutional layers and activation functions (optionally batch normalized).

[0116] The output from the last augmentation module in the augmentation path may be fed to another processing / training block or layer, such as a classifier block, which may be trained with the U-Net architecture. Alternatively, or additionally, the output of the last upsampling block (at the end of the augmentation path) may be provided to another convolution (e.g., output convolution) operation, as indicated by the dotted arrow, before generating its output U-out. The kernel size of the output convolution may be selected to reduce the dimensions of the last upsampling block to a desired size. For example, a neural network may have multiple features per pixel just before reaching the output convolution, which may provide a 1×1 convolution operation to combine these multiple features into a single output value per pixel at the per-pixel level.

[0117] Computing Devices / Systems FIG. 31 illustrates an exemplary computer system (or computing device). In some embodiments, one or more computer systems may provide functionality described or illustrated herein and / or perform one or more steps of one or more methods described or illustrated herein. The computer system may take any suitable physical form. For example, the computer system may be an embedded computer system, a system-on-chip (SOC), or a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, a mesh of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the computer system may reside in a cloud, which may include one or more cloud components within one or more networks.

[0118] In some embodiments, a computer system may include a processor Cpnt1, a memory Cpnt2, a storage Cpnt3, an input / output (I / O) interface Cpnt4, a communication interface Cpnt5, and a bus Cpnt6. The computer system may also optionally include a display Cpnt7, such as a computer monitor or screen.

[0119] The processor Cpnt1 includes hardware for executing instructions, such as those that constitute a computer program. For example, the processor Cpnt1 may be a central processing unit (CPU) or a general-purpose computing-on-graphics processing unit (GPGPU). The processor Cpnt1 may read (or fetch) instructions from an internal register, an internal cache, memory Cpnt2, or storage Cpnt3, decode and execute the instructions, and write one or more results to the internal register, the internal cache, memory Cpnt2, or storage Cpnt3. In particular embodiments, the processor Cpnt1 may include one or more internal caches for data, instructions, or addresses. The processor Cpnt1 may include one or more instruction caches and one or more data caches, for example, to hold data tables. Instructions in the instruction caches may be copies of instructions in memory Cpnt2 or storage Cpnt3, and the instruction caches may speed up retrieval of these instructions by the processor Cpnt1. Processor Cpnt1 may include any suitable number of internal registers and may include one or more arithmetic logic units (ALUs). Processor Cpnt1 may be a multi-core processor or may include one or more processors Cpnt1. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0120] The memory Cpnt2 may include a main memory that stores instructions for the processor Cpnt1 to execute processing or to hold intermediate data during processing. For example, the computer system may load instructions or data (e.g., a data table) from the storage Cpnt3 or from other sources (e.g., another computer system) into the memory Cpnt2. The processor Cpnt1 may load instructions and data from the memory Cpnt2 into one or more internal registers or internal caches. To execute an instruction, the processor Cpnt1 may read and decode the instruction from the internal register or internal cache. During or after the execution of an instruction, the processor Cpnt1 may write one or more results (which may be intermediate or final results) to an internal register, an internal cache, the memory Cpnt2, or the storage Cpnt3. The bus Cpnt6 may include one or more memory buses (each of which may include an ADDRESS bus and a DATA bus) and may couple the processor Cpnt1 to the memory Cpnt2 and / or the storage Cpnt3. Optionally, one or more memory management units (MMUs) facilitate data transfer between the processor Cpnt1 and the memory Cpnt2. The memory Cpnt2 (which may be a high-speed volatile memory) may include random access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM). The storage Cpnt3 may include long-term or high-capacity storage for data or instructions. The storage Cpnt3 may be internal or external to the computer system and may include one or more of a disk drive (e.g., a hard disk drive (HDD) or a solid-state drive (SSD)), flash memory, ROM, EPROM, optical disk, magneto-optical disk, magnetic tape, a universal serial bus (USB)-accessible drive, or other type of non-volatile memory.

[0121] The I / O interface Cpnt4 may be software, hardware, or a combination of both, and may include one or more interfaces (e.g., serial or parallel communication ports) for communicating with I / O devices, which may enable communication with a human (e.g., a user). For example, the I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, table, touch screen, trackball, video camera, other suitable I / O device, or a combination of two or more thereof.

[0122] The communication interface Cpnt5 may provide a network interface for communicating with other systems or networks. The communication interface Cpnt5 may include a Bluetooth interface or other types of packet-based communication. For example, the communication interface Cpnt5 may include a network interface controller (NIC) and / or a wireless NIC or wireless adapter for communication with a wireless network. The communication interface Cpnt5 may provide communication with a Wi-Fi network, an ad-hoc network, a personal area network (PAN), a wireless PAN (e.g., Bluetooth WPAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a cellular network (e.g., a Global System for Mobile Communications (GSM) network), the Internet, or a combination of two or more thereof.

[0123] Bus Cpnt6 may provide a communication link between the above-mentioned components of the computing system. For example, bus Cpnt6 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand bus, a low-pin-count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or any other suitable bus, or a combination of two or more thereof.

[0124] Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0125] As used herein, a computer-readable non-transitory storage medium may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, or any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more thereof, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.

[0126] While the present invention has been described in conjunction with several specific embodiments, as will be apparent to those skilled in the art in light of the foregoing description, many other alternatives, modifications, and variations will be apparent. Accordingly, the invention as described herein is intended to embrace all such alternatives, modifications, applications, and variations that may fall within the spirit and scope of the appended claims.

Claims

1. 1. A method for eye tracking, comprising: capturing a plurality of images of the eye, including a reference image and one or more live images; defining a reference anchor point in the reference image; defining one or more auxiliary points in the reference image; In the live image, select a) identifying an initial matching point that matches the reference anchor point; b) searching for a match of the selection assist point within a region based on a position of the selection assist point relative to the reference anchor point; and correcting tracking errors between the reference image and the selected live image based on those matched points.

2. A plurality of auxiliary points are defined in the reference image; searching for a match of the selected auxiliary point is part of searching for a match of a plurality of auxiliary points in the selected live image; The method of claim 1 , wherein in response to a number of matched auxiliary points in the selected live image not being greater than a predetermined minimum value, the selected live image is not corrected for tracking error.

3. The method of claim 2 , wherein the predetermined minimum value is greater than half of the plurality of auxiliary points.

4. The method of claim 1 , wherein the reference image and the live image are infrared images.

5. The method of claim 1 , wherein the captured images are of a retina of an eye.

6. further comprising identifying salient physical features within the reference image; The method of claim 1 , wherein the reference anchor points are defined based on the prominent physical features.

7. the fiducial anchor point is part of a fiducial anchor template consisting of a plurality of identifiers that together define the salient physical feature; The method of claim 6 , wherein identifying the initial matching point is part of identifying an initial matching template that matches the reference anchor template.

8. the one or more auxiliary points are part of one or more respective auxiliary templates comprising a plurality of identifiers that together define respective auxiliary physical features in the reference image; The method of claim 7 , wherein searching for a match of the selection assist point is part of searching for a match of a corresponding selection assist template within a region based on an offset position of the selection assist template relative to the reference anchor template.

9. The method of claim 8 , wherein the salient physical features in the reference image and the one or more live images are identified through the use of a neural network.

10. The method of claim 9 , wherein the prominent physical feature is a predetermined retinal structure.

11. 11. The method of claim 10, wherein the prominent physical feature is the optic nerve disc (or optic nerve head, ONH), a lesion, or a particular vascular pattern.

12. The method of claim 6 , wherein the prominent physical feature is the optic nerve head, the pupil, the iris border, or the center of the eye.

13. identifying a plurality of candidate anchor points in the reference image; searching for matches of the plurality of candidate anchor points in the selected live image; The method of claim 1 , further comprising the step of: designating the candidate anchor point with the best match in the selected live image as the reference anchor point.

14. The method of claim 13 , wherein the best matching candidate anchor point is the candidate that has the highest confidence in matching within the selected live image.

15. The method of claim 13 , wherein the best matching candidate anchor point is the candidate for which a match is found most quickly.

16. searching for matches of the candidate anchor points in the live images; 16. The method of claim 13, further comprising the step of designating the candidate anchor point with the best match in the plurality of selected live images as the reference anchor point.

17. The method of claim 16, wherein the best matching candidate anchor point is the candidate anchor point for which a match is found most frequently in a series of consecutive selected live images.

18. The method of claim 17 , wherein the series of images is a predetermined number of live images.

19. 19. The method of any one of claims 13 to 18, wherein the candidate anchor points are identified based on their saliency in the reference image.

20. 20. The method of claim 13, wherein candidate anchor points that are not designated as reference anchor points are designated auxiliary points.

21. defining a plurality of said reference anchor points in said reference image; Within the selected live image, a) identifying a plurality of initial matching points that correspond to the plurality of reference anchor points, and transforming the live image into the reference image as a coarse registration based on the identified plurality of reference anchor points; The method of claim 1 , further comprising: b) searching for a match of the selected auxiliary point within a region based on a position of the selected auxiliary point relative to the plurality of reference anchor points.

22. further comprising defining an OCT acquisition field of view (hereinafter FOV) on the eye using the OCT system; the reference anchor point is defined within a tracking FOV that is movable within the reference image; The method of claim 1 , wherein the tracking FOV is moved around the reference image to a position determined to be optimal for image tracking while at least partially overlapping with the OCT FOV.

23. 23. The method of claim 22, wherein the optimal position is based on tracking algorithm outputs including one or more of tracking error, landmark distribution, and number of landmarks.

24. 1. A method for eye tracking, comprising: capturing a plurality of images of a retina of the eye, including a reference image and one or more live images; identifying salient physical features in the reference image; defining a reference anchor template based on the salient physical features; defining one or more auxiliary templates based on other physical features in the reference image; storing the position of the auxiliary template relative to the reference anchor template; Within each live image, a) identifying an initial matching region that matches the reference anchor template, the initial matching region defining a corresponding live anchor template whose position matches a position of the reference anchor template; b) searching for matches of the one or more auxiliary templates, each found match defining another corresponding template in the live image, the search for each auxiliary template being limited to a bounded region whose position relative to the live anchor template is based on the stored position of the auxiliary template relative to the reference anchor template; and correcting a tracking error between the reference image and the selected live image based on two or more corresponding templates of the reference image and the selected live image.

25. 25. The method of claim 24, wherein the salient physical features in the reference image and each live image are identified through the use of a neural network.