Gaze Tracking Using an Integrated Calibration Process

The system provides automatic eye-tracking calibration using video analysis to track gaze points on everyday devices, addressing the limitations of controlled laboratory settings by reducing costs and improving accessibility for neurological assessments.

JP2026502269APending Publication Date: 2026-01-21NEUROLIGHT LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025539809
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-01-03
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Existing eye-tracking technologies require expensive and time-consuming controlled laboratory settings for accurate gaze measurement, limiting accessibility and usability for patients due to high setup costs and limited availability of such settings.

Method used

A system for automatic eye-tracking calibration that uses video analysis to determine eye features and gaze points, eliminating the need for pre-calibration procedures and allowing gaze tracking on ubiquitous devices like smartphones and tablets, adjusting for changes in position and orientation during test sequences.

Benefits of technology

Enables accurate gaze tracking on everyday devices, reducing setup costs and increasing accessibility for neurological and mental health condition assessments, while maintaining high accuracy by recalculating mappings during test sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502269000001_ABST
    Figure 2026502269000001_ABST
Patent Text Reader

Abstract

Disclosed herein are aspects of systems, methods, and / or computer program products, and / or combinations and subcombinations thereof, for tracking a subject's gaze in a manner incorporating automatic eye-tracking calibration. In one aspect, a video of a subject viewing content presented by a display is acquired. The video is analyzed to determine a series of time points during which the subject's gaze remains on a visual target, and to determine a set of eye features for each time point in the series. The location of a stimulus presented by the display is also acquired for each time point in the series. This information is utilized to determine a mapping that associates the set of eye feature values ​​with the gaze point. The mapping is used to associate the set of eye features acquired by analyzing the video with the subject's gaze point. TIFF2026502269000003.tif166140
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 18 / 397,535, filed December 27, 2023, and U.S. Provisional Patent Application No. 63 / 437,136, filed January 5, 2023, the contents of each of which are incorporated herein by reference in their entirety.

[0002] background Field The present disclosure is directed generally to techniques for tracking a subject's gaze, and more particularly to techniques for tracking a subject's gaze that rely on or otherwise benefit from eye-tracking calibration. [Background technology]

[0003] background Eye tracking can be defined as the process of measuring either a subject's point of gaze (where the subject is looking) or the angle of the eye's gaze relative to the subject's position. A significant number of studies have used eye tracking techniques to correlate eye movements with various medical conditions, including neurological conditions.

[0004] Testing medical conditions using eye-tracking technology may involve measuring a subject's gaze over time as accurately as possible for directional oculometry tasks, which may include saccades, smooth pursuit, prolonged fixations, etc. To conduct these types of tests with sufficient accuracy, eye movements are typically measured in a well-controlled laboratory setting (e.g., head restraints, controlled ambient light, or other such parameters) using specialized devices (e.g., infrared eye trackers, pupillometers, or other such devices). However, setting up and maintaining such a controlled testing environment can be very expensive and require significant time and effort. Furthermore, such well-controlled laboratory settings exist in limited numbers, making it difficult for patients to make an appointment and / or travel to one. Summary of the Invention

[0005] overview This specification provides aspects of a system, apparatus, article of manufacture, method, and / or computer program product, and / or combinations and subcombinations thereof, for tracking a subject's gaze point in a manner incorporating automatic eye tracking calibration. One exemplary aspect includes: acquiring a video of at least a portion of a subject's face while the subject is viewing content presented in a display area of ​​a display; analyzing the video to determine a series of time points at which the subject's gaze remains on a visual target, and determining a set of eye features for at least each time point in the series of time points; acquiring a position of a visual stimulus in the display area of ​​the display for each time point in the series of time points; and determining a mapping that associates the set of eye features with the gaze points, the determination being based at least on (i) the set of eye features for each time point in the series of time points, and (ii) the position of the visual stimulus in the display area of ​​the display for each time point in the series of time points; and using the mapping to associate one or more sets of eye features acquired by analyzing the video with one or more gaze points of the subject, respectively.

[0006] The one or more sets of eye features obtained by analyzing the video may include one or more sets of right eye features, and the one or more gaze points of the subject may include one or more gaze points of the subject's right eye. Alternatively, the one or more sets of eye features obtained by analyzing the video may include one or more sets of left eye features, and the one or more gaze points of the subject may include one or more gaze points of the subject's left eye.

[0007] This exemplary embodiment may further operate to detect a neurological or mental health condition of the subject based at least on one or more gaze points of the subject.

[0008] Capturing video of at least a portion of the subject's face while the subject views content presented within the viewing area of ​​the display may include capturing video of at least a portion of the subject's face while the subject performs a directional oculometry task in response to visual stimuli presented within the viewing area of ​​the display. The directional oculometry task may include, for example, one of a saccade test, a smooth pursuit test, or a long fixation test.

[0009] Each set of ocular features may include one or more of the following: pupil position relative to the eye opening, canthus position of the eye, or length and orientation of the minor and major elliptical axes of the eye's iris.

[0010] Determining the mapping may include performing a regression analysis to determine the mapping. Performing the regression analysis may include one of performing a linear regression analysis, performing a polynomial regression analysis, or performing a decision tree regression analysis. Performing the polynomial regression analysis may include performing a response surface analysis.

[0011] Analyzing the video to determine a set of eye features for each of the time points in the series of time points may include extracting eye images from frames of the video corresponding to the time points in the series of time points, and providing the eye images to a neural network that outputs an estimated gaze point for the time point in the series of time points based at least on the eye images.

[0012] This exemplary embodiment may further operate to: analyze the video to determine one or more additional time points at which the subject's gaze remains on the visual target and to determine a set of eye features for each of the one or more additional time points; obtain a position of the visual stimulus within the display area for each of the one or more additional time points; and recalculate the mapping based at least on (i) the set of eye features for each of the one or more additional time points and (ii) the position of the visual stimulus within the display area for each of the one or more additional time points. [Brief explanation of the drawings]

[0013] The accompanying drawings are incorporated into and form a part of this specification.

[0014] [Figure 1] 1A and 1B show the viewing area of ​​a display and a fixation point that is presented sequentially at different locations within the viewing area as part of a saccade test. [Figure 2] FIG. 2 illustrates an example system for tracking a subject's gaze in a manner that incorporates automatic eye-tracking calibration, according to some embodiments. [Figure 3] FIG. 3 shows an image of the eye with markings superimposed to demonstrate one method of measuring pupil position relative to the eye opening. [Figure 4] Figures 4A and 4B show the display area of ​​the display and a fixation point that is presented sequentially at different locations within the display area as part of a saccade test, along with the (x,y) coordinates representing the location of the fixation point within the display area. [Figure 5] FIG. 5 illustrates a neural network that may be used to obtain estimated right and left eye coordinates based on at least right and left eye images, according to some embodiments. [Figure 6] FIG. 6 is a flow diagram of a method for tracking a subject's gaze in a manner that incorporates automatic eye-tracking calibration, according to some embodiments. [Figure 7] FIG. 7 is a flow diagram of a method for recalculating a mapping that associates a set of ocular features with a point of regard, according to some embodiments. [Figure 8] FIG. 8 illustrates an example of a computer system useful for implementing various aspects.

[0015] In the drawings, like reference numbers generally indicate the same or similar elements. Further, the left-most digit(s) of a reference number generally identifies the drawing in which that reference number first appears. DETAILED DESCRIPTION OF THE INVENTION

[0016] Overview As mentioned in the background section above, eye tracking can be defined as the process of measuring either a subject's point of gaze (where the subject is looking) or the angle of the eye's gaze relative to the subject's position. A significant number of studies have used eye tracking techniques to correlate eye movements with various medical conditions, including neurological conditions.

[0017] Testing medical conditions using eye-tracking technology can involve measuring a subject's gaze over time as accurately as possible for directional oculometry tasks, which may include saccades, smooth pursuit, and prolonged fixations. For example, a saccade is a rapid movement of the eyes between two fixation points. Saccade testing measures a subject's ability to move their eyes from one fixation point to another in a single, rapid movement. For ease of illustration, FIGS. 1A and 1B show a display area 100 and a fixation point 102 (e.g., a dot) that is sequentially presented at different locations within the display area 100 as part of a saccade test. To conduct the test, the subject can be instructed to gaze directly at the fixation point 102 within the display area 100 as the fixation point 102 momentarily switches from a first location 104 (FIG. 1A) to a second location 106 (FIG. 1B) within the display area 100. A complete test sequence typically involves sequentially rendering multiple similar images onto the display area 100, with the fixation point 102 appearing at a different location in each image, with all the different locations spanning a large portion of the display area 100. Throughout the test sequence, the subject's eye movements are captured to enable measurements.

[0018] To conduct these types of tests with sufficient accuracy, eye movements are typically measured in a well-controlled laboratory setting (e.g., head restraint, controlled ambient light, or other such parameters) using specialized devices (e.g., infrared eye trackers, pupillometers, or other such devices). However, setting up and maintaining such a controlled testing environment can be very expensive and require significant time and effort. Furthermore, such well-controlled laboratory settings exist in limited numbers, making it difficult for patients to schedule and / or travel to them.

[0019] To make the benefits of the aforementioned research more widely available for the detection and treatment of neurological conditions, it is desirable to make eye tracking very widely available at low cost, for example, by using ubiquitous cameras on smartphones, tablets, laptop computers, desktop computers, etc. to observe eye behavior in response to visual stimuli.

[0020] The features measured in eye images from camera video do not indicate absolute gaze direction, only changes in gaze direction. To accurately determine what a subject is looking at, some kind of pre-calibration procedure is usually performed in which the subject views a series of fixation points spanning the display's viewing area while an eye tracker records the eye features corresponding to each gaze position. A generalized mapping of eye features to points within the viewing area can then be calculated from the calibration measurements. The calibration mapping must accurately capture the appearance and behavior details of each specific subject, as well as the specifications and geometry of the system used for the measurements. The calibration mapping remains accurate as long as the relative positions and orientations of the subject, display, and camera remain essentially constant. If any of these relative positions and orientations change significantly after pre-calibration is completed, the gaze direction calculated using the calibration mapping will be rendered inaccurate.

[0021] The aforementioned pre-calibration procedures can be cumbersome, uncomfortable, or very difficult to complete. This is especially true for subjects for whom keeping their head still for extended periods of time is very difficult or even impossible due to medical conditions, young age, etc. Furthermore, when using handheld devices such as smartphones or tablets for eye tracking, frequent changes in the relative position and orientation of the handheld device and the subject can necessitate frequent recalibration.

[0022] Gaze Tracking Using an Integrated Calibration Process The embodiments described herein address some or all of the aforementioned problems by tracking a subject's gaze point in a manner that incorporates automatic eye tracking calibration. In particular, as described herein, the embodiments may acquire a video of a subject while the subject is viewing content presented within a display area of ​​a display, such as while the subject is performing a directional oculometry task in response to visual stimuli presented within the display area, analyze the video to determine a series of time points at which the subject's gaze remains on a visual target and to determine a set of eye features for at least each time point in the series of time points, acquire a position of the visual stimulus within the display area for each time point in the series of time points, determine a mapping that associates the set of eye features with the gaze points based at least on (i) the set of eye features for each time point in the series of time points and (ii) the position of the visual stimulus within the display area for each time point, and use the mapping to map one or more sets of eye features acquired by analyzing the video to one or more gaze points of the subject, respectively.

[0023] Because the embodiments described herein can perform eye-tracking calibration based on, for example, video of the subject captured while performing the test sequence itself, such embodiments eliminate the need for the aforementioned pre-calibration procedure. Furthermore, if eye-tracking calibration is performed once per test sequence, the relative positions and orientations of the subject, display, and camera need only remain constant for the duration of each test sequence (e.g., 30-60 seconds in some cases) to achieve accurate results. Furthermore, this time can be further reduced if automatic eye-tracking calibration is performed more than once per test sequence (a feature that may be implemented in certain embodiments). Furthermore, because the embodiments described herein can perform eye-tracking calibration one or more times during a test sequence, such embodiments can adjust their gaze estimation function based on any changes in the position / orientation of the subject, display, and camera that may have occurred since the previous calibration, thereby improving the overall accuracy of the test results.

[0024] In addition to determining the subject's gaze point, if the distance from the display to the subject's head is known, the embodiments described herein may also trigonometrically calculate the angle of the eye's gaze relative to the line of sight from the gaze point on the display to the central target.

[0025] To further explain the aforementioned concepts, reference is now made to Figure 2. Figure 2 illustrates, among other things, an example system 200 for tracking a subject's gaze in a manner that incorporates automatic eye-tracking calibration, according to some embodiments. As shown in Figure 2, system 200 includes a subject 202, a computing device 204, a display 206, and a camera 208. Each of these aspects of system 200 will now be described.

[0026] The subject 202 may be a person undergoing an oculometry test. The oculometry test administered to the subject 202 may be designed, for example, to measure the oculomotor abilities of the subject 202. For example, the oculometry test may be a saccade test designed to measure the subject's 202 ability to move their eyes from one designated focus point to another in a single, rapid movement. As another example, the oculometry test may be a smooth pursuit test designed to measure the subject's 202 ability to accurately track a visual target in a smooth, controlled manner. As yet another example, the oculometry test may include a long fixation test designed to measure the subject's 202 ability to accurately hold their eyes on a specific spot for a relatively long period of time. However, these are merely examples, and other types of oculometry tests may be administered to the subject 202 to measure their oculomotor abilities.

[0027] Conducting an oculometry test may involve instructing the subject 202 to perform an oculometry task in response to stimuli presented via the display 206. For example, as previously described, conducting a saccade test may involve instructing the subject 202 to gaze directly at a fixation point presented within the display 206 as the fixation point switches between various locations within the display. As another example, conducting a smooth pursuit test may involve instructing the subject 202 to follow with their eyes a visual target presented within the display 206 as the visual target moves from one side of the display to the other in a smooth, predictable motion. In either case, conducting a test may involve sequentially rendering multiple images or frames on the display 206, and therefore a particular test may be referred to herein as a test sequence.

[0028] Computing device 204 may include a device including one or more processors configured to perform at least the functions attributed to computing device 204 in the following description. The one or more processors may include, for example, but not limited to, one or more central processing units (CPUs), microcontrollers, microprocessors, signal processors, application-specific integrated circuits (ASICs), and / or other physical hardware processor circuitry. Computing device 204 may include, by way of example only, but not limited to, a smartphone, laptop computer, notebook computer, tablet computer, netbook, desktop computer, smart television, video game console, or wearable device (e.g., smart watch, smart glasses, augmented reality headset).

[0029] Display 206 may include a device capable of rendering images generated by computing device 204 in a manner that allows subject 202 to visually perceive the images. Display 206 may include, for example, a display screen integrated with computing device 204 (e.g., the integrated display of a smartphone, tablet computer, or laptop computer) or a monitor separate from but connected to computing device 204 (e.g., a monitor connected to a desktop computer via a wired connection). Display 206 may also include a display panel of a standalone or tethered augmented reality headset, or may include a projector capable of projecting images onto a surface visible to subject 200. However, the above are merely examples of display 206 and other implementations are possible.

[0030] The camera 208 includes optical equipment capable of capturing and storing images and video. The camera 208 may include, for example, a digital camera that captures images and video via an electronic image sensor. The camera 208 may be integrated with the computing device 204 (e.g., an integrated camera in a smartphone, tablet computer, or laptop computer) or may be part of a device that is separate from but connected to the computing device 204 (e.g., a USB camera or webcam connected to a desktop computer via a wired connection). However, the above are merely examples of the camera 208, and other implementations are possible.

[0031] 2, computing device 204 includes a test conductor 210 and a test results generator 212. Each of these components may be implemented by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed by one or more processors of computing device 204), or a combination thereof. Each of these components will now be described.

[0032] The test conductor 210 is configured to conduct a test that involves tracking the gaze of a subject, such as the subject 202. The test may be conducted to measure, for example, but not limited to, the eye movement abilities of the subject 202. To conduct the test, the test conductor 210 may present a sequence of images to the subject 202 via the display 206, the sequence of images including stimuli visible to the subject 202. The subject 202 may be instructed (e.g., by the test conductor module 210 and / or a human test administrator) to perform a directional eye measurement task in response to the stimuli presented via the display 206. As described above, such directional eye measurement tasks may include, for example, but not limited to, a saccade test, a smooth pursuit test, or a long fixation test.

[0033] Test conductor 210 is further configured to operate camera 208 to capture video of at least a portion of subject's 202's face while subject 202 performs the directional oculometry task. Ideally, the video image captured using camera 208 includes both subject's 202's right and left eyes, although, according to certain embodiments, usable test results may still be achieved even if one or both eyes are not captured during portions of the test. Test conductor 210 is further configured to store the video captured using camera 208 so that it can be accessed by test results generator 212.

[0034] The test results generator 212 is configured to obtain video captured by the test conductor 210 and analyze such video to track the gaze of the subject 202 (e.g., to measure their eye movement capabilities) in a manner that incorporates automatic eye-tracking calibration. To accomplish this, the test results generator 212 may analyze the video captured by the test conductor 210 to determine a series of time points during which the gaze of the subject 202 remains on a visual target and, for each time point in the series of time points, determine a set of ocular features. For example, the test results generator 212 may analyze the video captured by the test conductor 210 to determine a series of frames during which the gaze of the subject 202 remains on a visual target and, for each frame in the series of frames, determine a set of ocular features.

[0035] The set of eye features may include, for example, but are not limited to, pupil position relative to the eye opening. An example of a technique for measuring pupil position relative to the eye opening will now be described with reference to FIG. 3. FIG. 3 shows a close-up eye image 300 with overlaid markings, particularly illustrating the measurement of pupil position relative to the eye opening. A first curve 302 approximates the contour of the upper eyelid. A second curve 304 approximates the contour of the lower eyelid. The two curves intersect at two corners of the eye, one at the innermost part of the eye opening and one at the outermost part of the eye opening. Long dashed lines are drawn through these two corners. The eye opening is defined as all the area completely enclosed by the two curves.

[0036] In Figure 3, the variables a and b are used to specify the coordinates of the center of the pupil. The origin (0,0) is the center of gravity of the eye opening. The a-axis is a line parallel to the long dashed line and passing through the origin. The b-axis is a line perpendicular to the a-axis and passing through the origin.

[0037] In one embodiment, the test result generator 212 is configured to measure (a, b) of the viewing left or right eye for each of a plurality of time points (e.g., frames) at which the subject's 202 gaze is determined to remain on the visual target. The test result generator 212 also obtains the position of the visual stimulus within the display area of ​​the display 206 for each of the plurality of time points. The corresponding position of the visual stimulus within the display area may be represented using coordinates (x, y), where x is the horizontal coordinate and y is the vertical coordinate. As an example, FIGS. 4A and 4B show a display area 400 of the display 206 and a fixation point 402 (e.g., a dot) that may be sequentially presented at different positions thereon as part of a saccade test. According to this example, the position of the fixation point 402 at a given time point (or for a given frame) may be represented using (x, y) coordinates, where x is the horizontal coordinate and y is the vertical coordinate. When the fixation point 402 is centered within the display area 400, it is at the origin (0, 0). However, this is merely one example, and any of a wide variety of coordinate systems may be used to represent the location of the stimulus within the viewing area of ​​display 206.

[0038] After obtaining (i) a set of eye features for each of the time points in the series of time points, and (ii) the location of the visual stimulus within the display area of ​​the display 206 for each of the time points in the series of time points, the test results generator 212 may use such information to determine a mapping that associates the set of eye features with the subject's point of gaze.

[0039] For example, continuing with the above example, the test results generator 212 may perform a regression analysis to derive, for each of the left and right eyes, a continuous function of the input variables a and b and x and y from all [(a,b),(x,y)] pairs associated with fixations on a visual target. Note that a single instance or multiple instances of fixations on a visual target by the subject 202 may be utilized to accumulate these [(a,b),(x,y)] pairs.

[0040] The test results generator 212 may utilize any of a variety of different algorithms to perform the regression analysis. For example, any of several known algorithms from the field of supervised machine learning may be used to perform the regression analysis. These include, but are not limited to, linear regression, polynomial regression, and decision tree regression.

[0041] In one embodiment, the test results generator 212 utilizes response surface analysis to perform the regression analysis. Response surface analysis is the name for polynomial regression with two input variables and one output variable. For example, according to such an embodiment, the test results generator 212 may fit the following two quadratic polynomials to all of the [(a, b), (x, y)] pairs associated with gazes staying on the visual target to obtain two response surfaces: x=C 2,0 a 2 +C 1,1 ab+C 0,2 b 2 +C 1,0 a+C 0,1 b+C 0,0 y=K 2,0 a 2 +K 1,1 ab+K 0,2 b 2 +K 1,0 a+K 0,1 b+K 0,0 After the test results generator 212 determines the coefficients C and K, the test results generator 212 can then calculate the fixation point (x, y) from any single sample (a, b). According to this example, there are five C coefficients and five K coefficients, so the test results generator 212 must have at least five [(a, b), (x, y)] pairs to obtain two response surfaces. However, the test results generator 212 may be configured to use more than five (e.g., many more) [(a, b), (x, y)] pairs to help mitigate the effects of noise.

[0042] Once the test results generator 212 has determined the aforementioned mapping that associates a set of ocular features (e.g., pupil position represented by (a, b)) with a subject's gaze point (e.g., a position within the display area of ​​the display 206 represented as (x, y)), the test results generator 212 may then apply the mapping to the set of features respectively associated with some or all of the time points (e.g., frames) in the video of the test sequence at which the left or right eye is visible, thereby obtaining the corresponding gaze points of the subject 202 for such time points. These corresponding gaze points of the subject 202 may then be used by the test results generator 212, for example, to measure the eye movement performance of the subject 202.

[0043] To further illustrate, the following is an example of an algorithm by which the test conductor 210 and test result generator 212, according to some embodiments, may work together to determine left and right response surfaces and calculate the gaze points (x,y) of all eyes found during the test sequence: TIFF2026502269000002.tif248157

[0044] As described in step 1 of the algorithm, the test results generator 212 sequentially applies a face detection and eye-finding algorithm to each frame of video captured by the test conductor 210 to find either the left eye or the right eye, or both, of the subject 202 within each frame. If the test results generator 212 determines that a left eye has been found with a sufficiently high degree of confidence within a given frame, it sets the variable LEF to true for that frame; otherwise, it sets the variable LEF to false for that frame. Similarly, if the test results generator 212 determines that a right eye has been found with a sufficiently high degree of confidence within a given frame, it sets the variable REF to true for that frame; otherwise, it sets the variable REF to false for that frame. The face detection and eye-finding algorithm described above may similarly be used to generate cropped images of eyes, labeled as either left or right, if eyes are detected within a frame; these cropped images may then be used to determine eye characteristics. Any of a wide variety of well-known techniques for performing face detection and eye finding on images may be used to implement the foregoing aspects of this exemplary algorithm.

[0045] As noted in step 2 of the algorithm, the test conductor 210 positions the target in the center of the viewing area of ​​the display 206 for two seconds so that the subject 202 can fixate their gaze on it. The test results generator 212 analyzes frames captured by the camera 208 during this period to collect (a, b) for each left eye found in the frame and (a, b) for each right eye found in the frame. Any of a variety of well-known techniques for accurately and reliably locating the center of the pupil within an eye image may be utilized to implement this aspect of the exemplary algorithm. Further according to step 2 of the algorithm, the test results generator 212 then calculates a threshold t equal to 6*STD(a), where STD(a) is the standard deviation of a for all left eyes found in frames captured during the two-second period. laand a threshold t equal to 6*STD(b), where STD(b) is the standard deviation of b for all left eyes found in frames captured during a 2-second period. lb Similarly, the test results generator 212 calculates a threshold t that is equal to 6*STD(a), where STD(a) is the standard deviation of a for all right eyes found in frames captured during a 2-second period. ra and a threshold t equal to 6*STD(b), where STD(b) is the standard deviation of b for all right eyes found in the frame during a 2-second period. rb Calculate the following.

[0046] Step 2 of the algorithm effectively calculates the range of variation in the (a,b) measurements for each eye when the subject holds a fixed gaze. This range of variation can represent the "noise" in the (a,b) measurements, which may be caused by variations in the subject's ability to hold a steady gaze, limitations in the algorithm, lighting variations, camera performance limitations, or other factors. The largest contributor to noise may be variation in the subject's ability to hold a steady gaze. For example, approximately 2% of the world's population has strabismus, a condition that prevents a subject from steadily aligning both eyes on a gaze target. Thus, step 2 of the above algorithm may be used to detect strabismus, allowing some measurements of eye movements to be excluded or interpreted differently.

[0047] As noted in step 3 of the algorithm, the test conductor 210 begins the test sequence (e.g., begins presenting visual stimuli for the subject 202 to perform a directional oculometry task), at which point the variable ACTIVE is set to true. When the test conductor 210 finishes the test sequence, the variable ACTIVE is set to false. As further noted in step 3, presenting visual stimuli may include sequentially presenting dots (or other visual stimuli) at various dwell locations within the viewing area of ​​the display 206 for a particular duration. In this example, the duration is 48 frames, or 800 milliseconds when the camera 208 is operating at 60 frames per second (FPS). Note that to facilitate correlation between video frames captured using the camera 208 and images presented on the viewing area of ​​the display 206, the camera 208 and the display 206 may each operate, or be controlled to operate, at the same FPS (e.g., both the camera 208 and the display 206 may operate at 60 FPS).

[0048] As described in step 4 of the algorithm, the test result generator 212 sets up two empty first-in, first-out (FIFO) queues, one for the left eye and one for the right eye, each with the same depth. The test result generator 212 utilizes these FIFO queues to temporarily store the [(a, b), (x, y)] pairs being tested, thereby identifying samples captured when the subject's 202 gaze remained on a visual target. In an exemplary embodiment in which both the camera 208 and the display 206 operate at 60 FPS, a FIFO size of 15 may be selected to support the search for a stable dwell time of at least 250 ms. As previously mentioned, the actual realistic time span required to acquire 15 samples may be larger because the subject may blink for approximately 100–300 ms during the target dwell time of 800 ms set in step 3.

[0049] As mentioned in step 5 of the algorithm, the test results generator 212 also sets up two empty DATA lists, one for the left eye and one for the right eye, which the test results generator 212 uses to store [(a,b),(x,y)] pairs identified as samples captured when the subject's 202 gaze remained on a visual target.

[0050] Step 6 of the algorithm describes a process for selectively adding [(a,b),(x,y)] pairs stored in the left eye FIFO queue to the corresponding left eye DATA list and selectively adding [(a,b),(x,y)] pairs stored in the right eye FIFO queue to the corresponding right eye DATA list. For example, as shown in the algorithm, while ACTIVE is true, for each video frame for which LEF is true, the corresponding [(a,b),(x,y)] pair for the left eye is pushed into the left eye FIFO. If the left eye FIFO queue is full and the (a,b) values ​​stored therein are greater than the previously determined t la value and t lb If the left eye FIFO queue shows a sufficiently low variance compared to the left eye DATA list, then a sample in the left eye FIFO queue that has not yet been added to the left eye DATA list is added to the left eye data list. A similar process occurs for the right eye FIFO queue and the right eye DATA list while ACTIVE is true.

[0051] As described in step 7 of the algorithm, the test results generator 212 calculates left-eye and right-eye response surfaces (x, y) as functions of (a, b) from the left-eye DATA list and right-eye DATA list, respectively. The left-eye response surface collectively includes a mapping that maps a set of left-eye features (e.g., left-eye pupil positions represented by (a, b)) to the gaze points of the subject's left eye (e.g., the gaze points of the subject's left eye relative to the display area of ​​the display 206 represented by (x, y)). The right-eye response surface collectively includes a mapping that maps a set of right-eye features (e.g., right-eye pupil positions represented by (a, b)) to the gaze points of the subject's right eye (e.g., the gaze points of the subject's right eye relative to the display area of ​​the display 206 represented by (x, y)).

[0052] As noted in step 8 of the algorithm, the test results generator 212 uses the aforementioned left and right eye response surfaces to calculate the gaze points (x,y) for all eyes found during the entire test sequence. That is, for each eye detected during the test sequence, the test results generator 212 may obtain a set of eye features and utilize the appropriate response surface to generate a gaze point corresponding to the set of eye features.

[0053] It should be noted that the above algorithms may be modified to periodically or intermittently recalculate the left and right eye response surfaces, such that the left and right eye response surfaces may be recalculated multiple times during a single test sequence. According to such implementations, the test results generator 212 may be configured to use the last calculated response surface to calculate the gaze direction (x, y) of a given eye found during the test sequence.

[0054] Although the foregoing description describes the test conductor 210 and the test results generator 212 as being part of the same computing device (i.e., computing device 204), these components need not be implemented on the same computing device. For example, the test conductor 210 may be implemented on a first computing device (e.g., a smartphone, tablet, or laptop computer), and the test results generator 212 may be implemented on a second computing device (e.g., a server in a cloud computing network) communicatively connected to the first computing device. According to such a distributed implementation, the test conductor 210 may conduct an ophthalmometric test in the manner described above and then upload or otherwise make available to the second computing device a video of the subject undergoing the ophthalmometric test. The test results generator 212 running on the second computing device may then analyze the video and track the subject's gaze in the manner described above to determine the results of the ophthalmometric test.

[0055] Additionally, the test conductor 210 and the test results generator 212 may operate simultaneously to conduct an ophthalmometric test. For example, video of the subject 202 captured by the test conductor 210 while conducting the ophthalmometric test may be made accessible (e.g., streamed) to the test results generator 212 in near real time, and the test results generator 212 may analyze such video, track the subject's gaze in the manner described above, and determine the results of the ophthalmometric test. However, the test conductor 210 and the test results generator 212 may also operate at different times to conduct the ophthalmometric test. For example, the test conductor 212 may conduct the ophthalmometric test in the manner described above and then store the video of the subject undergoing the ophthalmometric test for later analysis. The test results generator 212 may then access and analyze the video at a later time (e.g., minutes, hours, or days later), track the subject's gaze in the manner described above, and determine the results of the ophthalmometric test.

[0056] In certain embodiments, the automated eye tracking calibration performed by the test results generator 212 may be performed in a manner that compensates for changes in the position and orientation of the subject 202. For example, the test results generator 212 may use the position of the canthi of both eyes within the camera's field of view, in addition to the pupil position relative to the eye opening, to incorporate information about the subject's position and orientation into the calculation of the gaze point (x, y). This may be achieved, for example, by adding a term dependent on the 4 × 2 = 8 coordinates of the canthi of the eyes to the above quadratic polynomial for x and y of each eye. This increases the number of coefficients C and K to a certain number N. Thus, the test results generator 212 can process any number of video frame images necessary to obtain at least N frames, from which all input variables, i.e., (a, b) from both eyes and the corresponding eight canthi coordinates, can be extracted. The test results generator 212 may then calculate the C and K coefficients through regression analysis.

[0057] The shapes of the pupil and iris in eye images have been observed to change with the subject's gaze direction and the camera's field of view. Iris shape may be a better choice for estimating gaze angle because relatively small pupils may be more easily occluded by the eyelids or may appear distorted due to insufficient camera resolution. For example, an ellipse may be fitted to the iris-sclera boundary in eye images. The length and orientation of the minor and major ellipsoidal axes of the iris may be used to estimate the gaze angle relative to the camera's optical axis. Thus, additional terms depending on the length and orientation of these axes may be added, similar to how the canthus position can be used to refine quadratic polynomials in x and y for each eye. As with adding canthus information, N increases, and a sufficiently large number of frames must be analyzed. Therefore, the test result generator 212 can process any number of video frame images necessary to obtain at least a sufficiently large number of frames, from which all input variables, i.e., (a, b) from both eyes, the corresponding eight coordinates of the canthi, and the lengths and orientations of the minor and major elliptical axes of the eyes, can be extracted. The test result generator 212 can then calculate the C and K coefficients by regression analysis.

[0058] Note that to compensate for changes in subject position and orientation, either or both of the eye canthus information or iris shape information may be added to the input variables for mapping to (x,y) gaze coordinates.

[0059] As described above with respect to step 1 of the example algorithm, the test results generator 212 may use face detection and eye detection to generate cropped images of eyes labeled as either left or right. The test results generator 212 may achieve this by using a convolutional neural network (CNN). Furthermore, instead of processing these images to extract (a, b) measurements as described above, the left and right eye images may be fed into two separate copies of a CNN trained to estimate gaze positions within the viewing area of ​​the display 206. The gaze estimation CNN may be trained on a dataset of diverse face images with many gaze angles obtained from many people. These two gaze estimation CNNs may be combined with an eye detection CNN to construct the neural network 500 shown in FIG. 5. As shown in FIG. 5, the right eye image may be input to a convolutional layer 502 (which may also include one or more pooling layers), the left eye image may be input to a convolutional layer 504 (which may also include one or more pooling layers), and one or two eye positions may be input to a fully connected layer 506. The outputs from the convolutional layer 502, the convolutional layer 504, and the fully connected layer 506 may be passed to a fully connected layer 508. The fully connected layer 508 may estimate the left and right eye coordinates (x e ,y e ), which can be used by test results generator 212 to determine response surfaces for the left and right eyes, as described below.

[0060] For example, if the estimated gaze point within the viewing area of ​​the display 206 generated by the neural network 500 is (x e ,y e ) Furthermore, the above response surface analysis formula may be rewritten as follows: x=C 2,0 x e 2 +C 1,1 x e y e 2 +C 0,2 y e 2 +C 1,0 x e +C0,1 y e +C 0,0 y=K 2,0 x e 2+K 1,1 x e y e +K 0,2 y e 2+K 1,0 x e +K 0,1 y e +K 0,0 With these modifications, the test results generator 212 can generate response surfaces for the left and right eyes using the algorithm described above, where all (a, b) references are (x e ,y e ) reference.

[0061] The above-described technique for tracking a subject's gaze point can be used to perform oculometry tests to detect the subject's neurological or mental health condition. However, the above-described technique can also be used in a variety of other eye-tracking applications that may involve or benefit from tracking a subject's gaze point. Such applications include, but are not limited to, applications in the fields of marketing, human-computer interaction, virtual reality and augmented reality, or any application in which the dwell position can be included in the application content presented on the display.

[0062] 6 is a flow diagram of a method 600 for tracking a subject's gaze in a manner incorporating automatic eye-tracking calibration, according to some embodiments. Method 600 can be performed by processing logic that can include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed by a processing device), or a combination thereof. It should be understood that not all steps are required to carry out the disclosure provided herein. Furthermore, as will be understood by one of ordinary skill in the art, some of the steps may be performed simultaneously or in a different order than that shown in FIG. 6.

[0063] Although the method 600 is described with reference to the system 200 of FIG. 2, the method 600 is not limited in that respect.

[0064] At 602, the test conductor 210 uses the camera 208 to capture video of at least a portion of the face of the subject 202 while the subject 202 views content of an eye-tracking application within the display area of ​​the display 206. The eye-tracking application may include, for example, but is not limited to, an oculometry test application that tracks the gaze of the subject 202 while the subject 202 performs a directional oculometry task. For example, the test conductor 210 may use the camera 208 to capture video of at least a portion of the face of the subject 202 while the subject 202 performs a directional oculometry task in response to visual stimuli presented within the display area of ​​the display 206. Such a directional oculometry task may include, for example, a saccade test, a smooth pursuit test, or a long fixation test. However, it should be noted that the eye-tracking application may include eye-tracking applications in the fields of marketing, human-computer interaction, or virtual and augmented reality, or any application in which application content presented by the display 206 may include dwell positions.

[0065] At 604, the test results generator 212 analyzes the video to determine a series of time points during which the subject's 202 gaze remains on the visual target, and to determine a set of ocular features for each time point in the series of time points. Each set of ocular features may include, for example, a set of right eye features or a set of left eye features. Additionally, each set of eye features may include one or more of the following: pupil position relative to the eye opening, canthus position of the eye, or length and orientation of the minor and major elliptical axes of the eye's iris.

[0066] In one aspect, the test results generator 212 may analyze the video to determine a set of eye features for each of the time points in the series of time points by extracting eye images from frames of the video corresponding to the time points in the series of time points and providing the eye images to a neural network that outputs, based at least on the eye images, an estimated eye gaze direction for the time points in the series of time points.

[0067] At 606, the test results generator 212 obtains the location of the visual stimulus within the display area of ​​the display 206 for each time point in the series of time points.

[0068] At 608, the test results generator 212 determines a mapping that associates the set of eye features with the gaze point, the determination being based on at least (i) the set of eye features for each of the time points in the series of time points, and (ii) the location of the visual stimulus within the display area of ​​the display 206 for each of the time points in the series of time points.

[0069] The test results generator 212 may determine the mapping by performing a regression analysis. As an example, the test results generator 212 may perform one of a linear regression analysis, a polynomial regression analysis, or a decision tree regression analysis to determine the mapping. Furthermore, according to an exemplary embodiment in which the test results generator 212 performs a polynomial regression analysis to determine the mapping, the test results generator 212 may perform a response surface analysis to determine the mapping.

[0070] At 610, the test results generator 212 uses the mapping to respectively map one or more sets of eye features obtained by analyzing the video to one or more gaze points of the subject 202. If the one or more sets of eye features include one or more sets of right eye features, the one or more gaze points include one or more gaze points of the right eye of the subject 202. Similarly, if the one or more sets of eye features include one or more sets of left eye features, the one or more gaze points include one or more gaze points of the left eye of the subject 202. In certain implementations, the test results generator 212 may detect a neurological or mental health condition of the subject 202 based at least on the one or more gaze points of the subject determined at 610.

[0071] 7 is a flow diagram of a method 700 for recalculating a mapping that associates a set of ocular features with a point of gaze, according to some embodiments. Similar to method 600, method 700 can be performed by processing logic that can include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed by a processing device), or a combination thereof. It should be understood that not all steps are required to carry out the disclosure provided herein. Additionally, as will be appreciated by those skilled in the art, some of the steps may occur simultaneously or in a different order than that shown in FIG.

[0072] Although the method 700 will be described with continued reference to the system 200 of FIG. 2, the method 700 is not limited in that respect.

[0073] Method 700 may be performed after the steps of method 600 have been performed. At 702, test results generator 212 analyzes the video acquired in 602 to determine one or more additional time points at which the subject's 202 gaze remains on a visual target and to determine a set of ocular features for each of the one or more additional time points.

[0074] At 704, the test results generator 212 obtains the location of the visual stimulus within the display area of ​​the display 206 for each of the one or more additional time points.

[0075] At 706, the test results generator 212 recalculates the mapping previously determined at 608 based on at least (i) the set of eye features for each of the one or more additional time points, and (ii) the location of the visual stimulus within the display area of ​​the display 206 for each of the one or more additional time points.

[0076] Example of a computer system Various aspects may be implemented using one or more well-known computer systems, such as, for example, computer system 800 shown in Figure 8. For example, computing device 204 may be implemented using a combination or sub-combination of computer systems 800. Similarly, or alternatively, one or more computer systems 800 may be used to implement, for example, any of the aspects described herein, as well as combinations and sub-combinations thereof.

[0077] Computer system 800 may include one or more processors (also referred to as central processing units or CPUs), such as processor 804. Processor 804 may be connected to a communication infrastructure or bus 806.

[0078] The computer system 800 may also include one or more user input / output devices 803, such as a monitor, keyboard, pointing device, etc., that may communicate with a communications infrastructure 806 via one or more user input / output interfaces 802.

[0079] One or more of the processors 804 may be a graphics processing unit (GPU). In one aspect, a GPU may be a processor that is a dedicated electronic circuit designed to process mathematically intensive applications. A GPU may have a parallel structure that is efficient for parallel processing of large blocks of data, such as mathematically intensive data common in computer graphics applications, images, video, etc.

[0080] Computer system 800 may also include a main or primary memory 808, such as random access memory (RAM). Main memory 808 may include one or more levels of cache. Main memory 808 may store control logic (i.e., computer software) and / or data.

[0081] Computer system 800 may also include one or more secondary storage devices or memories 810. Secondary memory 810 may include, for example, a hard disk drive 812 and / or a removable storage device or drive 814. Removable storage drive 814 may be a floppy disk drive, a magnetic tape drive, a compact disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.

[0082] The removable storage drive 814 may interact with a removable storage unit 818. The removable storage unit 818 may include a computer-usable or readable storage device on which computer software (control logic) and / or data is stored. The removable storage unit 818 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / or any other computer data storage device. The removable storage drive 814 may read from and / or write to the removable storage unit 818.

[0083] Secondary memory 810 may include other means, devices, components, implements, or other techniques for allowing computer programs and / or other instructions and / or data to be accessed by computer system 800. Such means, devices, components, implements, or other techniques may include, for example, removable storage unit 822 and interface 820. Examples of removable storage unit 822 and interface 820 include program cartridges and cartridge interfaces (such as those found in video game devices), removable memory chips (such as EPROMs and PROMs) and associated sockets, memory sticks and USB or other ports, memory cards and associated memory card slots, and / or any other removable storage unit and associated interface.

[0084] Computer system 800 may further include a communications or network interface 824. Communications interface 824 may enable computer system 800 to communicate and interact with any combination of external devices, external networks, external entities, etc. (individually and collectively referred to by reference numeral 828). For example, communications interface 824 may enable computer system 800 to communicate with external or remote devices 828 via communications path 826, which may be wired and / or wireless (or a combination thereof) and may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be transmitted to and from computer system 800 via communications path 826.

[0085] Computer system 800 may also be a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet, a smartphone, a smartwatch or other wearable, an appliance, part of the Internet of Things, and / or an embedded system, or any combination thereof, to name a few non-limiting examples.

[0086] The computer system 800 may be implemented using, but is not limited to, remote or distributed cloud computing solutions; local or on-premise software ("on-premise" cloud-based solutions); The client or server may access or host any application and / or data via any delivery paradigm, including a "service-as-a-service" model (e.g., Content-as-a-Service (CaaS), Digital Content-as-a-Service (DCaaS), Software-as-a-Service (SaaS), Managed Software-as-a-Service (MSaaS), Platform-as-a-Service (PaaS), Desktop-as-a-Service (DaaS), Framework-as-a-Service (FaaS), Backend-as-a-Service (BaaS), Mobile Backend-as-a-Service (MBaaS), Infrastructure-as-a-Service (IaaS), etc.); and / or a hybrid model including any combination of the foregoing or other service or delivery paradigms.

[0087] Any applicable data structures, file formats, and schemas in computer system 800 may be derived from standards, including, but not limited to, JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or any other functionally similar representations, alone or in combination. Alternatively, proprietary data structures, formats, or schemas may be used exclusively or in combination with known or open standards.

[0088] In some aspects, a tangible, non-transitory apparatus or article of manufacture comprising a tangible, non-transitory computer-usable or readable medium having control logic (software) stored thereon may also be referred to herein as a computer program product or program storage device. This includes, but is not limited to, a tangible article of manufacture embodying computer system 800, main memory 808, secondary memory 810, and removable storage units 818, 822, as well as any combination thereof. Such control logic, when executed by one or more data processing devices (e.g., computer system 800 and processor(s) 804), may cause such data processing devices to operate as described herein.

[0089] Based on the teachings contained herein, it will be apparent to one skilled in the art how to make and use aspects of the present disclosure using data processing devices, computer systems and / or computer architectures other than those shown in Figure 8. In particular, aspects may operate with software, hardware, and / or operating system implementations other than those described herein.

[0090] conclusion It is understood that the Detailed Description section, and not any other section, is intended to be used to interpret the claims. While the other sections may describe one or more exemplary aspects contemplated by the inventor(s), they cannot describe everything and are therefore not intended to limit the scope of the disclosure or the appended claims in any way.

[0091] While this disclosure describes exemplary embodiments for exemplary fields and applications, it should be understood that the disclosure is not limited thereto. Other embodiments and modifications thereof are possible and are within the scope and spirit of the disclosure. For example, without limiting the generality of this paragraph, the embodiments are not limited to the software, hardware, firmware, and / or entities shown and / or described herein. Moreover, the embodiments (whether or not explicitly described herein) have significant utility in fields and applications beyond the examples described herein.

[0092] Aspects are described herein with functional building blocks that illustrate implementations of specified functions and relationships thereof. The boundaries of these functional building blocks are arbitrarily defined herein for convenience of description. Alternative boundaries may be defined so long as the specified functions and relationships (or equivalents) are appropriately performed. Alternative aspects may also use an order of function blocks, steps, operations, methods, etc. different from that described herein.

[0093] References herein to “one embodiment,” “one embodiment,” “one exemplary embodiment,” or similar phrases indicate that the described embodiment may include a particular feature, structure, or characteristic, but not all embodiments necessarily include that particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Furthermore, if a particular feature, structure, or characteristic is described in the context of one embodiment, it is within the knowledge of one of ordinary skill in the art to incorporate such feature, structure, or characteristic into other embodiments, whether or not explicitly described. Furthermore, some embodiments may be described using the terms “coupled” and “connected,” along with their derivatives. These terms are not necessarily intended as synonyms for each other. For example, some embodiments may be described using the terms “connected” and / or “coupled” to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term “coupled” can also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0094] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

Claims

1. A computer-implemented method for tracking a subject's gaze, comprising: acquiring, by at least one computer processor, video of at least a portion of the subject's face while the subject views content of an eye-tracking application presented within a viewing area of ​​a display; analyzing the video to determine a series of time points during which the subject's gaze remains on a visual target, and to determine a set of ocular features for at least each time point in the series of time points; obtaining, for each of the time points in the series of time points, a position of a visual stimulus within the display area of ​​the display; determining a mapping that associates a set of ocular features with a point of gaze, said determining being based at least on (i) the set of ocular features for each of the time points in the series of time points, and (ii) the position of the visual stimulus within the display area of ​​the display for each of the time points in the series of time points; and Using the mapping to respectively map one or more sets of eye features obtained by analyzing the video to one or more gaze points of the subject.

2. the one or more sets of eye features obtained by analyzing the video include one or more sets of right eye features, and the one or more gaze points of the subject include one or more gaze points of the subject's right eye; or the one or more sets of eye features obtained by analyzing the video include one or more sets of left eye features, and the one or more gaze points of the subject include one or more gaze points of the subject's left eye; 10. The computer-implemented method of claim 1.

3. detecting a neurological or mental health state of the subject based at least on the one or more gaze points of the subject; 10. The computer-implemented method of claim 1, further comprising:

4. acquiring the video of the at least a portion of the face of the subject while the subject is viewing the content of the eye-tracking application presented within the viewing area of ​​the display, capturing the video of the at least a portion of the face of the subject while the subject is performing a directional oculometry task in response to visual stimuli presented within the viewing area of ​​the display; 10. The computer-implemented method of claim 1, comprising:

5. the directional oculometry task comprises: Saccade test, smooth pursuit test, or Long-term fixation test 5. The computer-implemented method of claim 4, comprising one of:

6. Each set of eye features is Pupil position relative to the eye opening, The canthus position of the eyes, or The length and orientation of the minor and major ellipsoidal axes of the eye's iris 10. The computer-implemented method of claim 1, comprising one or more of:

7. determining the mapping comprises: performing a regression analysis to determine said mapping; 10. The computer-implemented method of claim 1, comprising:

8. performing the regression analysis to determine the mapping; Conducting a linear regression analysis; Conducting polynomial regression analysis, or Conducting decision tree regression analysis 8. The computer-implemented method of claim 7, comprising one of:

9. analyzing the video to determine the set of ocular features for each of the time points in the series of time points; extracting an eye image from a frame of the video corresponding to the time point in the series of time points; and providing the eye images to a neural network that outputs an estimated point of gaze for the time point in the series of time points based on at least the eye images; 10. The computer-implemented method of claim 1, comprising:

10. analyzing the video to determine one or more additional time points during which the subject's gaze remains on a visual target, and to determine a set of ocular features for each of the one or more additional time points; acquiring a position of the visual stimulus within the display area of ​​the display for each of the one or more additional time points; and recalculating the mapping based at least on (i) the set of ocular features for each of the one or more additional time points, and (ii) the position of the visual stimulus within the viewing area of ​​the display for each of the one or more additional time points.

10. The computer-implemented method of claim 1, further comprising:

11. one or more memories; at least one processor each coupled to at least one of the memories; acquiring video of at least a portion of a subject's face while the subject views content of an eye-tracking application presented within a viewing area of ​​a display; analyzing the video to determine a series of time points during which the subject's gaze remains on a visual target, and to determine a set of ocular features for at least each time point in the series of time points; obtaining, for each of the time points in the series of time points, a position of a visual stimulus presented within the display area of ​​the display; determining a mapping that associates a set of ocular features with a point of gaze, said determining being based at least on (i) the set of ocular features for each of the time points in the series of time points, and (ii) the position of the visual stimulus within the display area of ​​the display for each of the time points in the series of time points; and Using the mapping to respectively map one or more sets of eye features obtained by analyzing the video to one or more gaze points of the subject. at least one processor configured to perform operations including a system for tracking and measuring a gaze point of the subject, comprising:

12. the one or more sets of eye features obtained by analyzing the video include one or more sets of right eye features, and the one or more gaze points of the subject include one or more gaze points of the subject's right eye; or the one or more sets of eye features obtained by analyzing the video include one or more sets of left eye features, and the one or more gaze points of the subject include one or more gaze points of the subject's left eye; 12. The system of claim 11.

13. The operation is Detecting a neurological or mental health state of the subject based at least on the one or more gaze points of the subject.

12. The system of claim 11, further comprising:

14. acquiring the video of the at least a portion of the face of the subject while the subject is viewing the content of the eye-tracking application presented within the viewing area of ​​the display; capturing the video of the at least a portion of the face of the subject while the subject is performing a directional oculometry task in response to the visual stimuli presented within the viewing area of ​​the display; 12. The system of claim 11, comprising:

15. the directional oculometry task comprises: Saccade test, smooth pursuit test, or Long-term fixation test 15. The system of claim 14, comprising one of:

16. Each set of eye features is Pupil position relative to the eye opening, The canthus position of the eyes, or The length and orientation of the minor and major ellipsoidal axes of the eye's iris 12. The system of claim 11, comprising one or more of:

17. determining the mapping performing a regression analysis to determine said first mapping; 12. The system of claim 11, comprising:

18. performing the regression analysis to determine the mapping; Conducting a linear regression analysis; Conducting polynomial regression analysis, or Conducting decision tree regression analysis 20. The system of claim 17, comprising one of:

19. analyzing the video to determine the set of ocular features for each of the time points in the series of time points; extracting an eye image from a frame of the video corresponding to the time point in the series of time points; and providing the eye images to a neural network that outputs an estimated point of gaze for the time point in the series of time points based on at least the eye images; 12. The system of claim 11, comprising:

20. The operation is analyzing the video to determine one or more additional time points during which the subject's gaze remains on a visual target, and to determine a set of ocular features for each of the one or more additional time points; obtaining a position of the visual stimulus within the display area of ​​the display for each of the one or more additional time points; and recalculating the mapping based at least on (i) the set of ocular features for each of the one or more additional time points, and (ii) the position of the visual stimulus within the viewing area of ​​the display for each of the one or more additional time points.

12. The system of claim 11, further comprising:

21. 1. A non-transitory computer-readable medium having stored thereon instructions that, when executed by at least one computer processor, cause the at least one computer processor to perform operations for tracking a subject's gaze, the operations comprising: acquiring video of at least a portion of the subject's face while the subject views content of an eye-tracking application presented within a viewing area of ​​a display; analyzing the video to determine a series of time points during which the subject's gaze remains on a visual target, and to determine a set of ocular features for at least each time point in the series of time points; obtaining, for each of the time points in the series of time points, a position of a presented visual stimulus within the display area of ​​the display; determining a mapping that associates a set of ocular features with a point of gaze, said determining being based at least on (i) the set of ocular features for each of the time points in the series of time points, and (ii) the position of the visual stimulus within the display area of ​​the display for each of the time points in the series of time points; and using the mapping to respectively map one or more sets of eye features obtained by analyzing the video to one or more gaze points of the subject. The non-transitory computer-readable medium comprising:

22. the one or more sets of eye features obtained by analyzing the video include one or more sets of right eye features, and the one or more gaze points of the subject include one or more gaze points of the subject's right eye; or the one or more sets of eye features obtained by analyzing the video include one or more sets of left eye features, and the one or more gaze points of the subject include one or more gaze points of the subject's left eye; 22. The non-transitory computer-readable medium of claim 21.

23. The operation is Detecting a neurological or mental health state of the subject based at least on the one or more gaze points of the subject.

22. The non-transitory computer-readable medium of claim 21, further comprising:

24. acquiring the video of the at least a portion of the face of the subject while the subject is viewing the content of the eye-tracking application presented within the viewing area of ​​the display; capturing the video of the at least a portion of the face of the subject while the subject is performing a directional oculometry task in response to the visual stimuli presented within the viewing area of ​​the display; 22. The non-transitory computer-readable medium of claim 21, comprising:

25. the directional oculometry task comprises: Saccade test, smooth pursuit test, or Long-term fixation test 25. The non-transitory computer-readable medium of claim 24, comprising one of:

26. Each set of eye features is Pupil position relative to the eye opening, The canthus position of the eyes, or The length and orientation of the minor and major ellipsoidal axes of the eye's iris 22. The non-transitory computer-readable medium of claim 21, comprising one or more of:

27. determining the mapping performing a regression analysis to determine said mapping; 22. The non-transitory computer-readable medium of claim 21, comprising:

28. performing the regression analysis to determine the mapping; Conducting a linear regression analysis; Conducting polynomial regression analysis, or Conducting decision tree regression analysis 28. The non-transitory computer-readable medium of claim 27, comprising one of:

29. analyzing the video to determine the set of ocular features for each of the time points in the series of time points; extracting an eye image from a frame of the video corresponding to the time point in the series of time points; and providing the eye images to a neural network that outputs an estimated point of gaze for the time point in the series of time points based on at least the eye images; 22. The non-transitory computer-readable medium of claim 21, comprising:

30. The operation is analyzing the video to determine one or more additional time points during which the subject's gaze remains on a visual target, and to determine a set of ocular features for each of the one or more additional time points; obtaining a position of the visual stimulus within the display area of ​​the display for each of the one or more additional time points; and recalculating the mapping based at least on (i) the set of ocular features for each of the one or more additional time points, and (ii) the position of the visual stimulus within the viewing area of ​​the display for each of the one or more additional time points.

22. The non-transitory computer-readable medium of claim 21, further comprising: