System and method for wearable eye tracking slip detection and correction
By calibrating the features captured by the eye image sensor in the eye tracking system, a mapping of the gaze vector is generated and the correction slip is detected, and the problems of inaccuracy and slip in the prior art are solved, and higher stability and accuracy are achieved.
Patent Information
- Application Number
- CN202380076264.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-15
- Filing Date
- 2023-11-14
- Publication Date
- 2025-06-27
AI Technical Summary
Existing eye tracking techniques are difficult to effectively detect and correct the relative position and orientation changes between the eye image sensor and the eye, resulting in inaccuracy and slip problems of the gaze vector.
By calibrating the first and second features in the eye image captured using the eye image sensor, a mapping from the feature to the gaze vector is generated and the sliding is detected and corrected in the eye tracking session is utilized.
More accurate gaze vector generation and slip correction are achieved, improving the stability and accuracy of the eye tracking system.
Smart Images

Figure CN120225975A_ABST
Abstract
Description
Background Art
[0001] Eye tracking can be used in a wide range of application fields such as human-machine interfaces, gaming, virtual reality, human behavior research, and medicine. Summary of the Invention
[0002] Disclosed herein is a method that includes performing a first eye tracking system calibration using a first feature in an eye image captured by an eye image sensor.
[0003] After the first calibration, a mapping from the first feature in the eye image to a gaze vector can be generated.
[0004] After the first calibration, a mapping from the gaze vector to the first feature in the eye image can be generated.
[0005] After the first calibration, in an eye tracking session, using the eye image, the first feature can be obtained, and using the first feature and the mapping from the first feature in the eye image to the gaze vector, the gaze vector can be obtained.
[0006] Perform a second eye tracking system calibration using a second feature in an eye image captured by an eye image sensor.
[0007] After the second calibration, a mapping from the second feature in the eye image to the gaze vector can be generated.
[0008] After the second calibration, in an eye tracking session, using the eye image, the second feature can be obtained, and using the second feature and the mapping from the second feature in the eye image to the gaze vector, the gaze vector can be obtained.
[0009] After calibration using the first feature and the second feature, in an eye tracking session, the gaze vector can be generated using the first feature, and at the same time, slippage can be detected and corrected using the second feature.
[0010] When using the second feature to detect and correct slippage in an eye tracking session, first obtain the second feature from the eye image. Then, use the mapping from the second feature in the eye image to the gaze vector to generate an expected gaze vector.
[0011] Using the expected gaze vector and the mapping from the gaze vector to the first feature in the eye image, the expected first feature in the same eye image can be obtained.
[0012] In the same eye image, obtain the actual first feature. Knowing both the expected first feature and the actual first feature from the same image, the difference between them can be obtained. This difference can be used to detect the amount of slippage.
[0013] By using this difference to offset the true first feature and obtaining a corrected first feature in each subsequent eye image, slippage can be corrected.
[0014] Using the corrected first feature and the mapping from the first feature in the eye image to the gaze vector, a corrected gaze vector can be generated from each subsequent eye image.
[0015] It is assumed that some eye parameters can be obtained during the calibration of the eye tracking system using the first or second feature. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A 3D coordinate system is shown.
[0017] Figure 2 Several 3D coordinate systems mentioned herein are shown.
[0018] Figure 3A A system with a head-mounted display is schematically shown.
[0019] Figure 3B A system with a head-mounted field camera is schematically shown.
[0020] Figure 4 The offset between the optical axis and the visual axis of the eye is shown.
[0021] Figure 5 An image of the pupil is schematically shown.
[0022] Figure 6 A data stream is shown. DETAILED DESCRIPTION
[0023] The gaze vector is a 3D vector that points from the center of the user's eye to the point the user is looking at. It can be represented in one of the Figure 2 coordinate systems shown (e.g., the field coordinate system CS-F). The gaze point is a physical point in the world or the image of a point on the display that the user is looking at.
[0024] In one embodiment, the user wears an eye image sensor in front of the eyes. The eye image sensor is used to detect the eye position by taking an external eye image. During calibration, the position and orientation of the eye image sensor relative to the user's head remain unchanged.
[0025] The system may also include a head-mounted display (as Figure 3A shown) or a head-mounted field camera (as Figure 3BAs shown). The head-mounted display can display images in front of the user's eyes so that the eyes can see those images. The head-mounted field camera can capture images of the area in front of the user. The head-mounted display or the head-mounted field camera has a fixed relative position and orientation with respect to the eye image sensor.
[0026] Slippage refers to a change in the relative position and / or orientation between the eye image sensor and the eye due to weight, movement, or other reasons, usually with a small change amount.
[0027] Calibration is to create a mapping between one or more features of the eye in the eye image and the gaze vector using calibration points in front of the user's eyes. For example, if the pupil position is used as a feature, the gaze vector can be obtained based on the pupil position and the mapping. In this case, calibration is to find the relationship between the feature position in the eye image and the gaze vector.
[0028] Before calibration, it is assumed that some eye parameters are known. These eye parameters can be used together with one or more eye features from the eye image to create a calibration result. These eye parameters can include the position of the eye center in one of the Figure 2 shown coordinate systems (such as the eye image sensor coordinate system CS-C), or the offset between the eye optical axis and the visual axis (as Figure 4 shown), or other information. During calibration, in one embodiment, the calibration point is a point displayed on the head-mounted display as Figure 3A shown. The calibration point is stationary relative to the user's eyes, and the image position of the calibration point can be obtained. In another embodiment, a physical point in front of the user is used as the calibration point. The physical point can be recorded by the head-mounted field camera (as Figure 3B shown), and the image position of the calibration point can be obtained. In both of these embodiments, one calibration point is used at a time during calibration. Multiple calibration points can be used at different time points. When the user views the calibration point, the calibration point is the fixation point.
[0029] During calibration, different sensors and methods can be used to obtain the gaze vector GazeVector based on the calibration point. In one embodiment, it is assumed that the calibration point is displayed on the head-mounted display, or its image is captured by the head-mounted field camera. The position of the calibration point image in the image plane of the head-mounted field camera or the head-mounted display is known. We can call this position CaliPosition. The confirmed gaze vector GazeVector can be obtained according to the position of the calibration point image as follows: GazeVector = CaliPosition_to_GazeVector(CaliPosition), see method C1 for details.
[0030] During calibration using the first feature, when the user views the calibration point, the calibration point image position can be obtained, and the first feature can be obtained in the eye image captured by the eye image sensor. By repeating this step multiple times, multiple pairs of calibration point image positions and the first feature are obtained. Using this data and the eye parameters, a mapping MAP_Feature1_to_GazeVector() from the first feature in the eye image to the gaze vector can be generated, as detailed in Methods C4 and C5. Moreover, using the same data, a mapping MAP_GazeVector_to_Feature1() from the gaze vector to the first feature in the eye image can be generated, as detailed in Methods C4 and C5.
[0031] After the first calibration, during the eye tracking session, the first feature Feature1 can be obtained from the eye image, as detailed in Method C2; using the first feature, the following can be used to obtain the gaze vector: GazeVector1 = MAP_Feature1_to_GazeVector(Feature1).
[0032] During calibration using the second feature, when the user views the calibration point, the second feature is obtained in the eye image captured by the eye image sensor. By repeating this step multiple times, multiple pairs of calibration point image positions and the second feature are obtained. Using this data and the eye parameters, a mapping MAP_Feature2_to_GazeVector() from the second feature in the eye image to the gaze vector can be generated.
[0033] After the second calibration, during the eye tracking session, the second feature Feature2 can be obtained from the eye image. Using the second feature, the following can be used to obtain the gaze vector: GazeVector2 = MAP_Feature2_to_GazeVector(Feature2).
[0034] When the second feature is used to detect and / or correct slippage in the eye tracking session, first the second feature Feature2 is obtained from the eye image. Then, the mapping from the second feature in the eye image to the gaze vector is used to generate the expected gaze vector: GazeVector2_d = MAP_Feature2_to_GazeVector(Feature2).
[0035] Using the expected gaze vector GazeVector2_d and the mapping from the gaze vector to the first feature in the eye image, the expected first feature in the eye image can be obtained as follows: Feature1_d = MAP_GazeVector_to_Feature1(GazeVector2_d).
[0036] Using the same eye image simultaneously, obtain the true first feature Feature1_r. Knowing both the desired first feature and the true first feature from the same image, the difference Dif_Feature1_dr between them can be obtained as follows: Dif_Feature1_dr = DIF_Feature1(Feature1_d, Feature1_r). In one embodiment, as Figure 5 shown, the pupil center position in the eye image is used as the first feature. If slippage occurs, the desired first feature will be the desired pupil center position, and the true first feature will be the true pupil center position. Due to slippage, these two positions in the eye image may be different.
[0037] The amount of the difference Dif_Feature1_dr can be used to detect the amount of slippage as follows: Slippage = Dif_Feature_to_Slippage(Dif_Feature1_dr), see Method C14 for details.
[0038] To correct for slippage, during the eye tracking session, use this difference Dif_Feature1_dr to offset the true first feature Feature1_s to obtain the corrected first feature Feature1_c as follows: Feature1_c = OFFSET_Feature1(Feature1_s, Dif_Feature1_dr), see Method C16 for details.
[0039] Using the corrected first feature Feature1_c and the mapping from the first feature in the eye image to the gaze vector, the corrected gaze vector GazeVector1_c can be generated as follows: GazeVector1_c = MAP_Feature1_to_GazeVector(Feature1_c).
[0040] For slippage correction, the same Dif_Feature1_dr can be used for each subsequent eye image to generate GazeVector1_c.
[0041] The first calibration and the second calibration can be performed simultaneously using the same set of calibration points. The first calibration and the second calibration can be performed at different times, or using different sets of calibration points, or both.
[0042] The first feature can be a more accurate and / or faster - processing feature. Some image features are more robust to noise, such as the average of pixel positions, and generate more stable and accurate results. In one embodiment, the position of the pupil center in the eye image can be used as the first feature. Examples can be found in Reference 2.
[0043] The second feature can be a feature that is insensitive to slippage and / or requires more processing time. In one embodiment, the contour and position of the pupil center in the eye image, or the shape of the iris, limbus, etc. can be used as the second feature. Examples can be found in Reference 3.
[0044] The mathematical utility functions used in this disclosure are listed in Part A2 of the appendix section.
[0045] The position, orientation, and pose of the 3D object in the 3D coordinate system are defined in A1.7. The head pose includes the head position and the head orientation.
[0046] The 3D coordinate system using the right-hand rule is defined in Part A1.1 of the appendix section. This disclosure refers to the following coordinate systems as Figure 2 shown. They are defined as follows:
[0047] Eye image sensor coordinate system Xc - Yc - Zc - Oc: CS - C is defined in A1.2
[0048] Field coordinate system Xf - Yf - Zf - Of: CS - F is defined in A1.3
[0049] Head coordinate system Xh - Yh - Zh - Oh: CS - H is defined in A1.4
[0050] World coordinate system Xw - Yw - Zw - Ow: CS - W.
[0051] System Hardware, Software, and Configuration
[0052] The imaging sensor measures at least the brightness of light. The imaging sensor can also measure the color of light. A camera is an imaging sensor. Other types of imaging sensors can be used here in a similar manner. The camera can be color, grayscale, infrared, or non - infrared, etc. The parameters of the camera include its physical size, resolution, and the focal length of the lens installed, etc. The imaging sensors mentioned in this disclosure can be color or grayscale cameras, or a single camera, a camera array, or a combination of different image sensing devices.
[0053] Headgear. A device for fixing the eye image sensor, head - mounted display, field camera, and other sensors and / or devices to the user's head. It can be a spectacle frame, headband, or helmet, etc., depending on the application.
[0054] Computer. The computer processes the output of the sensing unit and calculates the motion / gesture tracking results. It can either be an embedded system or a desktop system.
[0055] Software. Computer software implements algorithms and enables workflows.
[0056] As Figure 3A shown, in one embodiment, configured as an eye tracking system, an eye image sensor and a head-mounted display are mounted on a rigid platform and attached to a harness frame worn by a user. The eye image sensor is directed at the user's eyes. The position and orientation of the eye image sensor relative to the head-mounted display are invariant.
[0057] As Figure 3B shown, in another embodiment, configured as an eye tracking system, an eye image sensor and a head-mounted field camera are mounted on a rigid platform and attached to a head harness worn by a user. The eye image sensor is directed at the user's eyes. The position and orientation of the eye image sensor relative to the head-mounted field camera are invariant.
[0058] Method
[0059] C1. During calibration using the first feature, for each calibration point, a gaze vector can be obtained. In one embodiment, assuming that the calibration point image position CaliPosition is a two-dimensional (2D) vector in the image plane and GazeVector is a three-dimensional (3D) vector in CS-F, we have: GazeVector = CaliPosition_to_GazeVector(CaliPosition) = v_from_p(CaliPosition).
[0060] C2. During calibration using the first feature, when the user views the calibration point, obtain the gaze vector GazeVector. The first feature Feature1 is obtained from the eye image captured by the eye image sensor. In one embodiment, assuming that the position of the pupil center in the eye image is used as the first feature, the first feature can be represented as a two-dimensional vector.
[0061] C3. In each of the above steps, obtain a pair of gaze vector and the first feature. Repeat the above steps multiple times to obtain multiple pairs of gaze vectors and the first feature {GazeVector_i, Feature1_i}, i = 1…n.
[0062] C4. Using these data and eye parameters, a mapping from the first feature in the eye image to the gaze vector can be generated: GazeVector = MAP_Feature1_to_GazeVector(Feature1). And using the same data, a mapping from the gaze vector to the first feature in the eye image can be generated: Feature1 = MAP_GazeVector_to_Feature1(GazeVector). This mapping is used to transform variables from one space to another. Depending on the required accuracy, one or more samples are needed.
[0063] C5. The method for generating MAP_Feature1_to_GazeVector() and MAP_GazeVector_to_Feature1() can be one of many existing eye-tracking calibration methods, such as model-based algorithms, interpolation, or machine learning, etc. Examples can be found in Reference 1.
[0064] C6. During calibration using the second feature, when the user views the calibration point, the gaze vector GazeVector is obtained. At the same time, the second feature Feature2 is obtained in the eye image captured by the eye image sensor.
[0065] C7. Repeat the above steps multiple times to obtain multiple pairs of gaze vectors and second features {GazeVector_i, Feature2_i}, where i = 1…n.
[0066] C8. Using these data and eye parameters, a mapping from the second feature in the eye image to the gaze vector can be generated: GazeVector = MAP_Feature2_to_GazeVector(Feature2).
[0067] C9. The method for generating MAP_Feature2_to_GazeVector() can be one of many existing eye-tracking calibration methods, including model-based algorithms, interpolation, or machine learning, etc.
[0068] C10. When the second feature is used to detect and / or correct slippage in an eye-tracking session, the desired gaze vector is obtained as follows based on the second feature in the eye image: GazeVector2_d = MAP_Feature2_to_GazeVector(Feature2).
[0069] C11. And the desired first feature in the image can be obtained as follows: Feature1_d = MAP_GazeVector_to_Feature1(GazeVector2_d).
[0070] C12. Using the same eye image, obtain the true first feature Feature1_r. Given both the desired first feature and the true first feature in the same image, the difference between Feature1_d and Feature1_r can be obtained as follows: Dif_Feature1_dr = DIF_Feature1(Feature1_d, Feature1_r).
[0071] C13. In one embodiment, assume that the position of the pupil center in the eye image is used as the first feature. Thus, the first feature can be represented as a two-dimensional vector, where Feature1_d = (x1_d, y1_d) and Feature1_r = (x1_r, y1_r). The difference between Feature1_d and Feature1_r can be defined as: Dif_Feature1_dr = DIF_Feature1(Feature1_d, Feature1_r) = p_sub(Feature1_d, Feature1_r), see A2.2.2.
[0072] C14. Dif_Feature1_dr can be used to detect the slippage amount as follows: Slippage = Dif_Feature_to_Slippage(Dif_Feature1_dr). In one embodiment, Dif_Feature1_dr is a 2D vector, and its length is used to measure the slippage amount: Slippage = Dif_Feature_to_Slippage(Dif_Feature1_dr) = p_len(Dif_Feature1_dr) * k, where k can be a scale factor.
[0073] C15. When the slippage exceeds a certain threshold, the slippage is detected. In one embodiment, the threshold can be determined by the minimum error caused by the slippage that the system can accept.
[0074] C16. Dif_Feature1_dr can be used to correct slippage. First, the true first feature Feature1_s in each subsequent eye image is offset using Dif_Feature1_dr to obtain the corrected first feature Feature1_c as follows: Feature1_c = OFFSET_Feature1(Feature1_s, Dif_Feature1_dr). In one embodiment, Feature1_c, Feature1_s, and Dif_Feature1_dr are two-dimensional vectors, so it can be obtained that: Feature1_c = OFFSET_Feature1(Feature1_s, Dif_Feature1_dr) = p_add(Feature1_s, Dif_Feature1_dr), see A2.2.3.
[0075] C17. Using the corrected first feature Feature1_c and the mapping from the first feature in the eye image to the gaze vector, the corrected gaze vector GazeVector1_c can be generated as follows: GazeVector1_c = MAP_Feature1_to_GazeVector(Feature1_c).
[0076] C18. Figure 6 Shows the data flow from the eye image to GazeVector1_c when the second feature is used for slippage detection and correction and the first feature is used to generate the gaze vector.
[0077] Appendix
[0078] The mathematical tools listed in the appendix are used in the method section.
[0079] A1. Coordinate System
[0080] A1.1. The 3D coordinate system has three axes, X, Y, and Z, as Figure 1 shown. The right-hand rule applies to the order of the axes and the positive rotation direction.
[0081] A1.2. The sensing direction of the eye image sensor is forward. The x-axis of the eye image sensor coordinate system CS-C points to the right, the y-axis points to the top, and the z-axis points in the opposite direction of the lens. Its origin is located at the optical center of the camera.
[0082] A1.3. A field coordinate system CS-F is defined for both the head-mounted display and the head-mounted field camera. Its x-axis points to the right side of its image plane, and its y-axis points to the top of its image plane. For the head-mounted display, its z-axis points to the user's eyes, and its origin is located at the center of the eyeball. For the head-mounted field camera, its z-axis points in the opposite direction of the lens, and its origin is located at the optical center of the field camera.
[0083] A1.4. With the head facing forward, the x-axis of the head coordinate system points to the right side, the y-axis points to the top, and the z-axis points in the opposite direction of the nose.
[0084] A1.5. A point in 2D coordinates can be represented by a 2D vector p = (x, y). This vector is from the origin of the coordinate system to the position of the point.
[0085] A1.6. A point in 3D coordinates can be represented by a 3D vector v = (x, y, z). This vector is from the origin of the coordinate system to the position of the point.
[0086] A1.7. The position of a 3D object in a 3D coordinate system is described as (x, y, z). The orientation of a 3D object in a 3D coordinate system can be described by quaternions, 3x3 matrices, Euler angles, etc. The pose of a 3D object in a 3D coordinate system is described as its position and its orientation.
[0087] A1.8. The image plane is defined as a 2D coordinate system with its center as the origin. Points in the image plane are 2D vectors. The image plane can exist in CS-F, where their x-axis and y-axis are aligned. In one embodiment, the position of the center of the image plane is at the position (0, 0, -FOC_LEN) in CS-F, where FOC_LEN is called the focal length.
[0088] A2. 2D vectors, 3D vectors
[0089] A2.2.1 A 2D vector has two elements
[0090] p = (x, y)
[0091] A2.2.2 Subtraction of two 2D vectors pa = (xa, ya), pb = (xb, yb):
[0092] p_sub(x1, xb) = (xa - xb, ya - yb)
[0093] A2.2.3 Addition of two 2D vectors pa = (xa, ya), pb = (xb, yb):
[0094] p_add(x1, xb) = (xa + xb, ya + yb)
[0095] A2.2.4 The length of the 2D vector p = (x, y):
[0096] p_len(p) = sqrt(x * x + y * y)
[0097] A2.2.5 A 3D vector has three elements
[0098] v = (x, y, z)
[0099] A2.2.6 The length of the 3D vector v = (x, y, z):
[0100] V_len(v)) = sqrt(x * x + y * y + z * z)
[0101] 2.2.7 Convert the 3D vector v to the 2D vector p = (x, y):
[0102] v_from_p(p) = (x, y, FOCAL_LEN)
[0103] where FOCAL_LEN is the distance from the center of the 2D plane to the origin of the 3D coordinate system.
[0104] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for illustrative purposes and are not intended to be limiting, and the true scope and spirit are indicated by the appended claims.
[0105] References
[0106] The disclosure of each reference in this section is hereby incorporated by reference in its entirety.
[0107] Reference 1: US Patent Application No. 16 / 753,907.
[0108] Reference 2: US Patent Application No. 15 / 387,024.
[0109] Reference 3: Wang, Rui, et al., "Eye gaze tracking based on the shape of pupil image", 2017 International Conference on Optical Instruments and Technology: Optoelectronic Imaging / Spectroscopy and Signal Processing Technology, Vol. 10620, SPIE, 2018.
Claims
1. A method, comprising: When a person's eye gazes at a first point in the 3D space where the person is located, capturing a first image of the eye using a first imaging sensor and capturing a first reference image of the first point using a second imaging sensor, wherein the first imaging sensor and the second imaging sensor are stationary relative to the person's head; Based on the position of a first feature of the eye in the first image and the position of the first point in the first reference image, determining a first mapping between the gaze vector of the eye and the position of the first feature of the eye in the field of view of the first imaging sensor; When the eye gazes at a second point in the 3D space, capturing a second image of the eye using the first imaging sensor and capturing a second reference image of the second point using the second imaging sensor; Based on the position of a second feature of the eye in the second image and the position of the second point in the second reference image, determining a second mapping between the gaze vector and the position of the second feature of the eye in the field of view of the first imaging sensor.
2. A method, comprising: When a person's eye gazes at a first point on a display, capturing a first image of the eye using a first imaging sensor, wherein the first imaging sensor is stationary relative to the person's head, the display is stationary relative to the person's head, and the first point is stationary relative to the display; Based on the position of a first feature of the eye in the first image and the position of the first point on the display, determining a first mapping between the gaze vector of the eye and the position of the first feature of the eye in the field of view of the first imaging sensor; When the eye gazes at a second point on the display, capturing a second image of the eye using the first imaging sensor, wherein the second point is stationary relative to the display; Based on the position of a second feature of the eye in the second image and the position of the second point on the display, determining a second mapping between the gaze vector of the eye and the position of the second feature of the eye in the field of view of the first imaging sensor.
3. The method according to claim 1 or claim 2, wherein the first feature is the pupil center of the eye, and the second feature is the iris or limbus of the eye.
4. The method according to claim 1 or claim 2, wherein determining the first mapping or determining the second mapping is further based on the geometric characteristics of the eye.
5. The method according to claim 4, wherein the geometric characteristics include the position of the center of the eyeball of the eye, or the offset between the optical axis and the visual axis of the eye, or both.
6. A computer program product, comprising a non-transitory computer-readable medium having instructions recorded thereon, the instructions, when executed by a computer, implementing the method according to claim 1 or claim 2.
7. A method, comprising: Obtaining a first mapping between the gaze vector of a person's eye and the position of a first feature of the eye in the field of view of a first imaging sensor, wherein the first imaging sensor is stationary relative to the person's head; Obtaining a second mapping between the gaze vector and the position of a second feature of the eye in the field of view of the first imaging sensor; Determine the fixation vector based on the position of the second feature of the eye in the field of view of the second mapping and the first imaging sensor; Determine the predicted position of the first feature of the eye in the field of view of the first imaging sensor based on the fixation vector and the first mapping; Obtain the actual position of the first feature of the eye in the field of view of the first imaging sensor; Determine the slip of the first imaging sensor relative to the person's head based on the predicted position and the actual position; 8. The method according to claim 7, wherein the first feature is the pupil center of the eye and the second feature is the iris or limbus of the eye.
9. The method according to claim 7, further comprising: Determine the corrected position of the first feature of the eye in the field of view of the first imaging sensor based on the predicted position and the actual position; 10. The method according to claim 8 further comprises: Determine the corrected fixation vector of the eye based on the corrected position and the first mapping; 11. A computer program product comprising a non-transitory computer-readable medium having instructions recorded thereon, the instructions when executed by a computer implementing the method according to claim 7.
Citation Information
Patent Citations
Systems and methods for calibrating an eye tracking system
US11573630B2
Systems and methods for tracking motion and gesture of heads and eyes
US9785249B1