A line-of-sight calibration, motion tracking and accuracy testing method based on adaptive time series analysis prediction
By employing eye-controlled nine-point calibration and adaptive temporal analysis prediction algorithms, the problem of unstable gaze positioning during head movement is solved, achieving stable gaze tracking and real-time accuracy measurement. This makes it suitable for gaze calibration and motion tracking in various scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENYANG AIRCRAFT DESIGN & RES INST YANGZHOU COLLABORATIVE INNOVATION RES INST CO LTD
- Filing Date
- 2023-03-07
- Publication Date
- 2026-05-22
AI Technical Summary
Existing pupil-corneal gaze localization algorithms suffer from algorithm recognition errors and system random errors when processing single-frame images. This leads to unstable gaze localization during head movement, with the gaze point position fluctuating randomly, making it difficult to achieve real-time accuracy quantification.
An eye-controlled nine-point calibration method combined with an adaptive temporal analysis and prediction algorithm is adopted. The screen position is coarsely located through image color recognition technology, distortion calibration and mapping relationship determination are performed, the current gaze positioning is optimized by combining historical gaze information, and the parameters of the temporal analysis and prediction algorithm are adaptively adjusted to achieve stable tracking and real-time accuracy measurement.
It achieves stable tracking and real-time accurate measurement of the line of sight during head movement, without the need for additional equipment. It automatically calibrates and quantifies to obtain the line of sight positioning accuracy, and has real-time performance and stability.
Smart Images

Figure CN116382473B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction. Existing pupil-corneal gaze localization algorithms only process single-frame images. Due to factors such as algorithm recognition errors and system random errors, the gaze localization accuracy of a single frame image is not ideal. This leads to highly unstable gaze tracking in video streams, resulting in severe jitter in the gaze point position. To address this problem, this invention achieves automated calibration through eye-controlled nine-point calibration and employs an adaptive temporal analysis and prediction method to ensure stable motion gaze tracking. Furthermore, a quantitative accuracy testing method is proposed. Background Technology
[0002] In recent years, gaze localization has become a research hotspot and has received widespread attention. Gaze localization can be applied to multiple fields such as medicine, education, entertainment and military.
[0003] Hutchinson (Reference: Hutchinson. Human-computer interaction using eye-gaze input. IEEE Trans. on System, Man and Cybernetics, 1989, 19(6): 1527-1533) calculated the gaze direction using the vector formed by the center of the infrared light spot and the center of the pupil, but this method only applies to scenarios where the test subject's head is stationary. Morimoto et al. (Reference: Morimoto et al., Keeping an eye for HCI, in: Proc. on Computer Graphics and Image Processing, 1999, 171-176) expressed the relationship between the glint-pupil vector and the position of the display gaze point using a quadratic polynomial. This method works well when the user is seated and stable, but its accuracy is poor when the head moves. Zhang Dengyin (Reference: Zhang Dengyin, A gaze tracking and localization method based on iris recognition, Chinese invention patent, application patent number: CN103136519A) calculates the position of the test subject's gaze based on the current frame of human eye image captured by the camera, without considering the influence of head movement under continuous eye movement recognition.
[0004] The above methods mainly study how to accurately obtain the gaze point position in a single frame image. However, in engineering applications, the systematic error of gaze localization in a single frame will cause the gaze point position deviation to be random. When used for a long time, when the head moves continuously to gaze at the target, the gaze point position identified by the gaze localization algorithm will randomly jitter near the target, which seriously affects the effect of gaze localization and eye movement control.
[0005] Unlike the methods described above, when using a glasses-type eye tracker to control eye movement on a screen, this invention first uses color positioning to achieve real-time screen position locking in various indoor and outdoor scenarios, mapping the gaze point coordinates to the screen coordinates. Secondly, after completing the model mapping through nine-point calibration, accuracy testing is performed to quantitatively calculate the real-time accuracy of gaze positioning. Thirdly, an adaptive temporal analysis and prediction algorithm, combined with historical gaze information, optimizes the current gaze positioning and predicts the future gaze point position. Simultaneously, the parameters in the temporal analysis and prediction algorithm can be adaptively adjusted according to real-time accuracy or the severity of jitter. Finally, eye state recognition is used to obtain eye control information, achieving the desired eye movement control effect. Summary of the Invention
[0006] This invention achieves automated calibration through eye-controlled nine-point calibration, and realizes stable tracking of the gaze point during head movement based on an adaptive temporal analysis and prediction algorithm. It can automatically and in real-time quantify and obtain the gaze positioning accuracy without the need for any additional equipment. It solves the problems of random jitter of the gaze point and difficulty in real-time quantification and calculation of gaze positioning accuracy during head movement and gaze. The designed solution has the characteristics of real-time performance, stability and reliability.
[0007] The technical solution of the present invention:
[0008] A method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis prediction, comprising the following steps:
[0009] Step 1: Coarse screen position positioning: Using image color recognition technology, identify the color blocks and calibration point positions on the screen in the scene image, and combine this with the actual screen size to coarsely position the screen.
[0010] Step 2: Precise screen position positioning: Based on calibration point information, distortion calibration is performed to achieve precise screen position positioning and determine the mapping relationship between the scene image and the precisely positioned screen.
[0011] Step 3: Perform eye-controlled nine-point calibration to determine the mapping model between the physical information of the eyeball and the spatial position of the test subject's gaze. Combined with the mapping relationship in Step 2, obtain the coordinates of the screen gaze point.
[0012] Step 4: Perform a line-of-sight positioning accuracy quantification test. If the accuracy does not meet the requirements, return to step 3.
[0013] Step 5: Optimize the current gaze positioning and predict the future gaze point position by using an adaptive temporal analysis and prediction algorithm combined with historical gaze information. At the same time, the parameters in the temporal analysis and prediction algorithm can be adaptively adjusted according to the real-time accuracy or the severity of jitter.
[0014] Step 6: Detect whether the human eye is open or closed using a human eye state detection algorithm, and then perform eye movement control.
[0015] In step 1, the screen position is roughly located. In order to adapt to multiple application scenarios such as indoor and outdoor, bright and dark fields, the color recognition technology is used to identify the color blocks and calibration point positions in the scene image. Combined with the actual screen size, the screen position is roughly located, and the positions of nine calibration points with the same special color are located at the same time.
[0016] Step 2, screen position fine localization, mainly includes two parts: First, screen position fine localization technology based on screen distortion calibration, which calculates distortion and updates the precise screen position using nine equally spaced calibration points. Second, it calculates the mapping relationship between the scene image and the finely positioned screen using homography.
[0017] In step 3, the eye-controlled nine-point calibration requires the subject to observe calibration points appearing at specific locations on the screen according to system instructions. The coordinates of these calibration points in the screen coordinate system are known and fixed. Combining this with the scene image-screen coordinate transformation obtained in step 2, the coordinate positions of the calibration points in the scene image coordinate system can be obtained. By mapping the pupil-corneal reflection vectors corresponding to these calibration points to the coordinates of the scene image calibration points, the mapping relationship between the pupil-corneal reflection vectors and the scene image coordinate system is obtained. During the subsequent actual fixation process, the pupil-corneal reflection vector corresponding to the fixation point is first calculated. Then, using the obtained mapping relationship between the pupil-corneal reflection vectors and the scene image, the coordinates of the fixation point in the scene image are calculated. Finally, the position of the fixation point on the display screen is obtained through the foreground image-display screen coordinate transformation.
[0018] In step 4, the gaze positioning accuracy is quantitatively measured. First, an accuracy test image is generated, which can be a test point or a test line. Then, the gaze is directed onto the test image in sequence. During the test, the system software collects the coordinates of the gaze point on the screen. At the same time, without adding any other equipment, the vertical distance between the eye and the screen is calculated solely based on the principle of similar triangles in imaging characteristics. Finally, the gaze positioning accuracy is statistically calculated. Combined with the distance between the coordinates of the target test point and the coordinates of the gaze point, the deviation angle of gaze positioning is calculated, and the root mean square error of the deviation angle is calculated as the actual error.
[0019] In step 5, the gaze point coordinate anti-shaking mechanism based on the adaptive temporal analysis and prediction algorithm filters the historical gaze point coordinates by multiplying them by weights, smoothing the current gaze point coordinates and predicting the gaze point coordinates for the next moment. The adaptive aspect is reflected in the automatic adjustment of the weights. A queue is used to store the gaze point coordinates from the previous m moments. The presence of a target button is determined by the gaze point coordinates. If a target button is present, the weights are adjusted based on the error accuracy calculated in step 4. If no target button is present, the weights are adjusted based on the severity of the gaze point jitter.
[0020] In eye-tracking control software systems, adaptive temporal analysis and prediction algorithms are mainly applied in the following two aspects: first, obtaining the gaze coordinates in the scene image coordinate system using the pupil-corneal reflection gaze localization algorithm; and second, obtaining the gaze coordinates in the screen coordinate system after the scene-screen coordinate system transformation in step 3.
[0021] In step 6, the human eye state detection first extracts the edge, then fits the edge curve to obtain the user's eye contour curve, and finally selects the state of open eye, start blinking and close eye, closed eye, and end blinking and close eye by the eye pixel area, and then calculates the duration of eye closure and the number of blinks, which facilitates the design of subsequent eye movement control commands.
[0022] The beneficial effects of this invention are:
[0023] 1. This invention proposes an eye-controlled nine-point calibration method that eliminates the need for additional personnel to assist in clicking calibration points with a mouse, achieving fully automatic and intelligent eye-tracking calibration;
[0024] 2. The algorithm of this invention is stable in real time and can achieve fast and stable gaze point tracking and prediction even during continuous head movements;
[0025] 3. This invention provides a convenient and quick method for quantifying and measuring line-of-sight positioning accuracy. It requires no additional equipment and can detect line-of-sight positioning accuracy in real time, which is beneficial for optimizing line-of-sight positioning algorithms. Attached Figure Description
[0026] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.
[0027] Figure 1 Flowchart of a line-of-sight calibration, tracking, and testing method based on adaptive temporal analysis prediction
[0028] Figure 2 For eye-controlled nine-point calibration screen display
[0029] Figure 3 A schematic diagram of a distorted screen in a scene image.
[0030] Figure 4 A diagram illustrating nine-point calibration and eye movement.
[0031] Figure 5 Schematic diagram of test points for quantifying line-of-sight positioning accuracy
[0032] Figure 6 Schematic diagram of test line for quantifying line-of-sight positioning accuracy
[0033] Figure 7 A schematic diagram of the method for quantifying line-of-sight positioning accuracy.
[0034] Figure 8Diagram of the human eye blinking sequence
[0035] Figure 9 Schematic diagram of edge extraction for human eye state detection Detailed Implementation
[0036] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0037] The flowchart of the method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis prediction is as follows: Figure 1 As shown. Considering that eye trackers are worn on the user's head and move with the head, and that eye trackers are generally more accurate than desktop eye trackers, this invention uses an eye tracker for illustration. Eye trackers are typically equipped with three high-definition cameras: one scene camera to capture images of the user's field of vision, and two eye cameras to acquire information from the left and right eyes.
[0038] Firstly, while eye-tracking devices can directly output the coordinates of the gaze point in the scene image, these coordinates need to be converted to screen coordinates for screen control. Secondly, conventional eye-tracking calibration requires additional personnel to click on the target calibration point with a mouse. Thirdly, measuring the current gaze positioning accuracy is complex and difficult to quantify, lacking a method that allows for convenient real-time quantification of gaze positioning accuracy without additional equipment. Finally, current mainstream pupil-corneal reflex gaze positioning algorithms only acquire the gaze point at the current moment. During head movements, due to systematic and random errors in the algorithm, the gaze point may jitter near the target, making it difficult to stably lock onto the target. To overcome these problems, we propose a method combining image recognition and temporal analysis. First, the scene image color blocks are used to locate the screen. Then, an automated calibration method using eye-controlled nine-point calibration is performed. Next, the gaze positioning accuracy is quantitatively measured. Finally, an adaptive temporal analysis prediction algorithm is used to optimize the current gaze point position and predict the future gaze point position.
[0039] A method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis prediction, comprising the following steps:
[0040] Step 1: Coarse screen position positioning. Using image color recognition technology, the color blocks and calibration point positions on the screen are identified in the scene image. Combined with the actual screen size, the screen position is coarsely positioned.
[0041] Step 2: Fine-tuning the screen position. Based on the calibration point information, distortion calibration is performed to achieve fine-tuning of the screen position and to determine the mapping relationship between the scene image and the finely calibrated screen.
[0042] Step 3: Eye-controlled nine-point calibration, determining the mapping model between the physical information of the eyeball and the spatial position of the test subject's gaze, and combining the information from Step 2 to obtain the coordinates of the screen gaze point;
[0043] Step 4: Quantitatively measure the line-of-sight positioning accuracy. If the accuracy does not meet the requirements, return to Step 3.
[0044] Step 5: Anti-shake gaze positioning based on adaptive temporal analysis and prediction algorithm. By combining historical gaze information, the current gaze positioning is optimized and the future gaze point position is predicted. At the same time, the parameters in the temporal analysis and prediction algorithm are adaptively adjusted according to the real-time accuracy or the severity of shaking. The gaze point position can be stably obtained when staring at the target during head movement, effectively preventing shaking.
[0045] Step 6: Detect whether the human eye is open or closed using a human eye state detection algorithm, and then perform eye movement control.
[0046] The screen position coarse localization technique based on image color recognition in step 1 is as follows:
[0047] To adapt to various application scenarios, including indoor and outdoor environments, and bright and dark fields, the coarse screen position localization technology based on image color recognition mainly includes the following steps:
[0048] 1) When selecting the HSV variation range of the color block, the color difference caused by changes in light and dark in the usage scenario must be taken into account;
[0049] 2) Convert the scene image to HSV format, and extract the specified color graphic through the HSV range;
[0050] 3) Binarize the image and remove salt-and-pepper noise;
[0051] 4) Calculate the center position of the color block based on the binarized image;
[0052] 5) Based on the actual screen size, calculate the pixel coordinates of the four vertices of the screen in the scene image to achieve coarse positioning.
[0053] Specifically, in step 1), during the design phase, select a color block of a special color and place it in the upper left corner of the screen; this could also be an icon or logo. Consider the light and dark variations in the usage scenario and select the range of HSV variations for the special color.
[0054] Specifically, in step 3), the image is binarized and salt-and-pepper noise is removed by Gaussian filtering or morphological filtering, with the core function being the morphologyEx function.
[0055] Specifically, in step 4), the pixel area s and center position of the color block are calculated based on the binarized image. The cv2.moments() function can obtain the image distance, and then the centroid is calculated through the image distance.
[0056] Specifically, in step 5), the actual area S of the color block is combined with the actual area M of the screen, and the pixel area of the screen is m = s * M / S. Contour extraction is performed on the scene image, and only the corner points near the screen pixel area can be considered as screen vertices. Thus, the pixel coordinates of the four vertices of the screen in the scene image can be roughly obtained, achieving coarse positioning.
[0057] The location of the nine calibration points is also determined by using special color recognition to obtain their specific positions in the scene image.
[0058] Step 2 consists of two parts: one is the screen position fine positioning technology based on screen distortion calibration, and the other is calculating the mapping relationship between the scene image and the finely positioned screen.
[0059] Screen position precision positioning technology based on screen distortion calibration:
[0060] The screen precision positioning step is completed within the eye-tracking nine-point calibration interface, such as... Figure 2 As shown. Although this system requires the user's head to be perpendicular to the screen, head movements make it impossible to be perfectly perpendicular. Therefore, the screen will inevitably have some distortion in the scene image and will not be a perfectly square rectangle. Screen distortion calibration calculates the screen distortion in the scene image using nine equally spaced points. The positions of the nine calibration points are determined using special color recognition to obtain their specific locations in the scene image, and then the distances between the calibration points are calculated. For example... Figure 3 As shown, the actual distance L1 between calibration point 1 and calibration point 3 corresponds to a pixel distance l1 in the scene image. The actual screen size is M*N mm, so the screen's pixel length in the scene image is l1*M / L1. Similarly, the actual distance L2 between calibration point 1 and calibration point 7 corresponds to a pixel distance l2 in the scene image, so the left pixel width of the screen in the scene image is l2*N / L2. The actual distance L3 between calibration point 3 and calibration point 9 corresponds to a pixel distance l3 in the scene image, so the right pixel width of the screen in the scene image is l3*N / L3. The top-left vertex of the screen is located by the center of the color block. The other three vertices of the screen in the scene image are located using a combination of coarse and fine positioning. Thus, the screen's position in the scene image is accurately located. When the screen does not use the nine-point calibration image in subsequent iterations, the screen's positioning uses coarse positioning in real-time, combined with fine positioning parameters to obtain the screen's position in the scene image in real-time.
[0061] Mapping relationship between scene image and precise positioning screen:
[0062] Once the screen position can be accurately located, image processing algorithms (homography) can be used to determine the coordinate transformation relationship between the screen and the scene image. This coordinate transformation relationship allows for the conversion of calibration point coordinates between the screen coordinate system and the scene image coordinate system. It also allows for the subsequent conversion of the gaze point coordinates on the scene image to coordinates in the screen coordinate system. Homography transformation is a two-dimensional projection transformation that maps a point in one plane to another. The `findHomography()` function in the OpenCV library can calculate the homography matrix. Thus, the mapping relationship between the scene coordinate system `oxy` and the screen coordinate system `o1x1y1` is obtained, as shown below. Figure 3 As shown.
[0063] The eye-controlled nine-point calibration technology described in step 3:
[0064] The eye-tracking software uses the pupil-corneal reflex method for gaze localization. The software SDK interface outputs coordinates of the pupil center and corneal reflex point. Through an eye-controlled nine-point calibration program, it finds the mapping function between the pupil-corneal vector and the gaze point in the scene image. Then, by detecting changes in the pupil-corneal vector, it tracks the user's gaze position on the screen in real time. Because each person's eye structure is unique, using the same parameters for all individuals will result in poor performance. To address this, eye-controlled nine-point calibration is necessary. The eye-controlled nine-point calibration model includes the mapping relationship between the eye's gaze projection point and the eye's rotation angle, such as... Figure 4 As shown.
[0065] Conventional nine-point calibration procedures are cumbersome and complex, requiring an additional operator to click on calibration points on the screen using a mouse within the scene image. Furthermore, inaccurate mouse clicks can easily introduce system errors. The eye-controlled nine-point calibration method proposed in this invention requires no additional personnel or equipment. It achieves eye-controlled nine-point calibration solely through the relationship between the scene image coordinate system and the screen coordinate system obtained in step 2.
[0066] During eye-controlled nine-point calibration, the subject needs to observe points appearing at specific locations on the screen according to system instructions. These points are called calibration points, and a nine-point calibration is typically used. Figure 2As shown. For each calibration point, its coordinates in the screen coordinate system are known and fixed. Combining the scene image-screen coordinate transformation relationship obtained in step 2, the coordinate position of the calibration point in the scene image coordinate system can be obtained. By mapping the pupil-corneal reflection vectors corresponding to these calibration points with the coordinates of the scene image calibration points, the mapping relationship between the pupil-corneal reflection vectors and the scene image coordinate system is obtained. In the subsequent actual fixation process, the pupil-corneal reflection vector corresponding to the fixation point is first calculated. Then, by using the obtained mapping relationship between the pupil-corneal reflection vectors and the scene image, the coordinates of the fixation point in the scene image are calculated. Finally, the position of the fixation point on the display screen is obtained through the foreground image-display screen coordinate transformation relationship.
[0067] Eye-tracking nine-point calibration specifically includes the following steps:
[0068] (1) The pupil-corneal vector V(V) can be obtained by processing the eye diagram using the pupil-corneal reflection gaze localization algorithm. x V y The coordinates of the calibration point P in the scene image coordinate system. xi ,P yi ), i = 1, 2, ..., 9, establish a mapping model between the pupil-cornea vector and the scene image:
[0069]
[0070] The unknowns a0~a5 and b0~b5 are calculated by eye-controlled nine-point calibration.
[0071] (2) The tester looks at the nine calibration points on the screen in sequence to obtain the pupil-corneal vector V. i =(V xi V yi ), i = 1, 2, ..., 9.
[0072] (3) V i Substitute the coordinates of the gaze point on the screen (in the scene coordinate system) for each time into formula (1).
[0073] (4) From this, we can obtain the following 9 sets of equations.
[0074]
[0075] To ensure high accuracy in eye-tracking, its minimum value is taken:
[0076]
[0077]
[0078] For R 2 By taking the partial derivative, we can obtain the parameters a0 to a5. Similarly, we can obtain the parameters b0 to b5.
[0079] Thus, a mapping model between the pupil-corneal vector and the scene image is obtained. In subsequent use, as long as the eye map of each frame is processed, the gaze coordinates in the scene image coordinate system can be obtained based on this model, and then the gaze coordinates in the screen coordinate system can be obtained through the coordinate transformation in step 2.
[0080] The line-of-sight positioning accuracy quantification measurement technology in step 4:
[0081] The quantitative measurement of line-of-sight positioning accuracy mainly includes three steps:
[0082] 1) Generate accuracy test images
[0083] 2) Test data collection
[0084] 3) Statistical calculation of line-of-sight positioning accuracy
[0085] Step 1) generates test images for accuracy verification. There are two types of test images: explicit fixation points and fixation trajectories. Both can be generated using code. Keypoints and key trajectories generated via code have more accurate real-world coordinates. For example... Figure 5 As shown, this is a test image containing 8 key points and their coordinates. Figure 6 The image shown contains a semicircle and a horizontal line segment.
[0086] Step 2) Test data collection.
[0087] The tester starts the system software and performs eye-controlled nine-point calibration. The system then automatically proceeds to the gaze positioning accuracy measurement stage, generating test images. The tester then sequentially focuses on key points, key line segments, or key curves within these images. During the test, the system software collects the coordinates of the gaze point on the screen and calculates the vertical distance between the eyes and the screen.
[0088] Traditional methods for calculating the vertical distance between the eyes and the screen involve mounting a distance measurement device on the test subject's head. However, to avoid the device interfering with the image captured by the eyes, it should be kept out of the viewfinder of the binocular data acquisition sensor and the scene camera. When using an additional distance measurement device, it is crucial to ensure that the distance measured is the distance D between the test subject's eyes and the gaze plane. If using a laser rangefinder, it is essential to ensure that the laser beam returns through the gaze plane. This method has several limitations and is not practically feasible. Therefore, a distance measurement method based on imaging characteristics is proposed.
[0089] The distance measurement method based on imaging characteristics is mainly obtained through the principle of similar triangles. The resolution of the screen in the scene image is m*n, the actual size of the screen is M*N mm, and the camera focal length F is a known quantity. Then the distance D between the screen and the scene camera is D = M*F / m.
[0090] Step 3) Calculation of gaze positioning accuracy. During test data acquisition, the system software will directly output the corresponding gaze point coordinates. However, for the eye control system, the eye control calibration data recorded by the system software needs to be used as the calibration data for eye control calibration. Then, the binocular data recorded by the system software is used as the input data for the eye control system, and the eye control system outputs the corresponding gaze point coordinates.
[0091] like Figure 7 As shown, the dots represent the actual output gaze point coordinates (x1, y1) and the target test point coordinates (x2, y2). Therefore, the actual distance between the gaze point and the target test point is... The vertical distance D between the human eye and the screen is obtained from step 2) and is in mm. Therefore, the deviation angle for line-of-sight positioning is α = tan -1 (L / D).
[0092] Since the data collected during the test is sequential data, the root mean square error of the deviation angle needs to be calculated as the actual error when calculating the accuracy. Let the coordinate sequence of the gaze point output by the system software or eye control system be Points, and the corresponding deviation angle sequence be P. Then the root mean square error of the accuracy can be calculated using the following formula:
[0093]
[0094] The adaptive time series analysis and prediction technique in step 5:
[0095] The gaze point coordinate anti-shake mechanism based on the adaptive temporal analysis and prediction algorithm filters the historical gaze point coordinates by multiplying them by weights, smoothing the current gaze point coordinates and predicting the gaze point coordinates for the next moment. The adaptive aspect is reflected in the automatic adjustment of the weights, which changes according to the severity of historical gaze point jitter or the accuracy of the error between the historical gaze point coordinates and the button to be controlled.
[0096] There are many methods for time series prediction, such as simple translation, simple averaging, moving average, weighted moving smoothing, simple smoothing, and Holt linear trend method. Taking weighted moving smoothing as an example, the entire eye-tracking software system uses a two-layer weighted moving smoothing method: first, it obtains the gaze coordinates in the scene image coordinate system using the pupil-corneal reflection gaze localization algorithm; second, it obtains the gaze coordinates in the screen coordinate system after the scene-screen coordinate system transformation in step 3.
[0097] In the weighted moving average smoothing method, different weights are assigned to the gaze coordinates at historical moments to predict the value at the current moment. The gaze coordinates at the current moment i (x...) i y i The calculation formula is as follows:
[0098] x i =ω1*x i-1 +ω2*x i-2 +ω3*x i-3 +…+ω m *x i-m
[0099] y i =ω1*y i-1 +ω2*y i-2 +ω3*y i-3 +…+ω m *y i-m
[0100] Among them, (x i-1 y i-1 ) is the coordinate of the gaze point at time i preceding time i, and so on (x i-m y i-m ) represents the coordinates of the gaze point m time steps before time i, and ω 1...m These are the weights at corresponding times, ω1+ω2+ω3+…+ω m =1. The initial value of the weight can be set to ω. i =0.5 i The adaptive weighting is based on the severity of historical gaze point jitter or the accuracy of the error between the historical gaze point coordinates and the target button. For example... Figure 1 The gaze coordinates of the previous m time steps are stored in a queue. The presence of a target button is determined by the gaze coordinates. If a target button is present, the weight is adjusted based on the error precision calculated in step 4. If no target button is present, the weight is adjusted based on the intensity of gaze jitter. The specific weight adjustment method is as follows:
[0101] The severity of historical gaze jitter is calculated as follows:
[0102]
[0103] Where, d i-1 It is the distance between the coordinates of the gaze point at time i-1 and time i-2, and so on, d i-mIt is the distance between the gaze point coordinates at time im and time im-1. After obtaining all the jitter levels α for the first m times, the gaze point coordinates at the time with the largest jitter level are removed as outliers, that is, the weight is reset to 0, and then the weights are redistributed to increase the weights of other times with smaller jitter levels.
[0104] Method for calculating the error accuracy between historical gaze point coordinates and the target button:
[0105] The method for calculating the accuracy error between the current gaze point coordinates and the target button is the same as in step 4, except that the accuracy verification point in step 4 is replaced with the target button. After obtaining all the accuracy errors for the first m time steps, the gaze point coordinates at the time with the largest accuracy error are treated as outliers and removed, that is, the weight is reset to 0. Then the weights are redistributed, and the weights of other time steps with smaller accuracy errors are increased.
[0106] The human eye state detection technology described in step 6:
[0107] After obtaining stable gaze coordinates, eye-tracking control of buttons on the screen is required. This necessitates other information about the eyes, such as whether they are open or closed, the duration of eye closure, and the number of blinks. However, the raw data for this information is obtained through human eye state detection.
[0108] An eye-tracking camera is used to capture images of the user's eyes; a sequence of blinking images represents one blink. Figure 8 As shown. A complete blink sequence includes the following states: 1-eye open, 2-begin blinking (closing eyes), 3-eye closed, 4-end blinking (closing eyes). A complete blink sequence may contain multiple consecutive identical states, but it must always cycle through the four states 1, 2, 3, and 4.
[0109] The effect of human eye condition detection technology is as follows Figure 9 As shown, edge extraction is first performed using methods such as the Canny operator, Sobel operator, or Robert operator. Then, the edge curve is fitted to obtain the user's eye contour curve. Finally, the eye pixel area is used to select the eye state: open, starting to blink (eyes closed), closed, and ending to blink (eyes closed). Considering the changes in eye pixel distance caused by head movement, the eye pixel area is normalized. The eye state values are as follows:
[0110]
[0111] Where z i It is the pixel area of the human eye, and the size of the human eye image is N×M.
[0112] When the human eye state value f mn When the eye opening threshold f2 is exceeded, the eye is determined to be in an open state; when f mnWhen the value is below the eye-closing threshold f1, the eyes are considered to be in a closed state. mn When the eye is between f1 and f2, it is necessary to combine historical information to determine whether it is the beginning or end of the blinking / closing state. If the previous different state of the human eye was the open state, then the current state of the human eye is the beginning of blinking / closing; if the previous different state of the human eye was the closed state, then the current state of the human eye is the end of blinking / closing.
[0113] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction, characterized in that, The steps are as follows: Step 1: Coarse screen position positioning: Using image color recognition technology, identify the color blocks and calibration point positions on the screen in the scene image, and combine this with the actual screen size to coarsely position the screen. Step 2: Precise screen position positioning: Based on calibration point information, distortion calibration is performed to achieve precise screen position positioning and determine the mapping relationship between the scene image and the precisely positioned screen. Step 3: Perform eye-controlled nine-point calibration to determine the mapping model between the physical information of the eyeball and the spatial position of the test subject's gaze. Combined with the mapping relationship in Step 2, obtain the coordinates of the screen gaze point. Step 4: Perform a line-of-sight positioning accuracy quantification test. If the accuracy does not meet the requirements, return to step 3. Step 5: Optimize the current gaze positioning and predict the future gaze point position by using an adaptive temporal analysis and prediction algorithm combined with historical gaze information. At the same time, the parameters in the temporal analysis and prediction algorithm can be adaptively adjusted according to the real-time accuracy or the severity of jitter. Step 6: Detect whether the human eye is open or closed using a human eye state detection algorithm, and then perform eye movement control.
2. The method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction according to claim 1, characterized in that, Step 1 The specific steps are as follows: 1) When selecting the HSV variation range of the color block, the color difference caused by changes in light and dark in the usage scenario must be taken into account; 2) Convert the scene image to HSV format, and extract the specified color graphic through the HSV range; 3) Binarize the image and remove salt-and-pepper noise; 4) Calculate the center position of the color block based on the binarized image; 5) Based on the actual screen size, calculate the pixel coordinates of the four vertices of the screen in the scene image to achieve coarse positioning.
3. The method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction according to claim 1, characterized in that, Step 2 consists of two parts: first, a screen position fine positioning technology based on screen distortion calibration, which calculates distortion and updates the precise screen position using nine equally spaced calibration points; second, a mapping relationship between the scene image and the finely positioned screen is calculated using homography. Screen position precision positioning technology based on screen distortion calibration: Screen distortion calibration calculates the screen distortion in the scene image using nine equally spaced points. The positions of these nine calibration points are determined using special color recognition to pinpoint their exact locations within the scene image, followed by distance calculations between the calibration points. The actual distance L1 between calibration point 1 and calibration point 3 corresponds to the pixel distance l1 in the scene image, and the actual screen size is M*N. If the distance between calibration points is mm, then the screen's pixel length in the scene image is l1*M / L1; similarly, the actual distance L2 between calibration points 1 and 7 is the pixel distance l2 in the scene image, so the left width of the screen in the scene image is l2*N / L2; the actual distance L3 between calibration points 3 and 9 is the pixel distance l3 in the scene image, so the right width of the screen in the scene image is l3*N / L3. The top left vertex of the screen is located by the center of the color block, and the other three vertices of the screen in the scene image are located by a combination of coarse and fine positioning. At this point, the screen's position in the scene image has been accurately located. When the screen does not use the nine-point calibration image in the future, the screen's positioning will use coarse positioning in real time, combined with the parameters of fine positioning, to obtain the screen's position in the scene image in real time. Mapping relationship between scene image and precise positioning screen: Once the screen position is accurately located, a homography image processing algorithm is used to determine the coordinate transformation relationship between the screen and the scene image. This coordinate transformation relationship is then used to convert the coordinates of the calibration point between the screen coordinate system and the scene image coordinate system, thus obtaining the mapping relationship between the scene image coordinate system oxy and the screen coordinate system o1x1y1.
4. The method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction according to claim 1, characterized in that, In step 3, eye-controlled nine-point calibration: the subject needs to observe the calibration points appearing at specific positions on the screen according to the system instructions; the coordinates of the calibration points on the screen coordinate system are known and fixed. Combining the scene image-screen coordinate transformation relationship obtained in step 2, the coordinate positions of the calibration points on the scene image coordinate system can be obtained. By mapping the pupil-corneal reflection vectors corresponding to these calibration points with the coordinates of the scene image calibration points, the mapping relationship between the pupil-corneal reflection vectors and the scene image coordinate system can be obtained. In the subsequent actual fixation process, the pupil-corneal reflection vector corresponding to the fixation point is first calculated. Then, the coordinates of the fixation point on the scene image are calculated by the mapping relationship between the obtained pupil-corneal reflection vector and the scene image. Finally, the position of the fixation point on the display screen is obtained by the coordinate transformation relationship between the foreground image and the display screen.
5. The method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction according to claim 4, characterized in that, The eye-controlled nine-point calibration specifically includes the following steps: (1) The pupil-corneal vector V(V) can be obtained by processing the eye diagram using the pupil-corneal reflection gaze localization algorithm. x V y The coordinates of the calibration point P in the scene image coordinate system. xi ,P yi ), i = 1, 2, ..., 9, establish a mapping model between the pupil-cornea vector and the scene image: The unknowns a0~a5 and b0~b5 are calculated by eye-controlled nine-point calibration; (2) The tester looks at the nine calibration points on the screen in sequence to obtain the pupil-corneal vector V. i =(V xi V yi ), i = 1, 2, ..., 9; (3) V i And the coordinates of the gaze point on the screen in the scene image coordinate system (V) each time. xi V yi Substitute into formula (1); (4) From this, we can obtain the following 9 sets of equations. To ensure high accuracy in eye-tracking, its minimum value is taken: For R 2 By taking the partial derivative, we can obtain the parameters a0 to a5. Similarly, we can obtain the parameters b0 to b5. Thus, a mapping model between the pupil-corneal vector and the scene image is obtained. In subsequent use, as long as the eye map of each frame is processed, the gaze coordinates in the scene image coordinate system can be obtained based on this model, and then the gaze coordinates in the screen coordinate system can be obtained through the coordinate transformation in step 2.
6. The method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction according to claim 1, characterized in that, In step 4, the gaze positioning accuracy is quantitatively measured as follows: First, an accuracy test image is generated, which can be a test point or a test line; then, the gaze is directed to the test image in sequence. During the test, the system software collects the coordinates of the gaze point on the screen and calculates the vertical distance between the eye and the screen using only the principle of similar triangles based on imaging characteristics; finally, the gaze positioning accuracy is statistically calculated, and the deviation angle of gaze positioning is calculated by combining the distance between the coordinates of the target test point and the coordinates of the gaze point, and the root mean square error of the deviation angle is calculated as the actual error.
7. The method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction according to claim 6, characterized in that, Method for calculating the vertical distance between the eye and the screen: The resolution of the screen in the scene image is m*n, the actual size of the screen is M*N mm, and the camera focal length F is a known quantity. Then the distance between the screen and the scene camera is D = M*F / m. Line-of-sight positioning accuracy statistical calculation: The dots represent the actual output gaze point coordinates (x1, y1) and the target test point coordinates (x2, y2). Therefore, the actual distance between the gaze point and the target test point is... Therefore, the deviation angle for line-of-sight positioning is α = tan -1 (L / D); When calculating accuracy, the root mean square error of the deviation angle needs to be calculated as the actual error. Let the coordinate sequence of the fixation point be Points, and its corresponding deviation angle sequence be P. The root mean square error of the accuracy can be calculated using the following formula:
8. The method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction according to claim 1, characterized in that, In step 5, the gaze point coordinate anti-shake mechanism based on the adaptive temporal analysis and prediction algorithm filters the historical gaze point coordinates by multiplying them by weights, making the current gaze point coordinates smooth and predicting the gaze point coordinates at the next moment. The adaptiveness is reflected in the automatic adjustment of the weight changes. The gaze point coordinates of the previous m moments are stored in a queue. The presence of a target button is determined by the gaze point coordinates. If a target button is present, the weight is changed based on the error accuracy calculated in step 4. If no target button is present, the weight is changed based on the severity of the gaze point jitter. The gaze point coordinates at the moment with the maximum jitter or the maximum accuracy error are treated as outliers and removed, that is, the weights are reset to 0. Then the weights are redistributed to increase the weights of other moments with less jitter or less accuracy error. In eye-tracking control software systems, adaptive temporal analysis and prediction algorithms are mainly applied in the following two aspects: first, obtaining the gaze coordinates in the scene image coordinate system using the pupil-corneal reflection gaze localization algorithm; and second, obtaining the gaze coordinates in the screen coordinate system after the scene-screen coordinate system transformation in step 3.
9. The method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction according to claim 8, characterized in that, The specific methods for changing the weights are as follows: The severity of historical gaze jitter is calculated as follows: Where, d i-1 It is the distance between the coordinates of the gaze point at time i-1 and time i-2, and so on, d i-m It is the distance between the gaze point coordinates at time im and time im-1.
10. The method for gaze calibration, motion tracking, and accuracy testing based on adaptive temporal analysis and prediction according to claim 1, characterized in that, In step 6, the human eye state detection first involves edge extraction, then fitting the edge curve to obtain the user's eye contour curve, and finally selecting the state of open eyes, blinking, closed eyes, and blinking / closed eyes by the eye pixel area, thereby calculating the duration of eye closure and the number of blinks, which facilitates the design of subsequent eye movement control commands.