Near-to-eye fixation point tracking method and device based on event sensor and medium
Through the near-eye gaze point tracking method based on event sensors, combined with infrared illumination, time window clustering and vector computing module and other technologies, the accuracy and real-time problems of traditional eye tracking methods in high dynamic scenarios are solved, and the eye tracking effect with high accuracy, low latency and strong environmental adaptability is achieved.
Patent Information
- Application Number
- CN202510444657.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Traditional eye tracking methods have problems with reduced data acquisition accuracy and reliability in high dynamic scenarios, and the overall real-time performance of hybrid eye tracking methods is not as good as that of pure event camera systems.
The near-eye gaze point tracking method based on event sensor is adopted to provide infrared illumination through infrared light sources. The event sensor collects eye information, combines time window clustering, ROI and pupil calculation module, vector calculation module and dynamic filtering method to calculate and update the pupil ellipse information, and finally obtains the gaze point coordinates through deep learning model calculation.
Eye tracking with high precision, low latency and strong environmental adaptability is achieved, overcome the shortcomings of traditional methods in high dynamic scenarios, and improve the real-time and computing efficiency of the system.
Smart Images

Figure CN119960605A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a field, and in particular to a near-eye gaze point tracking method, device and medium based on an event sensor. Background Art
[0002] With the continuous development of human-computer interaction technology, eye tracking technology, as an efficient method for detecting and analyzing biological signals, has been widely used in many fields, including psychological research, market research, virtual reality (VR), augmented reality (AR), intelligent control, etc. Eye tracking technology can accurately reflect an individual's visual attention, cognitive load, emotional response and other psychological states by real-time monitoring and analyzing parameters such as the movement trajectory, gaze point and gaze time of the human eye.
[0003] Most traditional eye tracking systems are based on video image processing and infrared sensor technology, and use high-speed cameras or infrared sensors to capture the movement trajectory of the eyeball. However, existing eye tracking methods based on video cameras have some significant shortcomings, especially in high-dynamic scenes, which are easily affected by factors such as ambient light changes, object occlusion, motion blur, etc., resulting in reduced data acquisition accuracy and reliability. In recent years, event cameras have emerged as a new type of sensor in the field of computer vision, especially in tasks such as dynamic tracking and fast motion detection. The difference between event cameras and traditional frame capture cameras is that they are based on an "event-driven" approach, and only record data when there is a significant change in the pixels in the scene (i.e., changes in light intensity), rather than periodically capturing image frames at fixed time intervals. This enables event cameras to have higher timing accuracy and lower latency in dynamic scenes, and can capture high-speed eye movements and small changes that traditional cameras cannot obtain.
[0004] The shortcomings of traditional eye tracking methods are obvious: high frame rate video processing requirements are high, real-time performance is poor, and it is easily affected by reflections and motion blur in dynamic scenes. Another hybrid eye tracking method that combines infrared cameras and event sensors retains the advantages of event sensors, but this method requires a high-precision optical system and a complex calibration process. In addition, the hybrid eye tracking method relies more on infrared sensors to obtain the position of the eyeball and pupil ellipse information. Although event cameras can provide higher temporal resolution, the temporal accuracy and latency of infrared cameras are usually lower than those of event cameras. Therefore, the overall real-time performance of the hybrid system is not as good as that of a pure event camera system.
[0005] Therefore, how to combine the advantages of event cameras to improve the accuracy, real-time performance and environmental adaptability of eye tracking technology has become an important research direction in this field. The present invention proposes an eye tracking method / device based on an event camera, which utilizes the excellent performance of event cameras in dynamic scenes, overcomes the limitations of traditional eye tracking methods, and provides an eye tracking solution with high accuracy, low latency and strong environmental adaptability. Summary of the invention
[0006] The purpose of the present invention is to provide a near-eye gaze point tracking method, device and medium based on event sensor to solve the problems raised in the above background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions: A near-eye gaze point tracking method based on an event sensor, comprising: S1, the infrared light source provides infrared illumination for the eyes and the surrounding area of the eyes, and the event sensor collects information about the eyes and the surrounding area of the eyes, and obtains a series of event stream information generated by the illumination changes caused by movement; S2, sending the time window clustering event set T to the ROI and pupil calculation module, and outputting pupil ellipse information P, where the time window length is t; S3, sending the event stream information to the vector calculation module, combining it with the pupil ellipse information P, and calculating the motion vector of the pupil center; S4, combining the motion vector obtained in S3, updating the pupil ellipse information P; S5, combining a dynamic filtering method to smooth the pupil ellipse information P; S6, the filtered pupil ellipse information P is transmitted to the sight line regression and compensation model, and the gaze point coordinates are calculated by combining the deep learning method.
[0008] Furthermore, the ROI and pupil calculation module calculates pupil ellipse information by combining the least square method and the random sampling consensus method, and the process includes: S2.1, calculate the event density based on the event flow information obtained in S1, event density = number of events in the time window / time window t; determine the center position of the ROI based on the event density; S2.2, calculating the geometric center of gravity of the events in the ROI area, and preliminarily determining the geometric center of gravity as the center point of the pupil; S2.3, using the center point of the pupil and the preset pupil ellipse information P, filter the events in the neighborhood of the pupil ellipse information P to obtain the filtered event set T′: in, are the spatial coordinates of the event, is the preset neighborhood screening threshold; S2.4, the least squares method is applied to fit the filtered event set T′ to obtain a quadratic equation of an ellipse representing the pupil shape. At the same time, a random sampling consistency algorithm is applied to automatically identify and remove outliers or abnormal values by evaluating the quality of the fitting results.
[0009] Further, the S3 includes: S3.1, Time domain division: Select the time window length as Splitting the event stream information into several event units Ei; S3.2, spatial domain partitioning: For an event in the segmented event unit Ei, select the events in the surrounding L×L neighborhood to form the event set Q; S3.3, applying the principal component analysis method to the event set Q, and calculating a number of eigenvalues and a number of eigenvectors; S3.4, using the eigenvalues and eigenvectors obtained in S3.3, fit the small plane ,in, It is expressed as: , a, b, c are unknown parameters. Among the eigenvalues calculated by S3.3, the eigenvector corresponding to the minimum eigenvalue is Corresponding to the required parameters [a, b, c], that is , Respectively represent the components of the eigenvector V in the x, y, and t directions; S3.5, calculate the motion vector of the pupil center: In the formula, Respectively represent the motion vector of the center of the pupil ellipse information P in the x and y directions: S3.6, for the obtained small plane All points within are evaluated: In the formula, is the evaluation value of the point, which is less than the neighborhood screening threshold Points with a value greater than the neighborhood filtering threshold are considered to be points inside the plane. The points outside the boundary are regarded as out-of-bounds points. When the number of out-of-bounds points is too large, the plane fitting result is discarded. S3.7, traverse the events in each event unit Ei, perform the operations of S3.2-S3.6 on each event, and calculate a series of vector information , Represents the vector information of the i-th point, represents the i-th point Directional component, represents the i-th point Directional component; S3.8, a series of vector information obtained in S3.7 , perform weighted summation, and finally obtain the pupil center motion vector , vector This is the output of the vector calculation module.
[0010] Furthermore, the specific method of updating the pupil ellipse information P in S4 in combination with the motion vector obtained in S3 is: in, is the pupil center information of the previous moment, is the updated pupil center information, is the motion vector of the pupil center obtained by the vector calculation module in S3.
[0011] Further, the S5 includes: The updated pupil coordinates are smoothed, and the filtering process is expressed as: in, is the filtered signal at the current moment, that is, the smoothed pupil center coordinates, is the signal after filtering at the previous moment, is the current input signal, that is, the current pupil center coordinate, is the filter coefficient, which is used to control the smoothness of the signal. Use the following formula: in is the dynamically adjusted cutoff frequency, is the signal sampling frequency; Dynamic cutoff frequency calculation method: in, is the basic cutoff frequency, min means taking the minimum value, is the adjustment factor of the response speed, is the rate of change of the input signal.
[0012] Further, the S6 includes: S6.1, the user gazes at the known calibration point coordinates on the screen in turn; the calibration point coordinates and the corresponding pupil center coordinates are recorded, and the matching data is transmitted to the sight line regression and compensation model, and the sight line regression and compensation model is trained and optimized to learn the complex mapping relationship between the pupil center and the screen gaze point coordinates, and obtain the optimal sight line regression and compensation model; S6.2, transmitting the pupil center coordinates to the optimal sight line regression and compensation model obtained in S6.1 to obtain the gaze point coordinates; S6.3, the gaze regression and compensation model is used to fit the gaze surface, and the gaze point coordinates obtained in S6.2 are compensated and regressed to obtain the final gaze point coordinates.
[0013] The present invention also provides a near-eye gaze point tracking device based on an event sensor, comprising one or more processors for implementing a near-eye gaze point tracking method based on an event sensor as described above.
[0014] The present invention also provides a readable storage medium having a program stored thereon, and when the program is executed by a processor, the method for near-eye gaze point tracking based on an event sensor as described above is implemented.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1) High-precision tracking: Compared with traditional eye tracking systems, this invention introduces event sensors, which can more accurately identify moving targets in dynamic scenes, especially in scenes with high-speed eye movements and rapid changes, and can effectively avoid the influence of motion blur and ambient light changes.
[0016] 2) Reduce the complexity of information processing: Compared with the hybrid system based on infrared sensors and event sensors, the present invention adopts a single-mode information processing method, which not only simplifies the complexity of information synchronization, but also reduces the burden of system design and improves processing speed.
[0017] 3) Improve real-time performance and low latency: The data stream driven mode based on the event camera enables the system to collect and process data with higher time accuracy and lower latency, overcoming the limitations of traditional cameras in high-dynamic scenes.
[0018] 4) Strong adaptability and robustness: By combining the dynamic filtering method and the vector calculation module, it can effectively eliminate noise interference and maintain the ability to respond to rapid eye movements. This makes the method more environmentally adaptable and suitable for different application scenarios, such as VR / AR, driving monitoring, etc.
[0019] 5) Higher computational efficiency: Compared with traditional multimodal eye tracking systems, the present invention greatly reduces the computational burden and can achieve smooth prediction of pupil displacement and accurate calculation of gaze point with less event data, thus optimizing real-time performance and system performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flow chart of a near-eye gaze point tracking method based on an event sensor of the present invention.
[0021] Figure 2 This is a schematic diagram of event data accumulated by an event camera within 33 ms in a near-eye gaze point tracking method based on an event sensor of the present invention, wherein green dots are events with positive polarity, and red dots are points with negative polarity.
[0022] Figure 3 The present invention is a flowchart of the vector calculation module in a near-eye gaze point tracking method based on an event sensor.
[0023] Figure 4 It is a structural schematic diagram of a near-eye gaze point tracking device based on an event sensor of the present invention. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] See also Figure 1-Figure 3 , a near-eye gaze tracking method based on event sensor, comprising the following steps: S1, the infrared light source provides infrared illumination for the eyes and the surrounding areas of the eyes, and the event sensor collects information about the eyes and the surrounding areas of the eyes, and obtains a series of event information e generated by the illumination changes caused by movement.
[0026] Among them, the event information e={x,y,p,t} includes the event coordinate information x, y, polarity information p and time window length t (timestamp information).
[0027] S2, clustering event set T based on time window, where the time window length is t (generally t = 33ms). In a near-eye gaze point tracking method based on event sensor of the present invention, the event camera clusters event set T based on time window. Figure 2As shown, in the event data accumulated by the event camera within 33 ms, the green points represent events with positive polarity, and the red points represent points with negative polarity. These events constitute the event set T.
[0028] The time window clustering event set T is sent to the ROI and pupil calculation module. ROI (region of interest) is the region of interest. The ROI and pupil calculation module outputs pupil ellipse information, where the pupil ellipse information includes the center position of the ellipse, the length of the major axis, the minor axis, and the rotation angle.
[0029] The input of the ROI and pupil calculation module is the event set T consisting of events in the time window. Each event captured by the event camera has a precise timestamp and spatial coordinates, and the spatiotemporal characteristics of the event are used to determine the region of interest (ROI). Specifically, the ROI and pupil calculation module combines the least squares method and the random sampling consensus method to calculate the pupil ellipse information. The process includes: S2.1, calculate the event density according to the event flow information obtained in S1, event density = number of events in the time window / time window length t. Determine the center position of the ROI according to the event density. The ROI area is a circle with a preset radius, which is close to the size of the pupil.
[0030] S2.2, after the ROI area is determined, the events in the area are selected, and the geometric center of gravity of these events is calculated, which is preliminarily determined as the center point of the pupil.
[0031] S2.3, using the center point of the pupil and the preset pupil ellipse information P, filter the events in the neighborhood of the pupil ellipse information P to obtain the filtered event set T′:
[0032] in, are the spatial coordinates of the event, is the preset neighborhood screening threshold.
[0033] S2.4, apply the least squares method to fit the filtered event set T′, with the goal of obtaining a quadratic equation of an ellipse representing the pupil shape. However, due to the interference of noise and outliers, the simple use of the least squares method may lead to inaccurate fitting results, especially when the pupil boundary is irregular or there is occlusion. In order to solve this problem, the present invention introduces a random sampling consistency algorithm, which automatically identifies and removes outliers or outliers by evaluating the quality of the fitting results. The robustness of the random sampling consistency algorithm enables the present invention to obtain more accurate ellipse fitting results in the presence of noise and imperfect data, and improve the anti-interference ability of the model.
[0034] It should be noted that the least squares method is a mathematical tool that is widely used in many disciplines of data processing such as error estimation, uncertainty, system identification, prediction, and forecasting, and is a well-known technology.
[0035] It should be noted that the ROI and pupil calculation modules are computer programs, and their specific implementation can be found in S2.1-S2.4 above, which will not be described in detail here.
[0036] S3, the input of the vector calculation module is the event stream information and the pupil ellipse information P, and the output is the motion vector of the pupil center, that is, the speed and direction of the pupil center in the camera space. Figure 3 As shown, the workflow of the vector calculation module includes the steps of receiving, analyzing and calculating event data and pupil ellipse data. The event stream information is sent to the vector calculation module, and the motion vector is calculated in combination with the pupil ellipse information P, which specifically includes: S3.1, time domain division: select a time window length of △T (optional 3ms, 5ms, etc.) to divide the event stream information, and divide the event stream information into several event units Ei.
[0037] S3.2, spatial domain partitioning: For an event in the segmented event unit Ei, select the events in the surrounding L×L (L can be 3 or 5) neighborhood to form an event set Q.
[0038] S3.3, apply the principal component analysis method to the event set Q, and calculate 3 eigenvalues and 3 eigenvectors; the eigenvector V corresponding to the smallest eigenvalue has special meaning and can be used to solve the vector information.
[0039] It should be noted that the principal component analysis method is a well-known technology. It combines the original variables into a few new variables through linear combination with the premise of minimizing information loss. New variables are used to replace the original variables in data modeling, which can greatly reduce the computational workload in the analysis process. The selection of new variables by the principal component is not a simple choice of the original variables, but the result of the reorganization of the original variables. Therefore, it will not cause a large amount of loss of the original variable information and can represent most of the information of the original variables. At the same time, the selected new variables are unrelated to each other, which can effectively solve many problems brought to the analysis application by overlapping variable information and multicollinearity.
[0040] S3.4, using the eigenvalues and eigenvectors obtained in S3.3, fit the small plane ,in, It is expressed as: , a, b, c, d are unknown parameters; the eigenvector corresponding to the minimum eigenvalue calculated by S3.3 Corresponding to the required parameters [a, b, c], that is , They represent the components of the eigenvector V in the x, y, and t directions respectively.
[0041] S3.5, calculate the motion vector of the pupil center:
[0042] In the formula, Represents the motion vector of the center of the pupil ellipse information P in the x and y directions.
[0043] S3.6, for the small plane obtained by fitting All points within are evaluated:
[0044] In the formula, is the evaluation value of the point, which is less than the neighborhood screening threshold Points with a value greater than the neighborhood filtering threshold are considered to be points inside the plane. The points outside the bounds are considered as out-of-bounds points. When the number of out-of-bounds points is too large, the plane fitting result is discarded.
[0045] S3.7, traverse the events in each event unit Ei, perform the operations of S3.2-S3.6 on each event, and calculate a series of vector information , Represents the vector information of the i-th point, represents the i-th point Directional component, represents the i-th point Directional component.
[0046] S3.8, a series of vector information obtained in S3.7 , perform weighted summation, and finally get the motion vector ; Vector This is the output of the vector calculation module.
[0047] It should be noted that the vector calculation module is a computer program, and its specific implementation can be found in S3.1-S3.8 above, which will not be repeated here.
[0048] S4, combining the motion vector obtained in S3, and updating the pupil ellipse information P.
[0049] The vector calculation module outputs the motion vector of the pupil center, which is used to update the pupil ellipse information. The specific method is as follows:
[0050] in, is the pupil center information of the previous moment, is the updated pupil center information, is the motion vector of the pupil center obtained by the vector calculation module in S3.
[0051] Predicting pupil displacement through the vector calculation module can significantly reduce the impact of noise interference and ambient light changes, and improve positioning accuracy and smoothness. Compared with directly using event data fitting, this method effectively reduces the interference of artifacts such as reflected light spots on the results through motion trend prediction. At the same time, combined with optimization strategies such as adaptive time step, acceleration vector correction and boundary constraints, its robustness and stability in dynamic scenes can be further enhanced, which is suitable for pupil positioning tasks in complex environments. When the reflective point of the illumination light source is close to the edge of the pupil, direct edge fitting may lead to unstable prediction, while vector update relies on the overall motion trend, which can effectively avoid such problems. In addition, the traditional method has a high jump in the pupil center coordinates and relies on a large amount of event information to obtain relatively accurate results, which not only reduces the tracking frequency of the eye tracking system, but also affects the real-time performance. Through the vector calculation module, only less event information is needed to generate a smooth motion vector, thereby achieving a more stable and smooth pupil center coordinate prediction.
[0052] S5, combining a dynamic filtering method, smoothing the pupil ellipse information P, specifically including: The vector calculation module outputs the predicted pupil displacement, and the updated pupil coordinates can be obtained by combining the pupil displacement. Then, the updated pupil coordinates are smoothed by combining the dynamic filtering method.
[0053] The filtering process can be expressed as: in, is the filtered signal at the current moment, that is, the smoothed pupil center coordinates, is the signal after filtering at the previous moment, is the current input signal, that is, the current pupil center coordinate, is the filter coefficient, which is used to control the smoothness of the signal. Use the following formula: in is the dynamically adjusted cutoff frequency, is the signal sampling frequency.
[0054] Dynamic cutoff frequency calculation method:
[0055] in, is the basic cutoff frequency, which determines the basic smoothness of the signal and is generally set to 1000; min means taking the minimum value, β is the adjustment factor of the response speed, its value is in the range of [0,1], controls the sensitivity of the filter to the rate of change of the signal, and can be set based on experience; is the rate of change of the input signal.
[0056] This dynamic filtering method is particularly suitable for real-time systems. It can effectively smooth noise signals while maintaining the ability to respond to rapid changes. Compared with the Kalman filter, it is simpler and less computationally intensive, and is suitable for embedded systems or scenarios that require real-time processing. The dynamic filtering method uses adaptive smoothing, which can dynamically smooth the degree of signal change according to the rate of change. For rapidly changing signals, the filter reduces smoothing and allows for rapid response.
[0057] S6, sending the filtered pupil ellipse information P to the sight line regression and compensation model, combining the deep learning method to calculate the gaze point coordinates. Specifically including: The sight line regression and compensation model used in the present invention is a nonlinear regression model based on a fully connected neural network. Specifically, the model takes the user's pupil center coordinates (two-dimensional coordinate values) as input and outputs the predicted gaze point coordinates on the corresponding screen.
[0058] The model consists of the following parts: (1) Input layer: accepts two-dimensional pupil center coordinates; (2) Hidden layer: The model contains three hidden layers, each of which consists of 128, 64, and 32 neurons, respectively. The ReLU activation function is used to enhance the nonlinear fitting ability of the network. (3) Output layer: Outputs two-dimensional predicted gaze point coordinates.
[0059] In order to improve the generalization performance of the model and prevent overfitting, batch normalization and Dropout layers are added after each hidden layer to enhance the model's generalization ability of the nonlinear relationship between pupil center coordinates and gaze point.
[0060] 9 or 16 evenly distributed calibration points are defined on the screen in advance as calibration references. These points cover different areas of the screen to enhance the prediction ability of the model in the full field of view. When the user gazes at these known points in turn, the camera captures the coordinates of the pupil center in real time and matches these coordinates with the calibration points on the screen to generate data for model training. The matching data is input into the sight regression and compensation model for training optimization. When the user's gaze point exceeds the area where the calibration point is located, the gaze point prediction error may increase. Therefore, the regression and compensation model believes that the position of the pupil center and the screen coordinates are not strictly linear, but conform to a nonlinear mapping with a gaze surface close to the screen. The regression and compensation model captures the complex mapping characteristics between the pupil center and the gaze point by fitting this gaze surface. Especially in the edge area or when the calibration points are sparsely distributed, the farther the gaze point is from the center of the screen, the more compensation is obtained. This dynamic compensation mechanism significantly improves the accuracy of gaze point calculation.
[0061] Through this method, the regression and compensation model improves the spatial accuracy of gaze point prediction, making it applicable to a wider range of practical scenarios, such as screen interaction, gaze tracking analysis, and gaze point control in virtual reality. The combination of the nonlinear mapping advantages of neural networks and the dynamic compensation mechanism lays an important foundation for future higher-precision gaze tracking technology.
[0062] See also Figure 4 An embodiment of the present invention provides a near-eye gaze point tracking device based on an event sensor, including one or more processors, which are used to implement a mobile robot trajectory tracking control method based on high-order full-drive theory in the above embodiment.
[0063] An embodiment of a near-eye gaze tracking device based on an event sensor of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From a hardware perspective, if Figure 4 As shown in FIG. 1 , a hardware structure diagram of a near-eye gaze point tracking device based on an event sensor of the present invention is a device with data processing capability, except that Figure 4 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in the embodiments may also include other hardware, usually based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0064] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0065] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0066] An embodiment of the present invention further provides a readable storage medium having a program stored thereon. When the program is executed by a processor, a mobile robot trajectory tracking control method based on high-order full-drive theory in the above embodiment is implemented.
[0067] The readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The readable storage medium may also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the readable storage medium may also include both an internal storage unit of any device with data processing capability and an external storage device. The readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.
[0068] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A near-eye gaze point tracking method based on event sensor, characterized in that: include: S1, the infrared light source provides infrared illumination for the eyes and the surrounding area of the eyes, and the event sensor collects information about the eyes and the surrounding area of the eyes, and obtains a series of event stream information generated by the illumination changes caused by movement; S2, sending the time window clustering event set T to the ROI and pupil calculation module, and outputting pupil ellipse information P, where the time window length is t; S3, sending the event stream information to the vector calculation module, combining it with the pupil ellipse information P, and calculating the motion vector of the pupil center; S4, combining the motion vector obtained in S3, updating the pupil ellipse information P; S5, combining a dynamic filtering method to smooth the pupil ellipse information P; S6, the filtered pupil ellipse information P is transmitted to the sight line regression and compensation model, and the gaze point coordinates are calculated by combining the deep learning method.
2. The method for near-eye gaze point tracking based on event sensor according to claim 1, characterized in that: The ROI and pupil calculation module calculates pupil ellipse information by combining the least square method and the random sampling consensus method. The process includes: S2.1, calculate the event density based on the event flow information obtained in S1, event density = number of events in the time window / time window t; determine the center position of the ROI based on the event density; S2.2, calculating the geometric center of gravity of the events in the ROI area, and preliminarily determining the geometric center of gravity as the center point of the pupil; S2.3, using the center point of the pupil and the preset pupil ellipse information P, filter the events in the neighborhood of the pupil ellipse information P to obtain the filtered event set T′: in, are the spatial coordinates of the event, is the preset neighborhood screening threshold; S2.4, the least squares method is applied to fit the filtered event set T′ to obtain a quadratic equation of an ellipse representing the pupil shape. At the same time, a random sampling consistency algorithm is applied to automatically identify and remove outliers or abnormal values by evaluating the quality of the fitting results.
3. The method for near-eye gaze point tracking based on event sensor according to claim 1, characterized in that: The S3 includes: S3.1, Time domain division: Select the time window length as Splitting the event stream information into several event units Ei; S3.2, spatial domain partitioning: For an event in the segmented event unit Ei, select the events in the surrounding L×L neighborhood to form the event set Q; S3.3, applying the principal component analysis method to the event set Q, and calculating a number of eigenvalues and a number of eigenvectors; S3.4, using the eigenvalues and eigenvectors obtained in S3.3, fit the small plane ,in, It is expressed as: , a, b, c are unknown parameters. Among the eigenvalues calculated by S3.3, the eigenvector corresponding to the minimum eigenvalue is Corresponding to the required parameters [a, b, c], that is , Respectively represent the components of the eigenvector V in the x, y, and t directions; S3.5, calculate the motion vector of the pupil center: In the formula, Respectively represent the motion vector of the center of the pupil ellipse information P in the x and y directions: S3.6, for the obtained small plane All points within are evaluated: In the formula, is the evaluation value of the point, which is less than the neighborhood screening threshold Points with a value greater than the neighborhood filtering threshold are considered to be points inside the plane. The points outside the boundary are regarded as out-of-bounds points. When the number of out-of-bounds points is too large, the plane fitting result is discarded. S3.7, traverse the events in each event unit Ei, perform the operations of S3.2-S3.6 on each event, and calculate a series of vector information , Represents the vector information of the i-th point, represents the i-th point Directional component, represents the i-th point Directional component; S3.8, a series of vector information obtained in S3.7 , perform weighted summation, and finally obtain the pupil center motion vector , vector This is the output of the vector calculation module.
4. The method for near-eye gaze point tracking based on event sensor according to claim 1, characterized in that: The specific method of updating the pupil ellipse information P in S4 in combination with the motion vector obtained in S3 is: in, is the pupil center information of the previous moment, is the updated pupil center information, is the motion vector of the pupil center obtained by the vector calculation module in S3.
5. The method for near-eye gaze point tracking based on event sensor according to claim 1, characterized in that: The S5 includes: The updated pupil coordinates are smoothed, and the filtering process is expressed as: in, is the filtered signal at the current moment, that is, the smoothed pupil center coordinates, is the signal after filtering at the previous moment, is the current input signal, that is, the current pupil center coordinate, is the filter coefficient, which is used to control the smoothness of the signal. Use the following formula: in is the dynamically adjusted cutoff frequency, is the signal sampling frequency; Dynamic cutoff frequency calculation method: in, is the basic cutoff frequency, min means taking the minimum value, is the adjustment factor of the response speed, is the rate of change of the input signal.
6. The method for near-eye gaze point tracking based on event sensor according to claim 1, characterized in that: The S6 includes: S6.1, the user gazes at the known calibration point coordinates on the screen in turn; the calibration point coordinates and the corresponding pupil center coordinates are recorded, and the matching data is transmitted to the sight line regression and compensation model, and the sight line regression and compensation model is trained and optimized to learn the complex mapping relationship between the pupil center and the screen gaze point coordinates, and obtain the optimal sight line regression and compensation model; S6.2, transmitting the pupil center coordinates to the optimal sight line regression and compensation model obtained in S6.1 to obtain the gaze point coordinates; S6.3, the gaze regression and compensation model is used to fit the gaze surface, and the gaze point coordinates obtained in S6.2 are compensated and regressed to obtain the final gaze point coordinates.
7. A near-eye gaze point tracking device based on an event sensor, characterized in that: It includes one or more processors for implementing a near-eye gaze point tracking method based on an event sensor as described in any one of claims 1-6.
8. A readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, a near-eye gaze point tracking method based on an event sensor as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Lightweight pulse wave signal noise removal and heart rate detection method
CN114098690A
First visual angle hand tracking system based on event camera and RGB camera and application
CN118505742A
Eye movement tracking method and system based on pupil shift
CN118506433A
Multi-mode user intention recognition method and system based on eye movement tracking
CN119148861A
Non-vision field imaging detection method and system of fusion event camera
CN119579644A