An Event Sensor-Based Near-Eye Gaze Point Tracking Method, Device, and Medium
The eye-tracking method using an event camera and deep learning enhances precision and adaptability in dynamic scenes by processing event sensor data for improved gaze point calculation, addressing the limitations of traditional eye-tracking systems.
Patent Information
- Application Number
- CN202510444657.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The existing eye tracking technology is susceptible to factors such as ambient light changes, object occlusion and motion blur in high dynamic scenarios, resulting in reduced data acquisition accuracy and reliability. The traditional method has poor real-time performance, and the hybrid eye tracking method relies on infrared sensors and has poor overall real-time performance.
The eye movement tracking method based on the event camera is adopted to provide eye illumination through infrared light sources, and the eye information is collected using event sensors. Combined with time window clustering, vector calculation, dynamic filtering and deep learning, the motion vector and gaze coordinates of the pupil center are calculated to simplify information processing and enhance environmental adaptability.
It realizes high-precision and low-latency eye tracking in dynamic scenarios, reduces the complexity of information processing, improves real-time and robustness, and is suitable for application scenarios such as VR/AR.
Smart Images

Figure CN119960605B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field, and specifically to a near-eye gaze point tracking method, device and medium based on an event sensor. Background Art
[0002] With the continuous development of human-computer interaction technology, eye tracking technology, as an efficient method for detecting and analyzing biological signals, has been widely applied in multiple fields, including psychological research, market research, virtual reality (VR), augmented reality (AR), intelligent control, etc. Eye tracking technology can accurately reflect psychological states such as an individual's visual attention, cognitive load, and emotional response by real-time monitoring and analyzing parameters such as the movement trajectory, fixation point, and gaze time of the human eye.
[0003] Most traditional eye tracking systems are based on video image processing and infrared sensor technology, and capture the movement trajectory of the eyeball through a high-speed camera or an infrared sensor. However, existing eye tracking methods based on video cameras have some significant deficiencies. Especially in high-dynamic scenes, they are easily affected by factors such as ambient light changes, object occlusion, and motion blur, resulting in a reduction in the accuracy and reliability of data acquisition. In recent years, the event camera, as a new type of sensor, has emerged in the field of computer vision and has shown excellent performance in tasks such as dynamic tracking and fast motion detection. The difference between an event camera and a traditional frame-capturing camera is that it is based on an "event-driven" method and only records data when significant changes occur in the pixels of the scene (i.e., changes in light intensity), rather than periodically capturing image frames at fixed time intervals. This enables the event camera to have higher temporal accuracy and lower latency in dynamic scenes and can capture high-speed eye movements and minute changes that cannot be obtained by traditional cameras.
[0004] The deficiencies of traditional eye tracking methods are very obvious: for example, the processing requirements for high-frame-rate videos are high, the real-time performance is poor, and they are easily affected by factors such as reflections and motion blur in dynamic scenes. Another hybrid eye tracking method that combines an infrared camera and an event sensor retains the advantages of the event sensor, but this method requires a high-precision optical system and a complex calibration process, and the hybrid eye tracking method relies relatively heavily on the infrared sensor to obtain the position of the eyeball and pupil ellipse information. Although the event camera can provide a high time resolution, the time accuracy and latency of the infrared camera are usually lower than those of the event camera. Therefore, the overall real-time performance of the hybrid system is not as good as that of a pure event camera system.
[0005] Therefore, how to combine the advantages of event cameras to improve the accuracy, real-time performance, and environmental adaptability of eye tracking technology has become an important research direction in this field. The present invention proposes an eye tracking method / device based on an event camera, which utilizes the excellent performance of the event camera in dynamic scenes, overcomes the limitations of traditional eye tracking methods, and provides a high-precision, low-latency, and strongly environmentally adaptable eye tracking solution. Summary of the Invention
[0006] The purpose of the present invention is to provide a near-eye fixation point tracking method, device, and medium based on an event sensor to solve the problems proposed in the above background technology.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] A near-eye fixation point tracking method based on an event sensor includes:
[0009] S1. An infrared light source provides infrared illumination for the eye and the surrounding area of the eye, and an event sensor collects information on the eye and the surrounding area of the eye to obtain a series of event stream information generated by the change in light caused by movement.
[0010] S2. Send the time window clustering event set T to the ROI and pupil calculation module, and output pupil ellipse information P, where the time window length is t.
[0011] S3. Send the event stream information to the vector calculation module, and combine the pupil ellipse information P to calculate the motion vector of the pupil center.
[0012] S4. Combine the motion vector obtained in S3 to update the pupil ellipse information P.
[0013] S5. Combine the dynamic filtering method to smooth the pupil ellipse information P.
[0014] S6. Send the filtered pupil ellipse information P to the line-of-sight regression and compensation model, and combine the deep learning method to calculate the fixation point coordinates.
[0015] Furthermore, the ROI and pupil calculation module calculates the pupil ellipse information by combining the least squares method and the random sample consensus method. The process includes:
[0016] S2.1. According to the event stream information obtained in S1, calculate the event density. Event density = the number of events within the time window / the time window t; determine the central position of the ROI according to the event density.
[0017] S2.2. Calculate the geometric centroid of the events within the ROI area, and initially determine the geometric centroid as the center point of the pupil.
[0018] S2.3, using the center point of the pupil and combining with the preset pupil ellipse information P, screen the events within the neighborhood of the pupil ellipse information P to obtain the screened event set T':
[0019]
[0020] wherein, is the spatial coordinate of the event, is the preset neighborhood screening threshold;
[0021] S2.4, for the screened event set T', apply the least squares method for fitting to obtain a quadratic equation of an ellipse representing the pupil shape. At the same time, apply the random sample consensus algorithm to automatically identify and remove outliers or anomalies by evaluating the quality of the fitting result.
[0022] Furthermore, the said S3 includes:
[0023] S3.1, time domain division: Select the time window length as Split the event stream information into several event units Ei;
[0024] S3.2, spatial domain division: For a certain event in the segmented event unit Ei, select the events in its surrounding L×L neighborhood to form an event set Q;
[0025] S3.3, apply the principal component analysis method to the event set Q to calculate several eigenvalues and several eigenvectors;
[0026] S3.4, use the eigenvalues and eigenvectors obtained in S3.3 to fit a microplane , wherein, is expressed as: , a, b, c are unknown parameters. Among the eigenvalues calculated in S3.3, the eigenvector corresponding to the smallest eigenvalue corresponds to the desired parameters [a, b, c], that is , respectively represent the components of the eigenvector V in the x, y, and t directions;
[0027] S3.5, calculate the motion vector of the pupil center:
[0028]
[0029] In the formula, respectively represent the motion vectors of the center of the pupil ellipse information P in the x and y directions:
[0030] S3.6, evaluate all the points within the obtained microplane :
[0031]
[0032] Wherein, is the evaluation value of the point. Points less than the neighborhood screening threshold are regarded as points inside the plane, and points greater than the neighborhood screening threshold are regarded as points outside the boundary. When the number of points outside the boundary is too large, the plane fitting result is discarded;
[0033] S3.7. Traverse the events in each event unit Ei, perform the operations of S3.2 - S3.6 on each event, and calculate a series of vector information , represents the vector information of the i-th point, represents the direction component of the i-th point, represents the direction component of the i-th point;
[0034] S3.8. Perform weighted summation on the series of vector information obtained in S3.7, and finally obtain the pupil center motion vector . The vector is the output of the vector calculation module.
[0035] Furthermore, the specific method for updating the pupil ellipse information P by combining the motion vector obtained in S3 in S4 is as follows:
[0036]
[0037] Wherein, is the pupil center information at the previous moment, is the updated pupil center information, is the motion vector of the pupil center obtained by the vector calculation module in S3.
[0038] Furthermore, the S5 includes:
[0039] Perform smoothing processing on the updated pupil coordinates, and the filtering process is expressed as:
[0040]
[0041] Wherein, is the filtered signal at the current moment, that is, the smoothed pupil center coordinate, is the filtered signal at the previous moment, is the currently input signal, that is, the current pupil center coordinate, is the filtering coefficient used to control the signal smoothing degree, Adopt the following formula:
[0042]
[0043] wherein is the dynamically adjusted cut-off frequency, is the signal sampling frequency;
[0044] Method for calculating the dynamic cut-off frequency:
[0045]
[0046] wherein, is the basic cut-off frequency, and min represents taking the minimum value, is the adjustment factor of the response speed, is the change rate of the input signal.
[0047] Furthermore, the S6 includes:
[0048] S6.1, the user sequentially gazes at the known calibration point coordinates on the screen; records the calibration point coordinates and the corresponding pupil center coordinates, transmits the matching data to the line-of-sight regression and compensation model, trains and optimizes the line-of-sight regression and compensation model, learns the complex mapping relationship between the pupil center and the screen fixation point coordinates, and obtains the optimal line-of-sight regression and compensation model;
[0049] S6.2, transmits the pupil center coordinates to the optimal line-of-sight regression and compensation model obtained in S6.1 to obtain the fixation point coordinates;
[0050] S6.3, the line-of-sight regression and compensation model fits the gaze surface, performs compensation regression on the fixation point coordinates obtained in S6.2, and obtains the final fixation point coordinates.
[0051] The present invention also provides a near-eye fixation point tracking device based on an event sensor, including one or more processors for implementing a near-eye fixation point tracking method as described above.
[0052] The present invention also provides a readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements a near-eye fixation point tracking method as described above.
[0053] Compared with the prior art, the beneficial effects of the present invention are:
[0054] 1) High-precision tracking: Compared with traditional eye movement tracking systems, the present invention introduces an event sensor, which can more accurately identify moving targets in dynamic scenarios, especially in scenarios with rapid changes in high-speed eye movements, and can effectively avoid the influence of motion blur and ambient light changes.
[0055] 2) Reduce the complexity of information processing: Compared with the hybrid system based on infrared sensors and event sensors, the present invention adopts a single-modal information processing method. This not only simplifies the complexity of information synchronization, but also reduces the burden of system design and improves the processing speed.
[0056] 3) Improve real-time performance and low latency: Based on the data stream-driven mode of the event camera, the system can collect and process data with higher time accuracy and lower latency, overcoming the limitations of traditional cameras in high-dynamic scenes.
[0057] 4) Strong adaptability and robustness: By combining the dynamic filtering method and the vector calculation module, it can effectively eliminate noise interference and maintain the ability to respond to rapid eye movements. This makes the method have stronger environmental adaptability and is applicable to different application scenarios, such as VR / AR, driving monitoring, etc.
[0058] 5) Higher computational efficiency: Compared with the traditional multi-modal eye movement tracking system, the present invention greatly reduces the computational burden and can achieve smooth prediction of pupil displacement and accurate calculation of the fixation point with less event data, optimizing real-time performance and system performance. Brief Description of the Drawings
[0059] Figure 1 It is a flowchart of a near-eye fixation point tracking method based on event sensors of the present invention.
[0060] Figure 2 It is a schematic diagram of the event data accumulated by the event camera within 33 ms in the near-eye fixation point tracking method based on event sensors of the present invention, where the green dots are events with positive polarity and the red dots are events with negative polarity.
[0061] Figure 3 It is a flowchart of the operation of the vector calculation module in the near-eye fixation point tracking method based on event sensors of the present invention.
[0062] Figure 4 It is a schematic structural diagram of a near-eye fixation point tracking device based on event sensors of the present invention. Detailed Embodiment
[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] Please refer to Figures 1-3, A near-eye gaze point tracking method based on an event sensor, comprising the following steps:
[0065] S1, An infrared light source provides infrared illumination for the eye and the surrounding area of the eye, and an event sensor collects information on the eye and the surrounding area of the eye to obtain a series of event information e generated by the change in light caused by movement.
[0066] Among them, the event information e = {x, y, p, t}, including the coordinate information x, y of the event, the polarity information p, and the time window length t (timestamp information).
[0067] S2, Cluster the event set T based on the time window, where the time window length is t (generally set t = 33ms). In a near-eye gaze point tracking method based on an event sensor of the present invention, the event camera clusters the event set T based on the time window. As Figure 2 shown, among the event data accumulated by the event camera within 33ms, the green dots represent events with a positive polarity, and the red dots represent events with a negative polarity. These events constitute the event set T.
[0068] The time window-clustered event set T is sent to the ROI and pupil calculation module. ROI (region of interest) is the region of interest. The ROI and pupil calculation module outputs pupil ellipse information, where the pupil ellipse information includes the center position of the ellipse, the lengths of the major axis and minor axis, and the rotation angle.
[0069] The input of the ROI and pupil calculation module is the event set T composed of events within the time window. Each event captured by the event camera carries an accurate timestamp and spatial coordinates, and the spatio-temporal characteristics of the events are used to determine the region of interest (ROI). Specifically, the ROI and pupil calculation module, combining the least squares method and the random sample consensus method, calculates the pupil ellipse information. The process includes:
[0070] S2.1, According to the event stream information obtained in S1, calculate the event density. The event density = the number of events within the time window / the time window length t. Determine the center position of the ROI according to the event density. The ROI region is a circle with a radius of a preset size, close to the size of the pupil.
[0071] S2.2, After determining the ROI region, select the events within this region and calculate the geometric centroid of these events, which is initially determined as the center point of the pupil.
[0072] S2.3, Using the center point of the pupil, combined with the preset pupil ellipse information P, screen the events within the neighborhood of the pupil ellipse information P to obtain the screened event set T':
[0073]
[0074] wherein, are the spatial coordinates of the event, is a preset neighborhood filtering threshold.
[0075] S2.4. For the filtered event set T′, the least squares method is applied for fitting, and the goal is to obtain a quadratic equation of an ellipse representing the pupil shape. However, due to the interference of noise and outliers, simply using the least squares method may lead to inaccurate fitting results, especially when the pupil boundary is irregular or there is occlusion. To solve this problem, the present invention introduces the Random Sample Consensus (RANSAC) algorithm, which automatically identifies and removes outliers or abnormal values by evaluating the quality of the fitting result. The robustness of the RANSAC algorithm enables the present invention to obtain a more accurate ellipse fitting result and improve the anti-interference ability of the model in the presence of noise and imperfect data.
[0076] It should be noted that the least squares method is a mathematical tool widely used in many disciplines of data processing such as error estimation, uncertainty, system identification, prediction, and forecasting, and is a well-known technology.
[0077] It should be noted that the ROI and pupil calculation module are computer programs, and their specific implementation can be referred to the above S2.1 - S2.4, which will not be elaborated here.
[0078] S3. The input of the vector calculation module is the event stream information and the pupil ellipse information P, and the output is the motion vector of the pupil center, that is, the magnitude and direction of the motion speed of the pupil center in the camera space. As Figure 3 shown, the workflow of the vector calculation module includes steps such as receiving, analyzing, and calculating the event data and the pupil ellipse data. The event stream information is sent into the vector calculation module, and combined with the pupil ellipse information P, the motion vector is calculated, specifically including:
[0079] S3.1. Time domain division: Select a time window length of △T (optional 3ms, 5ms, etc.) to divide the event stream information, and divide the event stream information into several event units Ei.
[0080] S3.2. Spatial domain division: For a certain event in the divided event unit Ei, select the events in its surrounding L×L (L can be 3 or 5) neighborhood to form an event set Q.
[0081] S3.3. Apply the principal component analysis method to the event set Q, and calculate 3 eigenvalues and 3 eigenvectors; among them, the eigenvector V corresponding to the smallest eigenvalue has a special meaning and can be used to solve the vector information.
[0082] It should be noted that the principal component analysis method is a well-known technology. On the premise of minimizing information loss, it synthesizes the original variables into a few new variables through linear combination; uses the new variables to replace the original variables to participate in data modeling, which can greatly reduce the computational workload in the analysis process; the selection of the new variables by the principal component is not a simple selection or rejection of the original variables, but the result of the recombination of the original variables. Therefore, it will not cause a large amount of loss of the original variable information and can represent most of the information of the original variables; at the same time, the selected new variables are not related to each other, which can effectively solve many problems brought by variable information overlap, multicollinearity, etc. to the analysis and application.
[0083] S3.4, using the eigenvalues and eigenvectors obtained in S3.3, fit the micro-plane , where is expressed as: , a, b, c, d are unknown parameters; the eigenvector corresponding to the minimum eigenvalue calculated in S3.3 corresponds to the parameters [a, b, c] to be found, that is , respectively represent the components of the eigenvector V in the x, y, and t directions.
[0084] S3.5, calculate the motion vector of the pupil center:
[0085]
[0086] In the formula, represents the motion vector of the center of the pupil ellipse information P in the x and y directions.
[0087] S3.6, evaluate all the points in the micro-plane fitted:
[0088]
[0089] In the formula, is the evaluation value of the point. Points less than the neighborhood screening threshold are regarded as points inside the plane, and points greater than the neighborhood screening threshold are regarded as points outside the boundary. When the number of points outside the boundary is too large, the plane fitting result is discarded.
[0090] S3.7, traverse each event in each event unit Ei, perform the operations of S3.2 - S3.6 on each event, and calculate a series of vector information , represents the vector information of the i-th point, represents the direction component of the i-th point, represents the Direction component.
[0091] S3.8, perform weighted summation on a series of vector information obtained in S3.7 , and finally obtain the motion vector ; vector is the output of the vector calculation module.
[0092] It should be noted that the vector calculation module is a computer program, and its specific implementation can be referred to the above S3.1 - S3.8, which will not be elaborated here.
[0093] S4, combine the motion vector obtained in S3 to update the pupil ellipse information P.
[0094] The vector calculation module outputs the motion vector of the pupil center, and this vector is used to update the pupil ellipse information. The specific method is as follows:
[0095]
[0096] Among them, is the pupil center information at the previous moment, is the updated pupil center information, is the motion vector of the pupil center obtained by the vector calculation module in S3.
[0097] Predicting the displacement of the pupil through the vector calculation module can significantly reduce the influence of noise interference and environmental light changes, and improve the positioning accuracy and smoothness. Compared with directly using event data fitting, this method effectively reduces the interference of artifacts such as reflected light spots on the results through motion trend prediction. At the same time, combined with optimization strategies such as adaptive time step, acceleration vector correction, and boundary constraint, its robustness and stability in dynamic scenarios can be further enhanced, making it suitable for pupil positioning tasks in complex environments. When the reflection point of the illumination source is close to the pupil edge, direct edge fitting may lead to unstable prediction, while vector update relies on the overall motion trend and can effectively avoid such problems. In addition, traditional methods have a relatively high jump in the pupil center coordinates and rely on a large amount of event information to obtain relatively accurate results, which not only reduces the tracking frequency of the eye movement tracking system but also affects the real-time performance. Through the vector calculation module, only less event information is required to generate a smooth motion vector, thereby achieving more stable and smooth prediction of the pupil center coordinates.
[0098] S5, combine the dynamic filtering method to smooth the pupil ellipse information P, specifically including:
[0099] The vector calculation module outputs the predicted pupil displacement, and the updated pupil coordinates can be obtained by combining the pupil displacement. Then, combine the dynamic filtering method to smooth the updated pupil coordinates.
[0100] The filtering process can be expressed as:
[0101]
[0102] Wherein, is the filtered signal at the current moment, that is, the smoothed pupil center coordinates, is the filtered signal at the previous moment, is the currently input signal, that is, the current pupil center coordinates, is the filtering coefficient, which is used to control the signal smoothing degree. The following formula is adopted:
[0103]
[0104] Wherein is the dynamically adjusted cut-off frequency, is the signal sampling frequency.
[0105] Method for calculating the dynamic cut-off frequency:
[0106]
[0107] Wherein, is the basic cut-off frequency, which determines the basic smoothing degree of the signal, generally set to 1000; min represents taking the minimum value, β is the adjustment factor of the response speed, and its value is in the range of [0,1], which controls the sensitivity of the filtering to the change rate of the signal and can be set according to experience; is the change rate of the input signal.
[0108] This dynamic filtering method is particularly suitable for real-time systems, can effectively smooth noise signals, and at the same time maintain the response ability to rapid changes. Compared with the Kalman filter, it is simpler and has less computational complexity, and is suitable for embedded systems or scenarios that require real-time processing. By adopting the dynamic filtering method and adaptive smoothing, the smoothing degree can be dynamically adjusted according to the change rate of the signal. For rapidly changing signals, the filter reduces smoothing and allows for rapid response.
[0109] S6. Transmit the filtered pupil ellipse information P to the line-of-sight regression and compensation model, and calculate the gaze point coordinates by combining the deep learning method. Specifically, it includes:
[0110] The line-of-sight regression and compensation model adopted by the present invention is a non-linear regression model based on a fully connected neural network. Specifically, the model takes the user's pupil center coordinates (two-dimensional coordinate values) as the input and outputs the predicted gaze point coordinates corresponding on the screen.
[0111] This model includes the following parts:
[0112] (1) Input layer: Accepts two-dimensional pupil center coordinates;
[0113] (2) Hidden layer: The model contains three hidden layers, each composed of 128, 64, and 32 neurons respectively, using the ReLU activation function to enhance the network's non-linear fitting ability;
[0114] (3) Output layer: Outputs two-dimensional predicted fixation point coordinates.
[0115] To improve the model's generalization performance and prevent overfitting, batch normalization and Dropout layers are added after each hidden layer to enhance the model's generalization ability for the non-linear relationship between pupil center coordinates and fixation points.
[0116] Nine or sixteen uniformly distributed calibration points are predefined on the screen as calibration references. These points cover different areas of the screen to enhance the model's prediction ability within the full field of view. When the user gazes at these known points in sequence, the camera captures the pupil center position coordinates in real time and matches these coordinates with the screen calibration points to generate data for model training. By inputting the matching data into the gaze regression and compensation model, training and optimization are performed. When the user's fixation point exceeds the area where the calibration points are located, the gaze point prediction error may increase. Therefore, the regression and compensation model believes that the relationship between the position of the pupil center and the screen coordinates is not a strictly linear relationship, but conforms to a non-linear mapping with a gaze surface close to the screen. The regression and compensation model captures the complex mapping characteristics between the pupil center and the fixation point by fitting this gaze surface. Especially in the edge area or when the calibration points are sparsely distributed, the farther the fixation point is from the center of the screen, the more compensation is obtained. This dynamic compensation mechanism significantly improves the accuracy of gaze point calculation.
[0117] Through this method, the regression and compensation model improves the spatial accuracy of gaze point prediction, making it applicable to a wider range of practical scenarios, such as screen interaction, gaze tracking analysis, and gaze point control in virtual reality. The combination of the non-linear mapping advantage of the neural network and the dynamic compensation mechanism lays an important foundation for future higher-precision gaze tracking technology.
[0118] See Figure 4 , an eye gaze point tracking device based on an event sensor provided in an embodiment of the present invention includes one or more processors for implementing a mobile robot trajectory tracking control method based on a high-order full drive theory in the above embodiment.
[0119] An embodiment of the near-eye gaze point tracking device based on event sensors of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. At the hardware level, as Figure 4 shown, it is a hardware structure diagram of any device with data processing capabilities where the near-eye gaze point tracking device based on event sensors of the present invention is located. In addition to Figure 4 the processor, memory, network interface, and non-volatile memory shown, generally, according to the actual functions of any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.
[0120] For the specific implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method, which will not be elaborated here.
[0121] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as these combinations of technical features do not conflict, they should be considered as within the scope described in this specification.
[0122] The embodiment of the present invention also provides a readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a mobile robot trajectory tracking control method based on the high-order full drive theory in the above embodiment.
[0123] The readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The readable storage medium can also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the readable storage medium can also include both an internal storage unit of any device with data processing capabilities and an external storage device. The readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or will be output.
[0124] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A near-eye fixation point tracking method based on event sensors, characterized in that, Including: S1. An infrared light source provides infrared illumination for the eye and the surrounding area of the eye. An event sensor collects information on the eye and the surrounding area of the eye, and obtains a series of event stream information generated by the change in light caused by movement. S2. The time window clustering event set T is sent to the ROI and pupil calculation module, and pupil ellipse information P is output, where the time window length is t. S3. The event stream information is sent to the vector calculation module, and combined with the pupil ellipse information P, the movement vector of the pupil center is calculated. S4. Combining the movement vector obtained in S3, the pupil ellipse information P is updated. S5. Combining the dynamic filtering method, the pupil ellipse information P is smoothed, including: Smoothing the updated pupil coordinates, and the filtering process is expressed as: Among them, is the signal after filtering at the current moment, that is, the smoothed pupil center coordinates, is the signal after filtering at the previous moment, is the currently input signal, that is, the current pupil center coordinates, is the filtering coefficient, used to control the signal smoothing degree, Adopt the following formula: wherein is the dynamically adjusted cut-off frequency, is the signal sampling frequency; Dynamic cut-off frequency calculation method: wherein, is the base cut-off frequency, and min represents taking the minimum value, is the adjustment factor of the response speed, is the change rate of the input signal; S6. The filtered pupil ellipse information P is sent to the gaze regression and compensation model, and combined with the deep learning method, the gaze point coordinates are calculated.
2. The method for near-eye fixation point tracking based on an event sensor according to claim 1, wherein The ROI and pupil calculation module combines the least squares method and the random sample consensus method to calculate the pupil ellipse information. The process includes: S2.
1. According to the event stream information obtained in S1, the event density is calculated, and the event density = the number of events within the time window / the time window t; the center position of the ROI is determined according to the event density. S2.
2. Calculate the geometric centroid of the events within the ROI area, and initially determine the geometric centroid as the center point of the pupil. S2.
3. Using the center point of the pupil, combined with the preset pupil ellipse information P, the events within the neighborhood of the pupil ellipse information P are screened to obtain the screened event set T'. Among them, is the spatial coordinate of the event, is a preset neighborhood screening threshold; S2.
4. For the screened event set T', the least squares method is applied for fitting to obtain a quadratic equation of an ellipse representing the pupil shape. At the same time, the random sample consensus algorithm is applied to automatically identify and remove outliers or anomalies by evaluating the quality of the fitting result.
3. The method for near-eye gaze point tracking based on an event sensor according to claim 1, wherein The said S3 includes: S3.1, Time domain division: Select the time window length to be Split the event stream information and divide the event stream information into several event units Ei; S3.
2. Spatial domain division: For a certain event in the segmented event unit Ei, select the events in its surrounding L×L neighborhood to form an event set Q. S3.
3. Apply the principal component analysis method to the event set Q to calculate a number of eigenvalues and a number of eigenvectors. S3.
4. Fit a tiny plane using the eigenvalues and eigenvectors obtained in S3.3 , where is expressed as: , where a, b, and c are unknown parameters. Among the eigenvalues calculated in S3.3, the eigenvector corresponding to the smallest eigenvalue corresponds to the parameters [a, b, c] to be found, that is , respectively represent the components of the eigenvector V in the x, y, and t directions; S3.
5. Calculate the movement vector of the pupil center: In the formula, respectively represent the motion vectors of the center of the pupil ellipse information P in the x and y directions: S3.6, evaluate all points within the obtained tiny plane : Wherein, is the evaluation value of the point. Points less than the neighborhood screening threshold are regarded as points inside the plane, and points greater than the neighborhood screening threshold are regarded as points outside the boundary. When the number of points outside the boundary is too large, discard this tiny plane ; S3.7, traverse the events within each event unit Ei, perform the operations of S3.2 - S3.6 on each event, and calculate a series of vector information , represents the vector information of the i-th point, represents the direction component of the i-th point, represents the direction component of the i-th point; S3.8, perform weighted summation on a series of vector information obtained in S3.7 , and finally obtain the pupil center motion vector , vector is the output of the vector calculation module.
4. A near-eye gaze point tracking method based on an event sensor according to claim 1, characterized in that In the said S4, the specific method of updating the pupil ellipse information P by combining the movement vector obtained in S3 is: wherein, is the pupil center information at the previous moment, is the updated pupil center information, is the motion vector of the pupil center obtained by the vector calculation module in S3.
5. A near-eye gaze point tracking method based on an event sensor according to claim 1, characterized in that, The said S6 includes: S6.
1. The user successively gazes at the known calibration point coordinates on the screen; records the calibration point coordinates and the corresponding pupil center coordinates, and sends the matching data to the gaze regression and compensation model to train and optimize the gaze regression and compensation model, and learn the complex mapping relationship between the pupil center and the screen gaze point coordinates to obtain the optimal gaze regression and compensation model. S6.
2. Send the pupil center coordinates to the optimal gaze regression and compensation model obtained in S6.1 to obtain the gaze point coordinates. S6.
3. The gaze regression and compensation model fits the gaze surface and compensates and regresses the gaze point coordinates obtained in S6.2 to obtain the final gaze point coordinates.
6. An event sensor-based near-eye gaze point tracking device, characterized in that, Comprising one or more processors for implementing an event sensor-based near-eye gaze point tracking method according to any one of claims 1-5.
7. A readable storage medium, characterized in that, Having a program stored thereon, which when executed by the processor, implements an event sensor-based near-eye gaze point tracking method according to any one of claims 1-5.
Citation Information
Patent Citations
Lightweight pulse wave signal noise removal and heart rate detection method
CN114098690A
First visual angle hand tracking system based on event camera and RGB camera and application
CN118505742A
Cited By
High-precision eye movement tracking system based on novel magnetoresistive effect quantum sensing
CN122195259A