Eye movement tracking method based on event camera

By reconstructing and training the Pulcim spot detection model using an event camera, and combining it with a 3D gaze calculation framework, the problem of insufficient accuracy in existing event camera-based eye tracking technologies is solved, achieving high-precision and low-power eye tracking suitable for complex scenarios.

CN121640554APending Publication Date: 2026-03-10ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing eye-tracking schemes based on event cameras mainly rely on pupil detection, which makes it difficult to achieve high-precision gaze estimation. Furthermore, traditional cameras are difficult to apply in low-power scenarios, and existing methods do not make full use of key points such as Purchin spots for 3D gaze estimation.

Method used

Eye-tracking event stream data is acquired by an event camera, three-channel frame images are reconstructed, Purchin spots are identified and labeled, a Purchin spot detection model is trained, and eye tracking is performed by combining a three-dimensional gaze calculation framework. The high temporal resolution and low power consumption of the event camera are used to adaptively track the eye movement process.

Benefits of technology

It achieves high-precision eye tracking, improves the accuracy of gaze estimation, enhances robustness in complex scenarios, reduces system power consumption, and meets application requirements under low light and high-speed motion conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640554A_ABST
    Figure CN121640554A_ABST
Patent Text Reader

Abstract

The invention discloses an eye movement tracking method based on an event camera, and belongs to the technical field of eye movement tracking. The method comprises the following steps: collecting eye movement event stream data of a human eye to be subjected to eye movement tracking by using an event camera, and performing event frame reconstruction on an event in the eye movement event stream data to obtain a plurality of three-channel frame images; inputting each three-channel frame image into a trained Purkinje spot detection model to obtain a candidate rectangular frame of Purkinje spots in each three-channel frame image; calculating the average coordinate of the centroids of the three channels of the image corresponding to each candidate rectangular frame, and taking the average coordinate as the estimated coordinate of the central point of the Purkinje spot in the candidate rectangular frame; and acquiring pupil center coordinates of the eyes to be subjected to three-dimensional sight line estimation, and performing three-dimensional sight line estimation by utilizing a three-dimensional sight line calculation framework in combination with the estimated coordinates of all the Purkinje spot center points to realize eye movement tracking. According to the invention, the application of the event camera in the eye movement tracking task is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of eye-tracking technology, specifically relating to an eye-tracking method based on an event camera. Background Technology

[0002] In recent years, the rapid development of computer vision technology has driven extensive research into non-contact eye-tracking technology based on video image analysis. The main task of eye tracking includes extracting feature information of the entire face or near-eye region from raw grayscale images or video streams for subsequent calculation of gaze points or gaze vectors (i.e., gaze). Research on gaze estimation methods can be mainly divided into three categories: methods based on two-dimensional mapping, methods based on three-dimensional models, and methods based on appearance. Among these, methods based on three-dimensional models have been proven to have good reliability and high accuracy. This method estimates the three-dimensional coordinates of the pupil center and the corneal curvature center by establishing a geometrical mathematical model of the human eye, thereby reconstructing the optical axis of the human eye. A typical approach is to first obtain the pixel coordinates of the pupil and the reflection points (Pulchin spots) of various light sources on the cornea in the camera image, then calibrate the three-dimensional world coordinates according to the actual physical positional relationship of the camera, light source, and human eye, and calculate the three-dimensional gaze vector.

[0003] Eye-tracking technology based on grayscale video streams is relatively mature. However, due to the limitations of frame rate, bandwidth, and power consumption of traditional CCD / CMOS cameras, eye-tracking methods based on camera images or video streams can only achieve a tracking frequency of a few hundred hertz, making it impossible to capture more refined eye movements. In addition, images captured by frame cameras contain a large amount of redundant information, which brings unnecessary overhead to downstream algorithm processing, making it difficult to further apply in low-power scenarios.

[0004] Event cameras, also known as dynamic vision sensors (DVS), are a new type of asynchronous sensor. Unlike traditional cameras that synchronously capture the entire scene at a fixed frequency, event cameras asynchronously record logarithmic intensity changes in brightness exceeding a threshold, generating a sparse spatiotemporal output event stream. Based on its unique sensing method, event cameras offer advantages over ordinary cameras, including high temporal resolution, high dynamic range, low latency, and low power consumption, making them highly promising for use in challenging and complex scenarios such as low-light conditions and high-speed motion.

[0005] Event cameras have brought new solutions to the field of eye tracking due to their unique advantages. Event cameras can dynamically respond to eye movements, adaptively outputting events based on movement rate and amplitude, significantly reducing power consumption compared to traditional cameras and making fuller use of camera bandwidth. However, due to the scarcity of eye-movement event datasets and the relatively late start of related research, the application of event cameras in this field is still relatively limited, and their accuracy still lags behind algorithms based on traditional grayscale video. Existing eye tracking schemes based on event cameras mainly detect the coordinates of the pupil center from the eye-movement event stream and use two-dimensional mapping methods for simple gaze point estimation, without fully utilizing key points such as the Pulcim spot for more accurate three-dimensional gaze estimation. Summary of the Invention

[0006] To address the problems in the prior art, this invention provides an eye-tracking method based on an event camera.

[0007] The technical solution of the present invention is as follows:

[0008] This invention discloses an eye-tracking method based on an event camera, comprising the following steps:

[0009] 1) Use an event camera to collect multiple first eye movement event stream data of the human eye, reconstruct event frames for each event in the first eye movement event stream data to obtain multiple first three-channel frame images, identify the Pulchin spot in each first three-channel event frame image, and label the candidate bounding box of each Pulchin spot.

[0010] 2) Obtain the Pulcyn spot detection model. Train the Pulcyn spot detection model using the labeled first three-channel event frame images to obtain the trained Pulcyn spot detection model.

[0011] 3) Use an event camera to collect second eye movement event stream data of the human eye to be tracked, and reconstruct event frames from the events in the second eye movement event stream data to obtain multiple second and third channel frame images;

[0012] 4) Input each second and third channel frame image into the trained Purchin spot detection model to obtain the candidate bounding boxes of Purchin spots in each second and third channel frame image;

[0013] 5) Calculate the coordinates of the pixels at the centroid positions of each channel of the image corresponding to each candidate rectangle in step 4), and then calculate the average coordinates of the centroids of the three channels. Use the average coordinates as the estimated coordinates of the center point of the Pulchin spot in the candidate rectangle.

[0014] 6) Obtain the pupil center coordinates of the eye to be estimated in 3D, and then combine them with the estimated coordinates of all the center points of the Pulcim spots. Use the 3D gaze calculation framework to perform 3D gaze estimation and realize eye tracking.

[0015] Further, in step 1), the first eye movement event stream data of the human eye is collected using an event camera, including:

[0016] Several LEDs are installed around the human eye for illumination, ensuring that each LED can produce a reflective spot in the corneal area of ​​the human eye;

[0017] The eye movement process of a human eye is captured by an event camera, and the event stream data triggered by the eye movement process is obtained. This event stream data is used as the first eye movement event stream data.

[0018] Further, in step 1), the event frame reconstruction includes:

[0019] 11) Initialize multiple three-channel RGB images with the same resolution as the event camera;

[0020] 12) On the time scale of the first eye-tracking event stream data, every n events are combined into an event set. An event set is used to perform event accumulation reconstruction on a three-channel RGB image to obtain a first three-channel frame image.

[0021] 13) Traverse each event set of the first eye movement event stream data to obtain multiple first three-channel frame images.

[0022] Further, in step 12), each event has the coordinates of the pixel that generated the event and information about the event polarity; the step of using an event set to perform event accumulation reconstruction on a three-channel RGB image includes:

[0023] Iterate through each event in the event set, and color the pixel corresponding to the event in the three-channel RGB image according to the event polarity. That is, if the event polarity is positive, add a preset value to the current value of the red channel and the blue channel of the pixel corresponding to the event in the three-channel RGB image. If the event polarity is negative, add a preset value to the current value of the green channel and the blue channel of the pixel corresponding to the event in the three-channel RGB image. Finally, obtain a first three-channel frame image corresponding to the event set.

[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0025] 1) This method fills a technical gap in existing event-camera-based eye-tracking solutions: Current event-camera-based eye-tracking solutions still primarily rely on pupil detection, and the gaze estimation accuracy they can achieve is still far from practical application levels. High-precision gaze estimation requires other key points besides the pupil center, with the most representative method being the use of the Pulcyn spot. This method can be integrated with existing event-driven pupil detection solutions, and is expected to achieve high-precision eye tracking.

[0026] 2) Adaptive eye tracking: The event camera dynamically responds to changes in brightness and can adaptively track the brightness changes caused by eye movements. It generates an event stream of corresponding density according to the frequency of eye movements. Compared with fixed frame rate global exposure imaging, it reduces visual redundancy information, makes full use of camera bandwidth, and is expected to further reduce system power consumption when combined with dedicated hardware (such as neural computing chips).

[0027] 3) Robustness in high dynamic and high-speed scenes: This invention utilizes the inherent advantages of event cameras to be robust to the detection of low-light, high dynamic range, and fast-moving scenes, meeting the technical requirements of eye-tracking in complex scenes. Attached Figure Description

[0028] Figure 1 This is a schematic diagram illustrating the Pulchin spot number and center coordinates in an eye-tracking grayscale video frame;

[0029] Figure 2 This is a schematic diagram showing the Pulchin spot number and center coordinates in a three-channel frame image.

[0030] Figure 3 This is a schematic diagram showing the change of the X-axis coordinate of the Pulchin spots (numbers 1 and 4) over time.

[0031] Figure 4 This is a schematic diagram showing the change of the Y-axis coordinate of the Pulchin spots (numbers 1 and 4) over time.

[0032] Figure 5 This is a schematic diagram of the distance error of the Pulchin spots (numbers 1 and 4) at a time threshold of 5000 microseconds.

[0033] Figure 6 This is a schematic diagram of the distance error of the Pulchin spots (numbers 1 and 4) at a time threshold of 2000 microseconds.

[0034] Figure 7 This is a schematic diagram showing the change of the X-axis coordinate of the Pulchin spots (numbers 2 and 3) over time.

[0035] Figure 8 This is a schematic diagram showing the change of the Y-axis coordinate of the Pulchin spots (numbers 2 and 3) over time.

[0036] Figure 9This is a schematic diagram of the distance error of the Pulchin spots (numbers 2 and 3) at a time threshold of 5000 microseconds.

[0037] Figure 10 This is a schematic diagram of the distance error of the Pulchin spots (numbers 2 and 3) at a time threshold of 2000 microseconds.

[0038] Figure 11 This is a flowchart of the eye-tracking method based on an event camera according to the present invention. Detailed Implementation

[0039] The present invention will be further described and illustrated below with reference to specific embodiments. The embodiments described are merely examples of the content of this disclosure and do not limit the scope of the invention. The technical features of each embodiment in the present invention can be combined accordingly, provided that there is no mutual conflict.

[0040] Current eye-tracking methods based on event cameras focus on extracting the pupil center position from event data and estimating the gaze point coordinates accordingly. To the best of our knowledge, there are no methods or systems specifically designed to extract the Pulcim spot from the event stream.

[0041] To fill the gap in existing research, this invention focuses on the design of an algorithm to directly estimate the coordinates of the Purchin spot from an event stream. It does not rely on fine hardware control and adaptively calculates the frequency based on eye movements. It can be combined with existing pupil detection methods based on event cameras to improve the accuracy of downstream gaze estimation tasks.

[0042] The technical solution of the present invention will be further described below with reference to the illustrative figures and embodiments.

[0043] This invention aims to combine the advantages of event cameras and traditional eye-tracking technology to estimate the coordinates of the Purchin spot from event stream data. It can also be easily combined with existing event stream pupil detection methods and 3D gaze estimation methods to achieve eye tracking based on pure event streams, which is expected to further expand the application of event cameras in eye-tracking tasks.

[0044] Example 1:

[0045] The steps of the eye-tracking method based on an event camera proposed in this invention are as follows: Figure 11 As shown, the specific steps are as follows:

[0046] (1) Construction of the event eye-tracking Pulcim patch dataset

[0047] This step aims to build a dataset for training the event stream Pulchin spot detection model (i.e., the Pulchin spot detection model).

[0048] (1.1) Eye-tracking data acquisition: Near-eye eye-tracking event data is acquired using an event camera such as DAVIS346 to obtain eye-tracking event stream data. During acquisition, the eye area should occupy as much of the event camera sensor's field of view as possible, and the eye area should be clearly focused within the host computer software screen captured by the event camera. The event camera output should include event data. At the same time, several LEDs are installed around the eye using a frame-like light fixture for illumination. Other lighting methods can also be used, as long as each LED can produce a clear reflective spot in the corneal area.

[0049] (1.2) Three-channel frame image annotation: The acquired eye-tracking event stream data is a data packet containing a series of (x,y,p,t) event formats, where x represents the x-axis coordinate of the pixel that generated the event; y represents the y-axis coordinate of the pixel that generated the event; p represents the event polarity, p=1 represents a positive event, i.e., brightness increases; p=0 represents a negative event, i.e. brightness decreases; and t is the timestamp of the event.

[0050] Using tools provided by the event camera manufacturer, the raw data can be converted into a file format suitable for program processing, such as a .npy file. The converted event file is then read, and a fixed number of events (n) are converted into three-channel frame images using the following event frame reconstruction algorithm:

[0051] 1) Initialize multiple three-channel RGB images frame(H, W, 3) with the same resolution as the event camera, i.e., create multiple pure black images. The size (height H, width W) of this pure black image is exactly the same as the sensor resolution of the event camera. It has three color channels: Red, Green, and Blue.

[0052] 2) On the time scale of eye-tracking event stream data, every n events form an event set. An event set is then used to perform event accumulation and reconstruction on a three-channel RGB image to obtain a three-channel frame image. Specifically, performing event accumulation and reconstruction on a three-channel RGB image using an event set to obtain a three-channel frame image includes:

[0053] Iterate through the n events in the event set. For each event (x,y,p,t), if p=1, then frame[x,y,0] =min(frame[x,y,0] += bias, 255); if p=0, then frame[x,y,1] = min(frame[x,y,1] += bias, 255). Regardless of whether p=1 or 0, frame[x,y,2] = min(frame[x,y,2] += bias, 255), where bias is generally set to 20-40. This means that the pixel at the location of each event in the three-channel RGB image is colored according to the event polarity. Specifically: if the event polarity is positive, a preset value is added to the current value of the red channel and the blue channel at the location of the event in the three-channel RGB image; if the event polarity is negative, a preset value is added to the current value of the green channel and the blue channel at the location of the event in the three-channel RGB image; finally, a three-channel frame image corresponding to the event set is obtained.

[0054] 3) Traverse each event set of the eye-tracking event stream data to obtain multiple three-channel frame images.

[0055] The number of events n depends on the specific event camera model. Generally, it is advisable to select events that can clearly distinguish the outline of the eyes, and the number is about 800-1500.

[0056] Next, candidate bounding boxes for each Pulchin blob (each three-channel frame image may have one or more Pulchin blobs) are labeled in each three-channel frame image. This can be done using common dataset annotation software in an object detection format, or manually. The labeled three-channel frame images are shown below. Figure 2 As shown, note that Figure 2 The index of each Pulchin spot can be arbitrarily assigned, as long as the index definition during annotation and training is consistent. In addition, if only some of the Pulchin spots generated by LEDs are visible in the event frame, only the visible spots need to be labeled, and attention should be paid to maintaining the consistency of the index. Training with this labeling allows the model to fully learn the spatial feature distribution of the spots, and can still give the index of the visible spots even in the case of partial missing spots.

[0057] Finally, the three-channel event frame images with all annotations are used to construct the event eye-tracking Purchin spot dataset.

[0058] (2) Training of the event flow Pulcim spot detection model

[0059] This step aims to enable neural network models, which are originally designed for object detection in frame images, to identify candidate bounding boxes for Purchin spots from three-channel frame images.

[0060] (2.1) Selection of Event Stream Pulchin Spot Detection Model Architecture: Considering the balance between model parameter count and detection accuracy, the CenterNet model with ResNet-18, ResNet-50, or MobileNetV2 as the backbone network was used during training, along with input images adapted to 512x512 and 416x416 resolutions respectively. Generally speaking, the larger the parameter count of the same type of neural network model and the higher the resolution of the input image, the higher the accuracy of the detection results.

[0061] (2.2) Training parameter settings: The event eye-tracking Purchin spot dataset is divided into training and testing sets in an appropriate ratio. If the amount of data in the event eye-tracking Purchin spot dataset is small, data augmentation operations such as rotation, cropping, and contrast adjustment can be performed on the data in the event eye-tracking Purchin spot dataset to enhance the generalization ability of the model.

[0062] During training, various hyperparameters, including learning rate, number of epochs, batch size, optimizer, warming up, and learning rate decay, should be set reasonably according to the actual training situation and dataset size to ensure stable convergence of the model. The specific training settings are based on those in the paper Zhou X, Wang D, Krähenbühl P. Objects as points[J]. arXiv preprint arXiv:1904.07850, 2019. The loss function used in the training of the event flow Pultchin spot detection model is the same as the CenterNet loss function in Zhou X, Wang D, Krähenbühl P. Objects as points[J]. arXiv preprint arXiv:1904.07850,2019.

[0063] The CenterNet network treats objects to be detected as points, predicting the location of the object's center point and its width and height. Given an input image... , Its size is Where W represents width, H represents height, and 3 represents the RGB three-channel array. The network output is a heatmap containing the distribution of center keypoints, where each keypoint corresponds to an object category, and the generated center point is denoted as . R is the output step size, and C is the number of keypoint types (i.e., object categories). In the heatmap, the position of each keypoint represents the center position of an object category. When the heatmap's predicted values... When, it indicates that the corresponding object has been detected at that location; if Then it is considered as the background.

[0064] Loss function during CenterNet training As shown in equation (1):

[0065] (1)

[0066] It mainly consists of three parts. and All are weights. Generally, 0.1 is used. Generally, the value is 1, and the definitions of each part of the loss are as follows:

[0067] During network training, the ground truth keypoints of the input image for each object category can be represented as follows: The coordinates of the key points after downsampling become Then, the real key points are mapped onto the heatmap using a Gaussian kernel. Above, that is ,in The target size adapts to the standard deviation. The coordinates are mapped onto the low-resolution feature map; when two objects of the same class have overlapping Gaussian kernels, the larger value is selected. The focal loss function is used as the keypoint prediction loss function. As shown in equation (2):

[0068] (2)

[0069] in, and It is a hyperparameter of focal loss. For the real key points, It is an image The number of key points.

[0070] To compensate for the discretization error caused by downsampling during feature extraction, it is necessary to predict and compensate for the local offset of each center point. In the prediction phase, a unified offset prediction model is used to compensate for the center point offsets of all target categories. The predicted center point offset is represented as... The predicted center point bias is trained using the L1 loss function, and the center point bias loss is... As shown in equation (3).

[0071] (3)

[0072] Assumption It is a category The The bounding box of an object, whose center point coordinates can be represented as: Key point estimator This is used to predict the center point location of all targets. For each detected target, its size information is further regressed, specifically represented as... and These represent the coordinates of the top-left and bottom-right corners of the target bounding box, respectively. To reduce computational complexity and improve efficiency, only a single size prediction is performed for all object categories, meaning that targets of different categories will share the same size prediction model. Loss assessment size prediction loss As shown in equation (4). The predicted size for each target k. The regression size value for each objective k.

[0073] (4)

[0074] After training, the performance of each model is evaluated on the test set at different epochs, and the best-performing model file is selected for subsequent processing.

[0075] Table 1

[0076]

[0077] Table 1 compares the classification accuracy (Average Precision, AP) and Mean Average Precision (mAP) of different backbones (the backbone network in a neural network model) for each Pulchin patch in the test set under two input sizes, as well as the required budget (GFLOPs, billions of floating-point operations per second) and number of parameters for each model. A larger number of model parameters and a higher input resolution result in greater computational overhead but also higher accuracy. AP1 represents... Figure 1 (or Figure 2 The average precision of the Pulchin spot in index 1; AP2 represents... Figure 1 (or Figure 2 The average precision of the Pulchin spot in index 2; AP3 represents... Figure 1 (or Figure 2 The average precision of the Pulchin spot in index 3; AP4 represents Figure 1 (or Figure 2 The average precision of the Pulchin spot in sequence 4 of the sample.

[0078] (3) Obtaining the three-channel frame images of the human eye to be tracked

[0079] An event camera is used to collect eye movement event stream data of the human eye to be tracked. Event frames are reconstructed from the events in the eye movement event stream data to obtain multiple three-channel frame images.

[0080] (4) Obtaining the candidate rectangle of the Pulchin spot

[0081] Input each three-channel frame image in step (3) into the trained event stream Purchin spot detection model in step (2) to obtain the candidate bounding box of each Purchin spot in each three-channel frame image.

[0082] (5) Regional post-processing

[0083] This step aims to further extract the center coordinates of each Pulchin blob from the candidate bounding boxes of Pulchin blobs obtained from the event flow Pulchin blob detection model trained in the previous step.

[0084] (5.1) Purchin Spot Feature Analysis: Due to the presence of noise during the event camera imaging process, and the accumulation of events within a certain time period in the three-channel frame images, it can be understood as the average eye movement during that time period. The event camera triggers events when the brightness changes to a threshold, and the corresponding region of the Purchin spot exhibits the characteristic of rapid alternation between positive and negative events, which is reflected in the three-channel frame images as a higher gray value. Therefore, the pixel centroid positions of the three channels of the image within each candidate rectangle can be used as the coordinates of the spot center.

[0085] (3.2) Average centroid calculation: For each candidate region (actually a portion of the image region cropped by several small rectangles) within a three-channel frame image, calculate the X-axis and Y-axis coordinates of the centroid of each channel pixel. Then calculate the average coordinates of the three-channel centroids and use the average coordinates as the estimated coordinates of the center points of each Pulchin spot, in pixels. Finally, traverse each three-channel frame image to obtain the estimated coordinates of the center points of all Pulchin spots.

[0086] In this three-channel frame image, the three channels are the red channel, the green channel, and the blue channel. The method for calculating the X-axis coordinate of the centroid of each pixel in each channel is as follows: multiply the X-axis coordinate of each pixel in the image corresponding to the candidate rectangle by the pixel value of that pixel in the red channel (or green channel or blue channel), then accumulate the multiplication results, and divide the accumulated result by the sum of the pixel values ​​of all pixels in the red channel (or green channel or blue channel) to obtain the X-axis coordinate of the pixel at the centroid position of the red channel (or green channel or blue channel).

[0087] The method for calculating the Y-axis coordinate of the centroid of each pixel in each channel is as follows: multiply the Y-axis coordinate of each pixel in the image corresponding to the candidate rectangle by the pixel value of that pixel in the red channel (or green channel or blue channel), then accumulate the multiplication results, and divide the accumulated result by the sum of the pixel values ​​of all pixels in the red channel (or green channel or blue channel) to obtain the Y-axis coordinate of the pixel at the centroid position of the red channel (or green channel or blue channel).

[0088] The method for calculating the average coordinates of the centroids of the three channels is as follows: For each candidate rectangle, add the x-axis coordinates of the pixels at the centroid positions of the three channels of the image corresponding to the candidate rectangle, and then divide the sum by three to obtain the average x-axis coordinates of the centroids of the three channels; add the y-axis coordinates of the pixels at the centroid positions of the three channels of the image corresponding to the candidate rectangle, and then divide the sum by three to obtain the average y-axis coordinates of the centroids of the three channels, and finally obtain the average coordinates.

[0089] (6) Three-dimensional line-of-sight estimation

[0090] Most existing pupil detection methods based on event cameras use three-channel frame images as input, such as E-track (Li N, Bhat A, Raychowdhury A. E-Track: Eye Tracking with Event Camera for Extended Reality (XR) Applications[C]. 2023 IEEE 5th International Conference on Artificial Intelligence Circuits and Systems (AICAS). Hangzhou, China:IEEE, 2023: 1-5.), thus allowing for parallel extraction of pupil center coordinates from three-channel frame images. After obtaining the estimated coordinates of the pupil center and the center points of each Purkinje spot, the coordinates can be directly provided to existing three-dimensional gaze calculation frameworks (such as Lidegaard M, Hansen DW, Krüger N. Head mounted device for point-of-gazeestimation in three dimensions[C]. Proceedings of the Symposium on EyeTracking Research and Applications. Safety Harbor Florida: ACM, 2014: 83-86.) to realize the complete eye tracking process, obtain the three-dimensional gaze direction vector of the eye and the gaze point coordinates of the eye in three-dimensional space (both of which are quantities that change over time). The gaze point coordinates of the eye in three-dimensional space indicate the specific location where the eye's gaze falls in the three-dimensional environment.

[0091] In a specific embodiment of the present invention, to verify the effectiveness of the present invention, therefore, during step 1), an event camera such as the DAVIS346 is used to simultaneously acquire eye-tracking grayscale video while acquiring eye-tracking event stream data.

[0092] Then add step (1.3) after step (1.2).

[0093] (1.3) Eye-tracking grayscale video annotation: The eye-tracking grayscale video output synchronously by event cameras such as DAVIS346 has the same spatial resolution as the eye-tracking event stream data, but the frame rate is low, only about 25fps. The eye-tracking process is very fast, and often only a few pictures can be taken for the entire process of the pupil moving from the leftmost to the rightmost. However, the number of three-channel frame images reconstructed in the same time period is as many as dozens, which makes it difficult for subsequent event-grayscale comparison verification. To address this, this invention introduces an existing advanced event-based grayscale video interpolation framework (Tulyakov S, Gehrig D, Georgoulis S, et al. Time Lens: Event-based Video Frame Interpolation[C]. 2021 IEEE / CVFConference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2021: 16150-16159.) to interpolate eye-tracking grayscale video to increase the frame rate to a level comparable to the equivalent frame rate of a three-channel frame image. This framework can be fine-tuned based on the eye-tracking grayscale video to further improve interpolation accuracy. Subsequently, the coordinates of the center points of each Pulchin spot are manually labeled in the interpolated high frame rate eye-tracking grayscale video, such as... Figure 1 As shown. The serial numbers 1, 2, 3, and 4 represent the serial numbers of the actual installed LEDs. This matching relationship can be used to calibrate the relative positions of each LED in the actual physical space, and then calculate the three-dimensional line of sight.

[0094] Then, after completing step (5) of the present invention, the estimated coordinates of the center points of the Pulcyn spots obtained in step (5) are compared and verified with the coordinates of the center points of each Pulcyn spot manually annotated in the high frame rate eye-tracking grayscale video to verify the detection accuracy level based on pure event stream. This step aims to verify whether the detection results based on pure event stream are truly reliable.

[0095] (a) Data selection: Due to the adaptive nature of the event camera, the detection results of fast eye movements are more concentrated and the equivalent frame rate is higher. Therefore, when performing the verification, the data are selected from the high frame rate eye movement grayscale video and eye movement event stream data within the significant eye movement time period.

[0096] (b) Specific verification: Calculate the Euclidean distance between the estimated coordinates of the center point of the Pulcyn spot obtained in step (5) and the coordinates of the center points of each Pulcyn spot manually annotated in the high frame rate eye-tracking grayscale video, and draw the motion trajectory of each Pulcyn spot coordinate over time in the two modes of eye-tracking event stream data and high frame rate eye-tracking grayscale video. Figure 3 , Figure 4 , Figure 7 and Figure 8The motion trajectories of the four Pulchin spots corresponding to the four LEDs were plotted within the same time period (i.e. Figure 1 and Figure 2 The motion trajectories of the four Pulchin spots numbered 1, 2, 3, and 4 in the diagram), where, Figure 1 and Figure 2 Pulchin spots with the same serial number belong to the same Pulchin spot, only... Figure 1 The Pulcyn spot in the eye-tracking grayscale video. Figure 2 This refers to the Pulchin spot in a three-channel frame image. Figure 5 , Figure 6 , Figure 9 and Figure 10 This represents the pixel distance changes of the coordinates of the two data modalities in the motion trajectory of the Pulcyn spot. When calculating the pixel distance, it is necessary to ensure that the timestamp deviation between the two data points from the event stream and the grayscale video stream is within a certain time threshold (usually a few milliseconds). Otherwise, rapid eye movements will cause significant positional shifts, resulting in large distance deviations.

[0097] (c) Results Analysis: Figures 3 to 10 The results show that the Puerchin spot detection of pure event stream in high-speed eye-tracking scenarios achieves an average detection error of less than 2 pixels and has good motion adaptability, which can provide relatively reliable estimation results for downstream tasks.

[0098] Accelerated inference deployment

[0099] This step aims to demonstrate that the overall algorithm is well-compatible with existing efficient deployment frameworks, facilitating rapid application in subsequent real-world scenarios.

[0100] (A) Model Format Conversion: The Purkinje spot detection model file trained in step (2) is usually in .pth or .pt format and can be quickly deployed using NVIDIA's TensorRT framework. TensorRT is a high-performance deep learning acceleration inference engine provided by NVIDIA, enabling low-latency deployment of neural network models on NVIDIA GPUs. TensorRT achieves lightweight and high-throughput computation through various optimizations, such as low-precision quantization, merging of model layers, and hardware optimization.

[0101] (B) Inference Acceleration: The selected hardware devices were an Intel i7-1165 G7 CPU and an NVIDIA RTX 4060 Laptop GPU. Table 2 shows the single inference time of six models with three different backbone networks and two different input resolutions on the CPU and GPU. Table 2 shows the running time of each model when deployed on the CPU and GPU. The model on the GPU was converted to TensorRT format to take advantage of the GPU's acceleration effect and meet the requirements of high-speed real-time eye tracking.

[0102] Overall, the acceleration effect is significant. The model architecture can be flexibly adjusted according to the actual computing power. Most models can complete inference within 3 milliseconds, meeting the requirements of real-time high frame rate eye tracking.

[0103] Table 2

[0104]

[0105] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. Those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.

Claims

1. An event camera based eye tracking method, characterized in that, The method comprises the following steps: 1) acquiring a plurality of first eye movement event stream data of a human eye by using an event camera, performing event frame reconstruction on events in each first eye movement event stream data, obtaining a plurality of first three-channel frame images, and identifying a Purkinje image in each first three-channel event frame image and labeling a candidate rectangular frame of each Purkinje image; 2) obtaining a Purkinje image detection model, training the Purkinje image detection model by using the labeled first three-channel event frame image, and obtaining a trained Purkinje image detection model; 3) acquiring second eye movement event stream data of a human eye to be tracked by using an event camera, performing event frame reconstruction on events in the second eye movement event stream data, and obtaining a plurality of second three-channel frame images; 4) inputting each second three-channel frame image into the trained Purkinje image detection model to obtain a candidate rectangular frame of a Purkinje image in each second three-channel frame image; 5) calculating coordinates of pixel points at a centroid position of each channel of an image corresponding to each candidate rectangular frame in step 4), and then calculating average coordinates of the centroids of the three channels, and taking the average coordinates as estimated coordinates of a Purkinje image center point in the candidate rectangular frame; 6) obtaining a pupil center coordinate of an eye to be tracked, and then combining the estimated coordinates of all Purkinje image center points to perform three-dimensional gaze estimation by using a three-dimensional gaze calculation framework to realize eye movement tracking.

2. The event camera based eye tracking method of claim 1, wherein, In step 1), the first eye movement event stream data of the human eye is acquired by using the event camera, comprising: a plurality of LEDs are installed around the human eye for illumination to ensure that each LED can generate a reflection spot in the corneal region of the human eye; an event camera is used to capture the eye movement process of the human eye to obtain event stream data triggered by the eye movement process, and the event stream data is taken as the first eye movement event stream data.

3. The event camera based eye tracking method of claim 1, wherein, In step 1), the event frame reconstruction comprises: 11) initializing a plurality of three-channel RGB images with the same resolution as the event camera; 12) in the time scale of the first eye movement event stream data, an event set is formed by every n events, and an event set is used to accumulate and reconstruct a three-channel RGB image to obtain a first three-channel frame image; 13) each event set of the first eye movement event stream data is traversed to obtain a plurality of first three-channel frame images.

4. The event camera based eye tracking method of claim 3, wherein, Each event in step 12) has the information of the coordinates of the pixel point generating the event and the event polarity; the event accumulation reconstruction of a three-channel RGB image by using an event set comprises: Traverse each event in the event set, color the pixel point where the event is located on the three-channel RGB image according to the event polarity of each event, that is, if the event polarity of the event is a positive event, increase the current value of the red channel of the pixel point where the event is located on the three-channel RGB image by a preset value, and at the same time, increase the current value of the blue channel of the pixel point where the event is located on the three-channel RGB image by a preset value; if the event polarity of the event is a negative event, increase the current value of the green channel of the pixel point where the event is located on the three-channel RGB image by a preset value, and at the same time, increase the current value of the blue channel of the pixel point where the event is located on the three-channel RGB image by a preset value; finally, a first three-channel frame image corresponding to the event set is obtained.

5. The event camera based eye tracking method of claim 1, wherein, In step 2), the Purkin spot detection model is a CenterNet model with a ResNet-18 model, a ResNet-50 model or a MobileNetV2 model as a backbone network.

6. The event camera based eye tracking method of claim 1, wherein, In step 4), each second three-channel frame image has one or more Purkin spots.

7. The event camera based eye tracking method of claim 1, wherein, The three channels in the second three-channel frame image are a red channel, a green channel and a blue channel respectively. In step 5), the coordinates of the pixel points at the channel centroid positions of the image corresponding to each candidate rectangular frame are calculated, including: 51) multiply the x-axis coordinate of each pixel point in the image corresponding to the candidate rectangular frame by the pixel value of the pixel point in a certain channel, then accumulate the multiplication results, divide the accumulation result by the sum of the pixel values of all pixel points in the channel, and obtain the x-axis coordinate of the pixel point at the channel centroid position; multiply the y-axis coordinate of each pixel point in the candidate rectangular frame by the pixel value of the pixel point in the channel, then accumulate the multiplication results, divide the accumulation result by the sum of the pixel values of all pixel points in the channel, and obtain the y-axis coordinate of the pixel point at the channel centroid position; finally, the coordinates of the pixel point at the channel centroid position of the candidate rectangular frame are obtained; 52) repeat step 51) to obtain the coordinates of the pixel point at the channel centroid position of the image corresponding to each candidate rectangular frame; 53) replace the channel, and steps 51)-52) are repeated to obtain the coordinates of the pixel points at the channel centroid positions of the image corresponding to each candidate rectangular frame.

8. The event camera based eye tracking method of claim 1, wherein, In step 5), the average coordinates of the centroids of the three channels are calculated, including: For each candidate rectangular frame, add the x-axis coordinates of the pixel points at the three channel centroid positions of the image corresponding to the candidate rectangular frame, then divide the addition result by three to obtain the x-axis average coordinate of the centroid of the three channels; add the y-axis coordinates of the pixel points at the three channel centroid positions of the image corresponding to the candidate rectangular frame, then divide the addition result by three to obtain the y-axis average coordinate of the centroid of the three channels.