A lightweight eye-tracking method based on half-eye image reconstruction

By combining half-eye image reconstruction and dynamic contextual memory with a 3D attention mechanism, the problems of high hardware cost, high power consumption and high computational complexity of traditional eye-tracking technology are solved, achieving high-precision, low-latency lightweight eye-tracking and improving system miniaturization and human-computer interaction performance.

CN120635497BActive Publication Date: 2025-10-28ANHUI AVATAR SANJIEWAI TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511113629.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-10-28
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing eye-tracking technologies require high-resolution, complete eye image acquisition at the image processing level, resulting in high hardware costs, high power consumption, large size, and high computational complexity, which affects lightweight design and real-time performance.

Method used

A half-eye image reconstruction method is adopted, which combines dynamic contextual memory and three-dimensional attention mechanism. The image acquisition unit acquires local half-eye images of the eyeball, performs low-complexity feature extraction and feature fusion, and generates a lightweight eye-tracking trajectory.

Benefits of technology

It achieves high-precision, low-latency eye tracking, significantly reduces hardware resource requirements, and improves the system's miniaturization, low power consumption, and natural perception performance in human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635497B_ABST
    Figure CN120635497B_ABST
Patent Text Reader

Abstract

This invention discloses a lightweight eye-tracking method based on half-eye image reconstruction, belonging to the field of image processing. The method includes: acquiring a half-eye image of the eyeball through an image acquisition unit; extracting low-complexity features to generate a basic feature map; querying a dynamic context memory based on the basic feature map to obtain relevant context feature vectors; dynamically fusing the basic feature map and context feature vectors to generate a task-oriented feature map; and performing eye movement parameter regression based on the task-oriented feature map to generate a lightweight eye tracking trajectory. This invention utilizes local half-eye images of the pupil and iris, combined with dynamic context memory and a three-dimensional attention mechanism for feature fusion, achieving lightweight, high-precision, and low-latency eye tracking, significantly improving human-computer interaction and perception performance. This method effectively solves the challenge of traditional eye-tracking technologies maintaining accuracy while simultaneously achieving system miniaturization, low power consumption, and a good user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and in particular to a lightweight eye-tracking method based on half-eye image reconstruction. Background Technology

[0002] Image processing technology plays a central role in machine vision, human-computer interaction, and various display systems. Particularly in the field of eye tracking, the acquisition and analysis of eye image data to accurately perceive user gaze points, intentions, and states has become a key technology for improving the naturalness of interaction and system intelligence. With the increasing popularity of head-mounted display systems such as virtual and augmented reality, higher demands are being placed on the image processing capabilities and system integration of embedded eye tracking modules.

[0003] However, existing eye-tracking technologies, especially in image processing, often face numerous challenges. To obtain sufficient accuracy, traditional solutions typically require acquiring high-resolution, complete eye images, necessitating image acquisition units with wide fields of view and large size. When integrated into lightweight head-mounted devices, such designs not only significantly increase hardware costs and power consumption but also negatively impact wearing comfort and experience due to obstruction of the user's vision. Furthermore, the highly complex feature extraction and processing of complete eye images introduces a significant computational burden and data transmission latency, contradicting the trend towards lightweight and real-time systems. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides a lightweight eye-tracking method based on half-eye image reconstruction. It employs an image acquisition unit to obtain partial half-eye images of the eyeball and combines dynamic contextual memory and a three-dimensional attention mechanism for feature fusion. This method enables high-precision, low-latency eye-tracking while significantly reducing hardware resource requirements and substantially improving system miniaturization, low power consumption, and natural perception performance in human-computer interaction.

[0005] The above objectives can be achieved through the following approach:

[0006] A lightweight eye-tracking method based on half-eye image reconstruction includes: acquiring a half-eye image of the eyeball in real time using an image acquisition unit; performing low-complexity feature extraction on the half-eye image to generate a basic feature map of the current frame, wherein the basic feature map is used to represent the local state information of the eyeball; querying a preset dynamic context memory based on the basic feature map to obtain relevant context feature vectors, wherein the memory stores eye-tracking context feature vectors of historical frames; dynamically fusing the basic feature map and relevant context feature vectors to generate a task-oriented feature map; and performing eye-tracking parameter regression based on the task-oriented feature map to generate a lightweight eye tracking trajectory.

[0007] Optionally, the method further includes: filtering valid tracking data for updating the dynamic context memory based on the confidence level or change magnitude of the eye tracking results; encoding the valid tracking data into a new context feature vector; integrating the new context feature vector into the dynamic context memory and adjusting the eye tracking context feature vectors of historical frames stored in the memory.

[0008] Optionally, low-complexity feature extraction of the half-eye image includes: identifying the pupil region and iris region in the half-eye image; determining whether the pupil region presents a bright pupil condition by comparing the brightness of the pupil region with a preset brightness threshold; if yes, extracting eyeball feature points from the pupil region and its surroundings; if no, performing contrast enhancement and edge detection on the half-eye image and extracting eyeball feature points from the pupil region and iris region; and generating a basic feature map of the current frame based on the eyeball feature points.

[0009] Optionally, querying the preset dynamic context memory includes: comparing the similarity between the basic feature map and the eye-tracking context feature vectors of historical frames; and selecting the eye-tracking context feature vector of the historical frame with the highest similarity as the relevant context feature vector.

[0010] Optionally, dynamically fusing the basic feature map and the relevant contextual feature vector includes: configuring a learnable attention mechanism based on the local key regions of the eye indicated by the eye feature points; and fusing the basic feature map and the relevant contextual feature vector through the learnable attention mechanism to generate a task-oriented feature map.

[0011] Optionally, generating a lightweight eye tracking trajectory includes: performing eye movement parameter regression using a pre-trained eye movement parameter regression model based on the task-oriented feature map to obtain the real-time three-dimensional gaze direction and gaze point of the eye; judging the motion characteristics of the three-dimensional gaze direction and gaze point, dynamically adjusting the output frequency and accuracy of the tracking data, and generating eye movement trajectory data that meets the lightweight requirements; compressing and encoding the eye movement trajectory data, and outputting it in a structured and compact format to generate a lightweight eye tracking trajectory.

[0012] Optionally, the learnable attention mechanism includes: inputting the information of key local regions of the eye indicated by the eye feature points into a preset lightweight neural network model, and outputting a dynamic attention weight map; and fusing the basic feature map and relevant context feature vectors based on the dynamic attention weight map.

[0013] Optionally, identifying the pupil region and iris region in the half-eye image includes: dynamically adjusting the output intensity of the supplementary light source integrated in the image acquisition unit based on the brightness of the pupil region and a preset brightness threshold, and adopting differentiated supplementary lighting strategies under bright pupil conditions and non-bright pupil conditions; implementing high-resolution feature encoding for the pupil region, implementing interference-resistant robust feature encoding for the iris region, and achieving cross-scale feature fusion through a feature pyramid structure.

[0014] Optionally, the dynamic fusion of the basic feature map and the relevant context feature vector further includes: constructing a three-dimensional attention mechanism that fuses spatial location weights and temporal decay factors, wherein the spatial location weights are determined by the topological distribution of eye feature points, and accurately focusing on key areas of the eye during the fusion process; inputting task-oriented feature maps from multiple consecutive frames into a temporal prediction network to generate predicted values ​​of eye movement trajectories and co-optimizing them with real-time detection results; and employing a differential coding scheme based on motion features to selectively compress and transmit lightweight tracking trajectory data, thereby ensuring low-latency data output and supporting dynamic perception of user gaze points, focal points, and attention regions.

[0015] Based on the same inventive concept, the present invention also provides a lightweight eye-tracking system based on half-eye image reconstruction, the system comprising:

[0016] The image acquisition module is used to acquire half-eye images of the eyeball in the current frame in real time;

[0017] The feature extraction module is used to perform low-complexity feature extraction on the half-eye image to generate the basic feature map of the current frame;

[0018] The context memory and query module is used to query a preset dynamic context memory based on the basic feature map to obtain relevant context feature vectors;

[0019] The feature fusion module is used to dynamically fuse the basic feature map and relevant context feature vectors to generate a task-oriented feature map;

[0020] The eye-tracking parameter regression module is used to perform eye-tracking parameter regression based on the task-oriented feature map to generate a lightweight eye tracking trajectory.

[0021] Compared with the prior art, the present invention has the following advantages:

[0022] 1. This invention acquires partial half-eye images of the eyeball through an image acquisition unit and introduces low-complexity feature extraction and a differentiated illumination strategy based on pupil brightness. This effectively solves the problems of excessive system size, weight, and power consumption caused by traditional eye tracking relying on whole-eye image acquisition. Simultaneously, by combining high-resolution encoding of the pupil region and robust feature encoding of the iris region, high-precision eye movement information can be reconstructed from limited half-eye images, significantly improving the system's miniaturization and lightweight design.

[0023] 2. This invention constructs a dynamic context memory based on fundamental feature maps. Through an intelligent compression management mechanism, it can store and update eye-tracking context feature vectors from historical frames, effectively overcoming the information isolation problem in traditional eye tracking when dealing with continuous and complex eye movements. Combined with a learnable 3D attention mechanism, it achieves dynamic deep fusion of fundamental feature maps and context feature vectors, enabling the system to accurately focus on key local regions of the eye, significantly improving the accuracy and robustness of eye-tracking parameter regression.

[0024] 3. In generating lightweight eye-tracking trajectories, this invention significantly improves the smoothness and robustness of eye-tracking trajectories through collaborative optimization using a temporal prediction network and real-time detection results. Simultaneously, it employs a motion-feature-based differential coding scheme to selectively compress and transmit the tracking trajectory data, ensuring low-latency data output. This effectively addresses the challenges of high transmission bandwidth and processing latency inherent in traditional methods, enabling lightweight eye-tracking data to be applied in real-time to various scenarios, significantly improving system response speed and overall performance.

[0025] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart illustrating a lightweight eye-tracking method based on half-eye image reconstruction according to an embodiment of the present invention.

[0028] Figure 2 This is a distribution diagram of the similarity of feature vectors in the dynamic context memory of this invention.

[0029] Figure 3 This is a convergence curve of the eye movement parameter regression model according to an embodiment of the present invention.

[0030] Figure 4 This is a heatmap of half-eye image feature coding intensity according to an embodiment of the present invention.

[0031] Figure 5 This is a schematic diagram of the structure of a lightweight eye-tracking system based on half-eye image reconstruction according to an embodiment of the present invention. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Reference Figure 1 One embodiment of the present invention proposes a lightweight eye-tracking method based on half-eye image reconstruction. The method involves acquiring half-eye images of the eyeball in real time through an image acquisition unit, performing low-complexity feature extraction on the half-eye images to generate a basic feature map of the current frame, querying a dynamic context memory based on the basic feature map to obtain relevant context feature vectors, and then dynamically fusing the basic feature map and relevant context feature vectors to generate a task-oriented feature map. Finally, eye-tracking parameter regression is performed based on the task-oriented feature map to generate a lightweight eye tracking trajectory. This method can achieve high-precision, low-latency eye tracking, while significantly reducing the demand for hardware resources and significantly improving the system's miniaturization, low power consumption, and natural perception performance in human-computer interaction.

[0034] The method described in this embodiment specifically includes:

[0035] The image acquisition unit acquires half-eye images of the current frame in real time.

[0036] Specifically, the step of acquiring a half-eye image of the eye in real time for the current frame is performed by a miniaturized high-speed camera. This camera is mounted near the user's eye, on the bridge of the nose of the wearable device, or on the edge of the frame. Its main task is to capture images of a local area of ​​the eye, i.e., a half-eye image. To achieve the goals of lightweight design and minimal occlusion, the camera is configured with optimized optical components to minimize its size and weight while ensuring the capture of clear local images including the pupil and iris. A unidirectional or offset image acquisition module can be used to capture a local effective area of ​​the eye to form a half-eye image. The optical path and sensor configuration of the image acquisition module are optimized to significantly reduce the size and weight of the optical components while maintaining eye-tracking accuracy and minimizing interference with user comfort and field of vision.

[0037] Low-complexity feature extraction is performed on the half-eye image to generate a basic feature map of the current frame, wherein the basic feature map is used to represent the local state information of the eyeball;

[0038] Specifically, after acquiring a half-eye image, the image is preprocessed, including grayscale conversion, noise filtering, and image correction. Then, image region recognition is performed to identify the pupil and iris regions in the half-eye image. Feature extraction is then performed, which involves determining whether the pupil region exhibits a bright pupil condition. If a bright pupil condition is observed, eyeball feature points are extracted from the pupil region and its surroundings; otherwise, the image undergoes contrast enhancement and edge detection before feature point extraction. Simultaneously, the intensity of the supplementary lighting integrated into the image acquisition unit is dynamically adjusted based on the brightness of the pupil region, employing a differentiated supplementary lighting strategy to optimize image acquisition quality. High-resolution feature encoding is implemented for the pupil region, and robust feature encoding for the iris region is implemented to resist interference. Cross-scale feature fusion is achieved through a feature pyramid structure, organizing these extracted features into a data structure that can concisely represent local eyeball movements or shape changes, generating a basic feature map.

[0039] Based on the aforementioned basic feature map, a preset dynamic context memory is queried to obtain relevant context feature vectors, wherein the memory stores eye-tracking context feature vectors of historical frames.

[0040] Specifically, after generating the basic feature map of the current frame, it is compared with the eye-tracking context feature vectors of historical frames stored in the dynamic context memory for similarity. The eye-tracking context feature vector of the historical frame with the highest similarity is selected as the relevant context feature vector. This process aims to quickly identify the historical eye-tracking pattern that best matches or is most relevant to the current local eye state through quantitative comparison, providing temporal contextual information for subsequent feature fusion. The dynamic context memory also filters effective tracking data based on the confidence level or magnitude of change in the eye-tracking results, encodes it into new context feature vectors and integrates them into the memory, and adjusts the eye-tracking context feature vectors of historical frames to ensure the dynamism and effectiveness of the memory. When the memory capacity reaches a preset condition, intelligent compression is performed based on the spatiotemporal correlation of the stored feature vectors, retaining representative feature patterns to maintain the lightweight and efficient nature of the memory.

[0041] The basic feature map and relevant context feature vectors are dynamically fused to generate a task-oriented feature map;

[0042] Specifically, the feature fusion step integrates the base feature map of the current frame with relevant contextual feature vectors obtained from a dynamic contextual memory to generate a more discriminative task-oriented feature map that can directly guide subsequent eye-tracking parameter regression. This process configures a learnable attention mechanism based on key local regions of the eye indicated by eye feature points. This attention mechanism inputs information from key local regions of the eye into a pre-defined lightweight neural network model, outputs a dynamic attention weight map, and uses this weight map to fuse the base feature map and relevant contextual feature vectors. In addition, a three-dimensional attention mechanism is constructed that integrates spatial location weights and a temporal decay factor, where the spatial location weights are determined by the topological distribution of eye feature points to accurately focus on key eye regions during the fusion process. Simultaneously, task-oriented feature maps from multiple consecutive frames are input into a temporal prediction network to generate predicted values ​​for eye-tracking trajectories and co-optimize with real-time detection results to improve the smoothness and robustness of the lightweight tracking trajectory.

[0043] Eye movement parameters are regressed based on the task-oriented feature map to generate a lightweight eye tracking trajectory.

[0044] Specifically, after generating the task-oriented feature map, the eye-tracking parameter regression step aims to transform it into specific eye-tracking parameters and ultimately generate a lightweight eye tracking trajectory. This process uses a pre-trained eye-tracking parameter regression model to perform eye-tracking parameter regression based on the task-oriented feature map, obtaining the real-time three-dimensional gaze direction of the eye, including pitch, yaw, and roll angles, as well as the gaze point, including two-dimensional coordinates on the screen or three-dimensional coordinates in three-dimensional space. Motion characteristics are judged on the three-dimensional gaze direction and gaze point by analyzing eye movement speed, acceleration, and the stability of the gaze area to determine whether the current eye movement is rapid saccade, smooth tracking, or fixed gaze. Based on the motion characteristic judgment results, the output frequency and accuracy of the tracking data are dynamically adjusted; for example, the output frequency and accuracy are reduced during rapid saccade and increased during fixed gaze, to generate eye movement trajectory data that meets the lightweight requirements. Finally, the eye movement trajectory data is compressed and encoded using a differential coding scheme based on motion features, transmitting only the changes between adjacent frames, rather than the complete data, thus ensuring low-latency data output. The encoded data is output in a structured and compact format to generate a lightweight eye tracking trajectory and supports dynamic perception of the user's gaze point, focus, and attention area.

[0045] By acquiring half-eye images in real time through the image acquisition unit, performing low-complexity feature extraction and dynamic context memory query, and combining a learnable 3D attention mechanism for feature fusion, a lightweight tracking trajectory is finally generated through eye movement parameter regression. This enables high-precision, low-latency, real-time, and robust tracking of eye movements, significantly improving the system's miniaturization, low power consumption, natural perception performance of human-computer interaction, and user experience.

[0046] Optionally, the method further includes:

[0047] Based on the confidence level or magnitude of change of the eye-tracking results, select valid tracking data for updating the dynamic context memory.

[0048] Specifically, the confidence level of eye-tracking results can be comprehensively evaluated using multi-dimensional indicators, including the output probability of the eye-tracking parameter regression model, the reprojection error of the tracking points, or the consistency of eye-tracking parameters across multiple frames. The magnitude of change measures the difference between the eye-tracking parameters of the current frame and the average value of the previous frame or multiple frames, determining whether a significant change in eye movement has occurred. Only when the confidence level is higher than a preset threshold and the magnitude of change is within an acceptable range is the corresponding tracking data marked as valid tracking data for subsequent updates to the memory bank.

[0049] For example, suppose the eye-tracking results include a 3D gaze direction and a gaze point. If the magnitude change of the gaze direction vector is within 5%, and the pixel displacement of the gaze point on the screen is less than 10 pixels, and the regression model gives a prediction confidence of more than 90%, then the tracking data of the current frame is considered valid.

[0050] Encode the effective tracking data into a new context feature vector;

[0051] Specifically, the effective tracking data is input into a pre-defined lightweight encoder. This encoder can compress multiple dimensions of information, such as the eye's three-dimensional gaze direction, gaze point coordinates, pupil diameter, blink state, and current local eye features, into a fixed-length numerical vector—the new context feature vector—through dimensionality reduction or nonlinear transformation. The encoding process aims to extract the most core and representative eye movement state information for easy storage and computation.

[0052] For example, the encoder can be a miniaturized fully connected neural network that maps multidimensional effective tracking data into a 64-dimensional or 128-dimensional feature space, and the output is the new context feature vector.

[0053] The new contextual feature vector is integrated into the dynamic contextual memory bank, and the eye-tracking contextual feature vectors of historical frames stored in the memory bank are adjusted.

[0054] Specifically, the process of integrating new contextual feature vectors into a dynamic contextual memory can employ weighted averaging, sliding window updates, or similarity-based replacement strategies. The memory can be a fixed-length queue or a storage structure with limited capacity. When the memory is not full, new vectors are added directly; when the memory is full, the oldest vectors can be discarded based on a time decay factor, or the least representative historical vectors can be replaced based on their similarity to the new vector. Adjusting the historical frame eye-tracking contextual feature vectors stored in the memory aims to maintain its timeliness and effectiveness.

[0055] For example, the memory maintains a queue of context feature vectors from the most recent N frames. When a new context feature vector is generated, it is added to the tail of the queue, while the oldest vector at the head of the queue that is similar to the newly added context feature vector is removed.

[0056] Optionally, low-complexity feature extraction of the half-eye image includes:

[0057] Identify the pupil region and iris region in the half-eye image;

[0058] Specifically, identifying the pupil and iris regions can be accomplished through preliminary analysis of the acquired half-eye image. This involves applying image segmentation algorithms, based on pixel intensity, color distribution, or a preset shape template, to locate and distinguish the circular pupil region and the surrounding iris ring within a general area.

[0059] For example, binarization based on the Otsu thresholding method is used, combined with circular Hough transform to detect the approximate position and size of the pupil, and then the Canny edge detection algorithm is used to delineate the inner and outer edges of the iris.

[0060] By comparing the brightness of the pupil region with a preset brightness threshold, it is determined whether the pupil region presents a bright pupil condition. If yes, eye feature points are extracted from the pupil region and its surroundings. If no, contrast enhancement and edge detection are performed on the half-eye image, and eye feature points are extracted from the pupil region and iris region.

[0061] Specifically, determining whether the pupil region exhibits a bright pupil condition is based on comparing the average pixel brightness of the pupil region with a preset brightness threshold. If the brightness exceeds the threshold, it is considered a bright pupil condition. In a bright pupil condition, light reflected from the fundus makes the pupil appear bright in the image. In this case, feature points are directly extracted from the pupil region and its surroundings; these points can be key pixels or points of maximum local intensity. If a bright pupil condition is not present, the half-eye image requires additional image processing, including contrast enhancement to highlight eye details and edge detection to accurately delineate the boundaries of the pupil and iris. Then, richer ocular feature points are extracted from the processed pupil and iris regions.

[0062] For example, the threshold is set to a grayscale value of 180 (255 being the maximum brightness). If the average brightness of the pupil region exceeds this value, feature points are uniformly sampled at and around the center of the pupil. If it is below this value, histogram equalization is performed to enhance contrast, and the Sobel operator is used to detect edges. Then, corner features of the pupil and iris are extracted from the image after contrast enhancement and edge detection.

[0063] Based on the eye feature points, a basic feature map of the current frame is generated.

[0064] Specifically, generating the base feature map for the current frame involves structuring the eye feature points extracted in the previous steps. These feature points can be their two-dimensional or three-dimensional coordinates, local texture descriptors, gradient information, or morphological parameters. This data is encoded and arranged according to a preset format to form a compact and efficient data structure, the base feature map, which represents the local state information of the eye, facilitating subsequent processing.

[0065] For example, the base feature map can be a fixed-size feature vector that contains the normalized pupil center coordinates, pupil diameter, coordinates of several sampling points on the inner and outer edges of the iris, and histogram descriptors of the local gradient orientations of these points.

[0066] Optionally, the preset dynamic context memory for querying includes:

[0067] Compare the similarity between the base feature map and the eye-tracking context feature vector of the historical frame;

[0068] Specifically, comparing the similarity between the base feature map and the eye-tracking contextual feature vectors of historical frames involves quantitatively evaluating the two feature representations. Assume the base feature map of the current frame is represented as... ∈ The eye-tracking context feature vectors of historical frames stored in the memory bank are represented as follows: ∈ ,in It is a feature dimension. The similarity measurement function can be defined as follows: Similarity can be calculated using cosine similarity. For two non-zero vectors... and ,

[0069] Their cosine similarity Defined as:

[0070] ,

[0071] in, and They are vectors and The There are several components. The range of cosine similarity values ​​is [...]. Similarity is defined as a range between 0 and 1, with larger values ​​indicating higher similarity. Another similarity measure is the Gaussian kernel function, which maps similarity to the range [0,1] and assigns higher weights to data points with high similarity.

[0072]

[0073] in, Representing vectors and The Euclidean distance between them The kernel width parameter controls the rate at which similarity decays. These similarity metrics reflect the distance or angular relationship between feature vectors to indicate their proximity.

[0074] like Figure 2As shown, the similarity distribution between the current base feature map and all historical feature vectors in the dynamic context memory is displayed. Peaks indicate highly relevant historical context information, while troughs indicate irrelevant information.

[0075] The eye-tracking contextual feature vector of the historical frame with the highest similarity is selected as the relevant contextual feature vector.

[0076] Specifically, selecting the eye-tracking context feature vector of the historical frame with the highest similarity is a process of sorting and selecting based on the similarity values ​​calculated in the previous step. In the dynamic context memory, assuming it stores the eye-tracking context feature vectors of N historical frames { , ,…, }. For the current basic feature map Calculate the similarity set between it and all historical feature vectors in the memory. From the set Find the historical feature vector corresponding to the maximum value in the data; this is the relevant context feature vector. .

[0077] ,

[0078] This selection process ensures that the acquired contextual information is most relevant to the current eye state.

[0079] For example, 1000 historical context feature vectors are stored in the memory. For the current base feature map, the system selects the historical feature vector with the highest similarity value (closest to 1) as the relevant context feature vector for the current frame. If multiple historical vectors have the same highest similarity, the one with the most recent or earliest timestamp can be selected.

[0080] Optionally, dynamically fusing the base feature map and the relevant context feature vector includes:

[0081] Based on the key local regions of the eye indicated by the eye feature points, a learnable attention mechanism is configured.

[0082] Specifically, configuring a learnable attention mechanism involves inputting information about key local regions of the eye, indicated by eye feature points, into a pre-defined lightweight neural network model. This model receives this information as input and outputs a dynamic attention weight map. The attention mechanism aims to learn, identify, and weight the most important local information for eye tracking, thereby precisely focusing on key eye regions during the fusion process.

[0083] For example, a 128-dimensional feature vector containing a local key region including the center of the pupil, the iris boundary point, and the corner of the eye is input into a fully connected neural network with two hidden layers, which outputs a weight distribution map corresponding to each feature point.

[0084] The learnable attention mechanism is used to fuse the basic feature map and the relevant contextual feature vectors to generate a task-oriented feature map.

[0085] Specifically, a learnable attention mechanism is used to fuse the base feature map and relevant contextual feature vectors. This is achieved by weighting these two types of features using a dynamic attention weight map generated in the previous step. The fusion operation can be element-wise multiplication followed by addition, or it can be performed through a fusion layer after a concatenation operation. A task-oriented feature map is generated, which is the fused result and can more effectively guide subsequent eye-tracking parameter regression because it contains local visual information from the current frame as well as relevant patterns extracted from historical context, with key information given higher weights.

[0086] For example, if the dynamic attention weight graph is ∈ The basic feature map is ∈ The relevant context feature vector is ∈ Then the task-oriented feature map It can be generated in the following way:

[0087] ,

[0088] Here, ⊙ represents element-wise multiplication. This fusion process ensures that key local information is fully utilized while also integrating temporal contextual information, thus improving the discriminative power of the features.

[0089] Optionally, the lightweight tracking trajectory for generating the eyeball includes:

[0090] Based on the task-oriented feature map, eye movement parameters are regressed using a pre-trained eye movement parameter regression model to obtain the real-time three-dimensional gaze direction and gaze point of the eyeball;

[0091] Specifically, the eye-tracking parameter regression model receives task-oriented feature maps. ∈ As input. The model is a lightweight neural network. By training with a large amount of eye-tracking data, the model can learn the mapping relationship from feature maps to eye-tracking parameters. The model outputs the real-time 3D gaze direction of the eyeball. and fixation point The gaze direction can be represented by a unit vector, showing the orientation of the eye in three-dimensional space. The gaze point can be in two-dimensional screen coordinates or three-dimensional spatial coordinates.

[0092] like Figure 3 As shown, the convergence trend of the loss function of the eye-tracking parameter regression model with the number of training iterations is demonstrated during the training process, reflecting the effectiveness of the model's learning.

[0093] For example, a pre-trained neural network model containing convolutional and fully connected layers receives a 128-dimensional task-oriented feature map and outputs a 5-dimensional vector containing eye pitch angle, yaw angle, and two-dimensional screen coordinates.

[0094] The motion characteristics of the three-dimensional gaze direction and gaze point are judged, and the output frequency and accuracy of the tracking data are dynamically adjusted to generate eye movement trajectory data that meets the requirements of lightweight design.

[0095] Specifically, motion characteristic determination involves the three-dimensional gaze direction of consecutive frames. and fixation point Analysis was performed. Eye movement speed was calculated. acceleration and the stability of the gaze area Kinematic parameters are used to determine whether the current eye movement is rapid saccadic, smooth tracking, or fixed fixation.

[0096] ,

[0097] ,

[0098] Based on the judgment results, the system dynamically adjusts the output frequency of the tracking data. and accuracy For example, during rapid saccades, the output frequency and accuracy can be appropriately reduced to conserve resources due to the high-speed eye movements. Conversely, during fixed gaze or smooth tracking, the output frequency and accuracy can be increased to ensure high precision.

[0099] > ,

[0100] ,

[0101] in and These are preset thresholds. These adjustments help generate eye movement trajectory data that meets lightweight requirements.

[0102] For example, when the eye is detected to have moved more than 50 pixels in 0.5 seconds and at a speed of less than 5 pixels per frame, the system will reduce the output frequency from 60Hz to 30Hz and reduce the precision of the gaze point coordinates from floating point to one decimal place.

[0103] The eye movement trajectory data is compressed and encoded, and output in a structured and compact format to generate a lightweight eye tracking trajectory.

[0104] Specifically, eye movement trajectory data is compressed and encoded, primarily using a differential coding scheme based on motion features. For continuous eye movement parameter sequences, such as fixation point coordinates... During encoding, what is transmitted is the difference between the current frame and the previous frame. Instead of the original coordinates.

[0105] ,

[0106] This differential coding significantly reduces data volume, especially during periods of stable eye movement where the difference is very small, allowing for further quantization or entropy coding to ensure low-latency data output. The encoded data is output in a structured, compact format that has been optimized to reduce redundant information and packet header overhead. The final output data stream is a lightweight eye tracking trajectory, supporting dynamic perception of the user's gaze point, focus, and attention area.

[0107] For example, the eye movement trajectory data contains 60 frames per second of two-dimensional coordinates of the gaze point. After differential coding, if the gaze point of the current frame is ( The previous frame was (). , If ), then what is actually transmitted is ( ),in , These differences are typically much smaller than the original coordinate values, thus achieving data compression.

[0108] Optionally, the learnable attention mechanism includes:

[0109] The information of key local regions of the eyeball indicated by the eyeball feature points is input into a preset lightweight neural network model, and a dynamic attention weight map is output.

[0110] Specifically, a lightweight neural network model is constructed capable of learning and identifying the importance of key information from partial, incomplete eye images. Instead of simply mapping inputs to weights, the model is designed to recognize the topological relationships and geometric distributions between eye feature points, inferring potential full-eye attention regions even with the limited field of view of a half-eye image. It receives information about key local eye regions indicated by eye feature points, which is standardized and normalized to form a compact input vector. The model learns the contribution of different local regions to the eye-tracking task through its internal low-computational-complexity layers (e.g., depthwise separable convolutional layers or fully connected layers trained with quantization perception) and outputs a fine-grained dynamic attention weight map. This weight map reflects which eye feature regions are most worthy of attention in the current frame, even with incomplete data.

[0111] For example, a geometric feature vector containing the pupil center, iris key points, and their relative positions is input into a small neural network composed of a multilayer perceptron. This network is pruned and quantized to generate an attention weight map with extremely low computational resources, matching the dimensions of the base feature map, where the weights for the pupil and iris activity areas are higher than those for areas such as the corner of the eye.

[0112] Based on the dynamic attention weight map, the basic feature map and the relevant context feature vector are fused.

[0113] Specifically, this fusion process goes beyond simple weighted summation; it employs an adaptive feature enhancement and context injection mechanism. The system utilizes a dynamic attention weight map to non-linearly weight each part of the base feature map, thereby strengthening the local information in the base feature map most relevant to the current eye-tracking task. Simultaneously, relevant contextual feature vectors, as supplementary information, are integrated with the weighted base feature map through a gating mechanism or soft thresholding function. This fusion is a dynamic adaptation across modalities (visual local features and temporal contextual features), ensuring that the final task-oriented feature map not only contains precise local information from the current frame but also incorporates prior knowledge of historical eye-tracking trends. Furthermore, all information is assigned appropriate weights according to its importance, achieving a compact and efficient representation of features and avoiding interference from redundant information in subsequent regression.

[0114] For example, a dynamic attention weight map is applied to the channel dimension of the base feature map, and channel-level features are weighted through element-wise multiplication. Then, the weighted base feature map and the relevant context feature vectors are fused through a learnable gating unit, which dynamically adjusts the proportion of information flow based on the similarity and importance of the two contents to generate the final task-oriented feature map.

[0115] Optionally, identifying the pupil region and iris region in the half-eye image includes:

[0116] Based on the brightness of the pupil region and a preset brightness threshold, the output intensity of the supplementary light source integrated in the image acquisition unit is dynamically adjusted, and a differentiated supplementary light strategy is adopted in bright pupil conditions and non-bright pupil conditions.

[0117] Specifically, the system analyzes the brightness of the pupil region in the half-eye image acquired by the image acquisition unit in real time and compares it with a preset brightness threshold that can be dynamically adjusted according to the ambient light intensity. When the pupil region brightness is higher than the threshold, indicating a bright pupil condition, the supplementary light source reduces its output intensity or even shuts down temporarily to avoid overexposure and save energy. Conversely, when the brightness is lower than the threshold, indicating a non-bright pupil condition (or dark pupil), the supplementary light source precisely increases its output intensity according to the degree of brightness deficiency, ensuring that eye details remain clear even in low-light environments. This differentiated supplementary lighting strategy aims to optimize the acquisition quality of the half-eye image while effectively controlling power consumption.

[0118] For example, when the average gray value of the pupil region detected by the infrared supplementary light source integrated in the image acquisition unit exceeds 200, the output power is reduced by 50%; when the average gray value is below 50, the output power is increased to the maximum, so as to obtain a clear half-eye image under different lighting conditions.

[0119] High-resolution feature encoding is performed on the pupil region, robust feature encoding is performed on the iris region, and cross-scale feature fusion is achieved through a feature pyramid structure.

[0120] Specifically, for the pupil region, the system performs high-resolution feature encoding to capture its fine shape, size, and position information. This is typically achieved through sub-pixel-level edge detection or circular / elliptical fitting algorithms to ensure the accuracy of eye-movement parameters. For the iris region, robust feature encoding is implemented to resist interference, focusing on extracting its unique texture patterns and structural features. These features are highly robust to illumination, blur, or partial occlusion, for example, using Gabor filters or local binary pattern descriptors. Subsequently, cross-scale feature fusion is achieved through a feature pyramid structure, effectively integrating features extracted from different resolution levels. This ensures that the final base feature map balances detail accuracy and global contextual information, thereby improving the accuracy and resistance to interference in eye state recognition under complex environments.

[0121] like Figure 4 As shown, the intensity distribution of feature encoding in a half-eye image is illustrated. The pupil region exhibits the highest intensity, representing the precision of its high-resolution encoding; the iris region has the second highest intensity, reflecting the focus of its robust feature extraction; the intensity gradually decreases in the surrounding regions, demonstrating the effective fusion of cross-scale information by the feature pyramid structure.

[0122] For example, the pupil center coordinates are fitted using a Gaussian kernel function, accurate to the sub-pixel level. Feature encoding of the iris region is achieved by applying multiple sets of Gabor filters at different scales to extract its orientation and frequency texture, generating multi-scale feature maps. These feature maps are then fused from top to bottom and horizontally connected through a feature pyramid network, ultimately generating a unified feature representation rich in semantic information.

[0123] Optionally, dynamically fusing the base feature map and the relevant context feature vector further includes:

[0124] A three-dimensional attention mechanism is constructed that integrates spatial location weights and temporal decay factors, wherein the spatial location weights are determined by the topological distribution of eye feature points, and the key regions of the eye are precisely focused on during the fusion process;

[0125] Specifically, a three-dimensional attention mechanism is constructed to achieve more refined feature fusion. This mechanism introduces a temporal decay factor on top of two-dimensional spatial attention (focusing on which region of the eye image to concentrate on). The spatial position weights are not fixed but dynamically generated based on the real-time topological distribution of eye feature points, ensuring that the most critical local regions in the current eye pose receive the highest weight. The temporal decay factor is adjusted based on the freshness of historical context feature vectors or their time interval with the current frame, making recent and highly relevant contextual information more influential during fusion. This three-dimensional attention mechanism can precisely focus on key eye regions during the fusion process, achieving efficient information aggregation and redundancy removal.

[0126] For example, the system learns a spatial weight map through a small network, which has higher weights in the pupil center and iris edge regions, and lower weights in regions such as the sclera. Meanwhile, the fusion weights of contextual feature vectors from further back in time decay exponentially.

[0127] The task-oriented feature maps of multiple consecutive frames are input into the temporal prediction network to generate predicted values ​​of eye-tracking trajectories and to perform collaborative optimization with real-time detection results.

[0128] Specifically, the co-optimization of prediction and correction improves the smoothness and robustness of the tracking trajectory. Task-oriented feature maps from multiple consecutive frames are input into a specially designed temporal prediction network. This network is pre-trained to identify temporal patterns and trends in eye movements and generates predicted values ​​for the eye movement trajectory based on the input sequence. These predictions are then co-optimized with the real-time detection results of the current frame. Co-optimization can employ Kalman filtering, particle filtering, or some form of sensor fusion algorithm to weight and fuse the prediction results with the real-time measurement results, thereby correcting for transient noise or delays in real-time detection and improving the coherence and accuracy of the eye movement trajectory.

[0129] For example, a sequence containing task-oriented feature maps from the most recent five frames is input into a Long Short-Term Memory (LSTM) network. The network outputs predicted gaze direction and gaze point for the next frame. These predictions are then fused with real-time detection results obtained from an eye-tracking parametric regression model for the current frame through an adaptive Kalman filter to generate smoother and more accurate eye-tracking trajectory data.

[0130] A differential coding scheme based on motion features is adopted to selectively compress and transmit lightweight tracking trajectory data, thereby ensuring low-latency data output and supporting dynamic perception of user gaze points, focus areas, and attention regions.

[0131] Specifically, to ensure low-latency data output, this step innovatively utilizes the inherent characteristics of eye-tracking data for compressed transmission. The system employs a motion-feature-based differential coding scheme, meaning that instead of transmitting the complete values ​​of eye-tracking trajectory data, it transmits the changes or differences in eye-tracking parameters between the current frame and the previous frame. These differences are typically much smaller than the original values, allowing for representation and transmission with fewer bits. Selective compression transmission means that the system dynamically adjusts the compression ratio based on the type of eye movement (e.g., saccades, smooth tracking, fixation) or its importance. For example, during rapid saccades, a higher compression ratio can be used due to relatively lower accuracy requirements; while during fixed fixations, a lower compression ratio or lossless compression is used to ensure high accuracy. This dynamic strategy significantly reduces transmission bandwidth requirements and lowers transmission latency, thereby ensuring real-time, low-latency data output and strongly supporting the dynamic perception of the user's gaze point, focus, and attention area.

[0132] For example, the eye-tracking trajectory data includes the coordinates of the gaze point and the pupil size. When a fixed gaze is detected, the system transmits only the complete data for every 10 frames, with intermediate frames transmitting the displacement difference compared to the previous frame, and the difference is quantized for further compression. If saccades occur, the transmission frequency is increased, but the displacement difference is also quantized with a coarser granularity.

[0133] Based on the same inventive concept, such as Figure 5 As shown, the present invention also provides a lightweight eye-tracking system based on half-eye image reconstruction, the system comprising:

[0134] The image acquisition module is used to acquire half-eye images of the eyeball in the current frame in real time;

[0135] The feature extraction module is used to perform low-complexity feature extraction on the half-eye image to generate the basic feature map of the current frame;

[0136] The context memory and query module is used to query a preset dynamic context memory based on the basic feature map to obtain relevant context feature vectors;

[0137] The feature fusion module is used to dynamically fuse the basic feature map and relevant context feature vectors to generate a task-oriented feature map;

[0138] The eye-tracking parameter regression module is used to perform eye-tracking parameter regression based on the task-oriented feature map to generate a lightweight eye tracking trajectory.

[0139] It should be noted that the electrical connections between the various units described above do not necessarily represent direct or indirect connections. Any indirect connection method can be applied to the embodiments of the present invention as long as it achieves the purpose of the present invention. The above descriptions are merely exemplary embodiments of the present invention and should not be construed as limiting the scope of the present invention.

[0140] All equivalent changes and modifications made in accordance with the teachings of this invention are still within the scope of this invention. Those skilled in the art will readily conceive of other embodiments of this invention upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this invention that follow the general principles of this invention and include common knowledge or conventional techniques in the art not described herein.

Claims

1. A lightweight eye-tracking method based on half-eye image reconstruction, characterized in that, The method includes: The image acquisition unit acquires half-eye images of the current frame in real time. Low-complexity feature extraction is performed on the half-eye image to generate a basic feature map of the current frame. This basic feature map represents the local state information of the eyeball. This includes: identifying the pupil region and iris portion in the half-eye image; determining whether the pupil region exhibits a bright pupil state by comparing its brightness with a preset brightness threshold; if yes, extracting eyeball feature points from the pupil region and its surroundings; otherwise, performing contrast enhancement and edge detection on the half-eye image, and extracting eyeball feature points from the pupil region and iris portion; and generating the basic feature map of the current frame based on these eyeball feature points. Based on the basic feature map, a preset dynamic context memory is queried to obtain relevant context feature vectors, wherein the memory stores eye-tracking context feature vectors of historical frames; including: comparing the similarity between the basic feature map and the eye-tracking context feature vectors of historical frames; and selecting the eye-tracking context feature vector of the historical frame with the highest similarity as the relevant context feature vector. The basic feature map and relevant context feature vectors are dynamically fused to generate a task-oriented feature map; Eye movement parameters are regressed based on the task-oriented feature map to generate a lightweight eye tracking trajectory; The process of identifying the pupil region and iris portion in the half-eye image includes: dynamically adjusting the output intensity of the supplementary light source integrated in the image acquisition unit based on the brightness of the pupil region and a preset brightness threshold, and adopting differentiated supplementary lighting strategies under bright pupil and non-bright pupil conditions; implementing high-resolution feature encoding for the pupil region, implementing robust feature encoding for the iris region to resist interference, and achieving cross-scale feature fusion through a feature pyramid structure.

2. The lightweight eye-tracking method based on half-eye image reconstruction according to claim 1, characterized in that, The method further includes: Based on the confidence level or the magnitude of change of the lightweight tracking trajectory of the eyeball, select effective tracking data for updating the dynamic context memory. Encode the effective tracking data into a new context feature vector; The new contextual feature vector is integrated into the dynamic contextual memory bank, and the eye-tracking contextual feature vectors of historical frames stored in the memory bank are adjusted.

3. The lightweight eye-tracking method based on half-eye image reconstruction according to claim 1, characterized in that, Dynamically fusing the base feature map and relevant contextual feature vectors includes: Based on the key local regions of the eye indicated by the eye feature points, a learnable attention mechanism is configured. The learnable attention mechanism is used to fuse the basic feature map and the relevant contextual feature vectors to generate a task-oriented feature map.

4. The lightweight eye-tracking method based on half-eye image reconstruction according to claim 1, characterized in that, The lightweight eye tracking trajectory includes: Based on the task-oriented feature map, eye movement parameters are regressed using a pre-trained eye movement parameter regression model to obtain the real-time three-dimensional gaze direction and gaze point of the eyeball; The motion characteristics of the three-dimensional gaze direction and gaze point are judged, and the output frequency and accuracy of the tracking data are dynamically adjusted to generate eye movement trajectory data that meets the requirements of lightweight design. The eye movement trajectory data is compressed and encoded, and output in a structured and compact format to generate a lightweight eye tracking trajectory.

5. A lightweight eye-tracking method based on half-eye image reconstruction according to claim 3, characterized in that, The learnable attention mechanism includes: The information of key local regions of the eyeball indicated by the eyeball feature points is input into a preset lightweight neural network model, and a dynamic attention weight map is output. Based on the dynamic attention weight map, the basic feature map and the relevant context feature vector are fused.

6. A lightweight eye-tracking method based on half-eye image reconstruction according to claim 3, characterized in that, The dynamic fusion of the basic feature map and the relevant context feature vector also includes: A three-dimensional attention mechanism is constructed that integrates spatial location weights and temporal decay factors, wherein the spatial location weights are determined by the topological distribution of eye feature points, and the key regions of the eye are precisely focused on during the fusion process; The task-oriented feature maps of multiple consecutive frames are input into the temporal prediction network to generate predicted values ​​of eye-tracking trajectories and to perform collaborative optimization with real-time detection results. A differential coding scheme based on motion features is adopted to selectively compress and transmit lightweight tracking trajectory data, thereby ensuring low-latency data output and supporting dynamic perception of user gaze points, focus areas, and attention regions.

7. A lightweight eye-tracking system based on half-eye image reconstruction, applied to a lightweight eye-tracking method based on half-eye image reconstruction as described in any one of claims 1-6, characterized in that, The system includes: The image acquisition module is used to acquire half-eye images of the eyeball in the current frame in real time; The feature extraction module is used to perform low-complexity feature extraction on the half-eye image to generate a basic feature map of the current frame. This includes: identifying the pupil region and iris portion in the half-eye image; determining whether the pupil region exhibits a bright pupil condition by comparing its brightness with a preset brightness threshold; if yes, extracting eyeball feature points from the pupil region and its surroundings; otherwise, performing contrast enhancement and edge detection on the half-eye image and extracting eyeball feature points from the pupil region and iris portion; and generating a basic feature map of the current frame based on the eyeball feature points. The context memory and query module is used to query a preset dynamic context memory based on the basic feature map to obtain relevant context feature vectors; including: comparing the similarity between the basic feature map and the eye-tracking context feature vectors of historical frames; and selecting the eye-tracking context feature vector of the historical frame with the highest similarity as the relevant context feature vector. The feature fusion module is used to dynamically fuse the basic feature map and relevant context feature vectors to generate a task-oriented feature map; The eye-tracking parameter regression module is used to perform eye-tracking parameter regression based on the task-oriented feature map to generate a lightweight eye tracking trajectory. The process of identifying the pupil region and iris portion in the half-eye image includes: dynamically adjusting the output intensity of the supplementary light source integrated in the image acquisition unit based on the brightness of the pupil region and a preset brightness threshold, and adopting differentiated supplementary lighting strategies under bright pupil and non-bright pupil conditions; implementing high-resolution feature encoding for the pupil region, implementing robust feature encoding for the iris region to resist interference, and achieving cross-scale feature fusion through a feature pyramid structure.

Citation Information

Patent Citations

  • Driver eye movement prediction method based on multi-scale attention

    CN119131766A