Self-adaptive eyeball tracking method and device based on deep learning

By combining the deep learning method of pulsed neural network and decision tree, high-precision eye tracking and reading behavior analysis under different light conditions is achieved, solving the problem of insufficient light adaptability and real-time in the existing technology, and improving the user experience.

CN120279591AActive Publication Date: 2025-07-08CHONGQING WANGYI INNOVATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510780486.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-08
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing eye tracking technology lacks accuracy and adaptability under different light conditions, which makes it difficult to meet real-time requirements, and has poor user experience.

Method used

Adaptive eye tracking method based on deep learning is adopted, combined with ViT and optimal decision tree on pulsed neural networks, and high-precision eye tracking and reading behavior analysis are achieved through adaptive tokens and ray condition adjustment.

Benefits of technology

Maintain high-precision eye tracking under various light conditions, adapt to different light environments, analyze reading behavior in real time, and provide support for personalized reading experience and content recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279591A_ABST
    Figure CN120279591A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a self-adaptive eyeball tracking method and device based on deep learning. The method comprises the following steps: acquiring eyeball images under various light conditions; preprocessing the eyeball image to obtain a preprocessed eyeball image; constructing an initial eyeball tracking model based on the spiking neural network, and training the initial eyeball tracking model to obtain a trained eyeball tracking model; inputting the preprocessed eyeball image into the trained eyeball tracking model, and obtaining a movement track of an eyeball in the eyeball image, wherein the movement track is output by the trained eyeball tracking model; and based on the motion track, analyzing eyeball data and reading behaviors of the target user to obtain an analysis result. The method not only can maintain high-precision eyeball tracking under various light conditions to adapt to different light environments, but also can analyze reading behaviors of readers in real time, and provides powerful support for personalized reading experience and content recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of eye tracking, and in particular, to an adaptive eye tracking method and device based on deep learning. Background Art

[0002] Eye tracking and analysis is a technology based on image processing and computer vision, used to real-time track and analyze the movement of the human eye in images or videos. Eye tracking technology has broad application prospects in fields such as reading analysis, human-computer interaction, and medical diagnosis.

[0003] In related technologies, a 3D dynamic model is displayed through a display screen to observe and induce pupil movements. However, the eye tracking technology in related technologies often has the following technical defects when facing different light conditions: errors are prone to occur when the light changes, reducing the accuracy and precision of tracking; calibration needs to be carried out in a fixed light environment, making it difficult to adapt to dynamically changing light conditions, with poor adaptability; some high-precision eye tracking systems often require a long calculation time when dealing with complex light conditions, unable to meet the real-time requirement; it is difficult to quickly adapt to the eye characteristics of different users, such as individual differences in eye shape, pupil size, etc., with low adaptability to individual differences; the eye tracking device may cause discomfort to users, resulting in eye fatigue after long-term use and affecting the long-term reading experience. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide an adaptive eye tracking method and device based on deep learning. By combining the ViT (Vision Transformer, a deep learning model based on the attention mechanism) adaptive token on a spiking neural network and the optimal decision tree method for approximate logic synthesis, precise capture and analysis of the eye movement of readers are achieved. This system can not only maintain high-precision eye tracking under various light conditions and adapt to different light environments, but also analyze the reading behavior of readers in real time, providing strong support for personalized reading experience and content recommendation.

[0005] In a first aspect, embodiments of the present disclosure provide an adaptive eye tracking method based on deep learning, adopting the following technical solutions: Obtain eye images under various light conditions; Preprocess the eye images to obtain preprocessed eye images; Construct an initial eye tracking model based on a spiking neural network and train the initial eye tracking model to obtain a trained eye tracking model; Input the preprocessed eye images into the trained eye tracking model and obtain the movement trajectory of the eyes in the eye images output by the trained eye tracking model; Analyze the eye data and reading behavior of the target user based on the movement trajectory to obtain an analysis result.

[0006] In some embodiments, obtain eye images under multiple light conditions, including: Set an image acquisition environment with different light source intensities or different light source angles; Collect the eye movements of a number of sampling users under various light conditions in the image acquisition environment through a photographing device; wherein, the number of sampling users have different ages and / or eye characteristics.

[0007] In some embodiments, preprocess the eye image to obtain a preprocessed eye image, including: Perform denoising processing on the eye image to obtain a denoised eye image; Perform enhancement processing on the denoised eye image to obtain an enhanced eye image; Detect the position and boundary of the eye region in the enhanced eye image, and perform segmentation processing on the eye region according to the detected position and boundary; Perform standardization processing and normalization processing on the segmented eye region to obtain a preprocessed eye image.

[0008] In some embodiments, construct an initial eye tracking model based on a spiking neural network, including; Define a spiking neuron model; Construct a Vision Transformer architecture based on the spiking neuron model; Initialize a number of learnable tokens, where each token is a D-dimensional vector; During each forward propagation, adaptively adjust these number of learnable tokens according to the light conditions of the input eye image; Connect the adjusted number of learnable tokens with the image patch embeddings of the Vision Transformer architecture; Input the connected image patch embeddings and the number of learnable tokens into the Transformer encoder of the Vision Transformer architecture to obtain the initial eye tracking model.

[0009] In some embodiments, train the initial eye tracking model to obtain a trained eye tracking model, including: Obtain a dataset of eye image samples of a number of sampling users, where the dataset of eye image samples contains eye image samples under multiple light conditions; Input the eye image samples under the multiple light conditions into the initial eye tracking model to obtain the motion trajectory samples of the eye samples output by the initial eye tracking model in the eye image samples; Calculate the loss between the motion trajectory samples and the target motion trajectory through the defined loss function to obtain a loss value; Based on the loss value, use an optimizer to optimize and adjust the relevant model parameters of the initial eye tracking model until the loss value is less than or equal to a preset threshold, and then output the trained eye tracking model; Among them, the loss function is expressed by the following formula: L total = α * L position + β * L gaze + γ * L adaptive + λ * L reg ; In the formula, L position represents the MSE loss of the feature point localization of the eye. The feature points of the eye include the eye corners and the pupil center; L gaze represents the angular error of the gaze direction prediction; L adaptive represents the consistency loss of the adaptive tokens; L reg represents the regularization term to prevent overfitting; α is the weight of L position ; β is the weight of L gaze ; γ is the weight of L adaptive ; λ is the weight of L reg .

[0010] In some embodiments, the method further includes: Extract several key eye features from the eye image samples. Among them, the key eye features include pupil center coordinates, pupil size, eye corner positions, eyelid positions, iris edge points, eye rotation angles, eye movement speeds, or pupil light reflection positions; Use the trained random forest model to calculate the importance of each of the key eye features; Based on the importance of the key eye features, select target eye features from the several key eye features; Construct a decision tree and select the target eye feature that can minimize the impurity or error to be the best splitting feature; Based on the best splitting feature, determine the best splitting threshold; Construct subtrees through the recursive method, and iteratively execute the steps of selecting the target eye feature that can minimize the impurity or error to be the best splitting feature and determining the best splitting threshold based on the best splitting feature through the subtrees until the iteration stop condition is reached; Optimize the decision tree through the post - pruning method to obtain the optimal decision tree; Perform ensemble learning on the optimal decision tree using a random forest model and a gradient - boosting tree model; Use k - fold cross - validation to verify and evaluate the initial eye - tracking model, and determine the trained eye - tracking model according to the verification and evaluation results.

[0011] In some embodiments, based on the motion trajectory, analyze the eye data and reading behavior of the target user to obtain analysis results, including: Based on the motion trajectory, calculate the eye fixation points and eye fixation duration of the target user; Identify the saccade behavior and regression behavior of the target user; Generate a heat map and a reading path based on the saccade behavior and regression behavior; Extract the reading behavior characteristics of the target user based on the heat map and reading path; Cluster - analyze to identify the reading pattern of the target user according to the reading behavior characteristics; Establish a personalized reading analysis system based on the reading pattern.

[0012] In a second aspect, an adaptive eye - tracking device based on deep learning provided by an embodiment of the present disclosure adopts the following technical solutions: An acquisition unit, configured to acquire eye images under various light conditions; A pre - processing unit, configured to pre - process the eye images to obtain pre - processed eye images; A model training unit, configured to construct an initial eye - tracking model based on a spiking neural network and train the initial eye - tracking model to obtain a trained eye - tracking model; An input - output unit, configured to input the pre - processed eye images into the trained eye - tracking model and obtain the motion trajectory of the eyes in the eye images output by the trained eye - tracking model; An analysis unit, configured to analyze the eye data and reading behavior of the target user based on the motion trajectory to obtain analysis results.

[0013] In a third aspect, an embodiment of the present disclosure also provides a computer device, adopting the following technical solutions: The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any of the above-described deep learning-based adaptive eye tracking methods.

[0014] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium storing computer instructions for causing a computer to execute any of the above-described deep learning-based adaptive eye tracking methods.

[0015] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of any of the above-described methods are implemented.

[0016] A deep learning-based adaptive eye tracking method provided by an embodiment of the present disclosure realizes precise capture and analysis of the eye movements of a reader by combining ViT (Vision Transformer, a deep learning model based on the attention mechanism) on a spiking neural network. The system can not only maintain high-precision eye tracking under various light conditions and adapt to different light environments, but also analyze the reading behavior of the reader in real time, providing strong support for personalized reading experiences and content recommendations.

[0017] The above description is only an overview of the technical solution of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present disclosure more obvious and understandable, the following preferred embodiments are specifically given and described in detail in conjunction with the drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart of a deep learning-based adaptive eye tracking method provided by an embodiment of the present disclosure; Figure 2 It is a schematic flowchart of optimizing the model training process of an initial eye tracking model provided by an embodiment of the present disclosure; Figure 3 It is a schematic flowchart of the implementation process of an optimal decision tree provided by an embodiment of the present disclosure; Figure 4Schematic flowchart for optimizing the performance of an eye tracking system under different lighting conditions using the optimal decision tree method provided by an embodiment of the present disclosure; Figure 5 Schematic structural diagram of an adaptive eye tracking device based on deep learning provided by an embodiment of the present disclosure; Figure 6 Schematic structural diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners

[0020] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0021] It should be clear that the following uses specific specific examples to illustrate the implementation manners of the present disclosure, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0022] It should also be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement a device and / or practice a method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.

[0023] It also needs to be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure in a schematic manner. The diagrams only show the components related to the present disclosure, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.

[0024] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0025] As Figure 1 shown, Figure 1 FIG. is a schematic flowchart of an adaptive eye tracking method based on deep learning provided by an embodiment of the present disclosure. An adaptive eye tracking method based on deep learning provided by an embodiment of the present disclosure includes the following steps: S101. Obtain eye images under various light conditions.

[0026] S102. Preprocess the eye images to obtain preprocessed eye images.

[0027] S103. Construct an initial eye tracking model based on a spiking neural network and train the initial eye tracking model to obtain a trained eye tracking model.

[0028] The embodiment of the present disclosure realizes precise capture and analysis of the eye movements of a reader (i.e., the target user in the embodiment of the present disclosure) by combining the adaptive tokens of the Vision Transformer deep learning model on the spiking neural network and the optimal decision tree method for approximate logic synthesis, and can maintain high-precision eye tracking under various light conditions.

[0029] S104. Input the preprocessed eye images into the trained eye tracking model and obtain the movement trajectory of the eyes in the eye images output by the trained eye tracking model.

[0030] S105. Analyze the eye data and reading behavior of the target user based on the movement trajectory to obtain an analysis result.

[0031] An adaptive eye tracking method based on deep learning provided by an embodiment of the present disclosure realizes precise capture and analysis of the eye movements of a reader by combining the ViT (Vision Transformer, a deep learning model based on the attention mechanism) on the spiking neural network for approximate logic synthesis. The system can not only maintain high-precision eye tracking under various light conditions and adapt to different light environments, but also analyze the reading behavior of the reader in real time, providing strong support for personalized reading experiences and content recommendations.

[0032] In some embodiments, obtaining eye images under various light conditions includes: Setting an image acquisition environment with different light source intensities or different light source angles; Collect the eye movements of a number of sampled users under various lighting conditions in the image acquisition environment; wherein, the number of sampled users have different ages and / or eye characteristics.

[0033] Optionally, the imaging device may be a high-resolution camera or other device, and the user may select a suitable imaging device according to actual needs. The embodiments of the present disclosure do not limit this.

[0034] In some embodiments, preprocessing the eye image to obtain a preprocessed eye image, including: Performing denoising processing on the eye image to obtain a denoised eye image; Performing enhancement processing on the denoised eye image to obtain an enhanced eye image; Detect the position and boundary of the eye region in the enhanced eye image, and perform segmentation processing on the eye region according to the detected position and boundary; Performing normalization processing and normalization processing on the segmented eye region to obtain a preprocessed eye image.

[0035] In some embodiments, constructing an initial eye tracking model based on a spiking neural network, including; Defining a spiking neuron model; Based on the spiking neuron model, constructing a Vision Transformer architecture; Initializing a number of learnable tokens, where each token is a D-dimensional vector; During each forward propagation, adaptively adjusting these learnable tokens according to the lighting conditions of the input eye image; Connecting the adjusted learnable tokens to the image patch embeddings of the Vision Transformer architecture; Inputting the connected image patch embeddings and the learnable tokens into the Transformer encoder of the Vision Transformer architecture to obtain the initial eye tracking model.

[0036] Optionally, the spiking neuron model is the basis of a bionic neural network, which simulates the behavior of biological neurons. In the embodiments of the present disclosure, the Leaky Integrate-and-Fire (LIF) model is adopted, and its mathematical expression is as follows: τ(dV / dt) = -(V - V rest ) + RI(t); where V represents the membrane potential, V restLet \(V\) represent the resting potential, \(\tau\) represent the time constant, \(R\) represent the membrane resistance, and \(I(t)\) represent the input current; when \(V\) reaches the threshold \(V\) th , the neuron fires a pulse and resets, that is, if \(V\geq V\) th , the fired pulse is reset to \(V = V\) reset .

[0037] Optionally, the core of the Vision Transformer architecture is to divide the eye image into fixed-size patches, and then input these divided patches into the Transformer encoder as a sequence. The main components of the Vision Transformer architecture include: 1. Patch embedding: used to divide the eye image into \(N\) patches, and each patch is converted into a \(D\)-dimensional vector through linear projection; 2. Position encoding: adding position information to each patch; 3. Transformer encoder: including multiple layers of self-attention mechanisms and feed-forward neural networks.

[0038] In the embodiments of the present disclosure, through the adaptive token mechanism, the initial eye tracking model dynamically adjusts its focus of attention to adapt to different lighting conditions In some embodiments, training the initial eye tracking model to obtain a trained eye tracking model includes: Obtaining an eye image sample data set of several sampled users, where the eye image sample data set contains eye image samples under various lighting conditions; Inputting the eye image samples under the various lighting conditions into the initial eye tracking model, and obtaining the movement trajectory samples of the eye samples output by the initial eye tracking model in the eye image samples; Calculating the loss between the movement trajectory samples and the target movement trajectory through a defined loss function to obtain a loss value; Based on the loss value, using an optimizer to optimize and adjust the relevant model parameters of the initial eye tracking model until the loss value is less than or equal to a preset threshold, and outputting the trained eye tracking model; Among them, the loss function is represented by the following formula: L total = α * L position + β * L gaze + γ * L adaptive + λ * L reg ; In the formula, \(L\) position represents the MSE loss of the feature point localization of the eye, and the feature points of the eye include the eye corners and the pupil center; \(L\) gaze represents the angular error of the gaze direction prediction; \(L\) adaptive represents the consistency loss of the adaptive token;reg Represents the regularization term to prevent overfitting; α is the weight of L position ; β is the weight of L gaze ; γ is the weight of L adaptive ; λ is the weight of L reg ; μ is the weight of L

[0039] Optionally, the eye image sample data set collected in the embodiments of the present disclosure is divided into a training set, a validation set, and a test set according to a ratio of 6:2:2 to ensure that each subset contains eye image sample data under various lighting conditions.

[0040] It should be noted that the user can divide the eye image sample data set according to the ratio of actual needs, and the embodiments of the present disclosure do not limit this.

[0041] Optionally, the optimizer adopts the Adam optimizer, the initial learning rate is set to 1e-4, and the cosine annealing scheduling strategy is used.

[0042] Optionally, the embodiments of the present disclosure monitor the model training process and implement the early stopping strategy. Specifically, tools such as TensorBoard (a set of visualization tools provided by TensorFlow) are used to monitor the model training process in real time, and metrics such as loss values and accuracies are recorded; if the performance of the validation set does not improve for 10 consecutive epochs (a general chart library for application developers and visualization designers), the model training is stopped to implement the early stopping strategy. Moreover, the model parameters with smaller loss values are saved on the validation set.

[0043] As Figure 2 shown, Figure 2 is a schematic flowchart of optimizing the model training process of the initial eye tracking model provided by the embodiments of the present disclosure. The embodiments of the present disclosure optimize the model training process of the initial eye tracking model, including the following steps: Step S1: Automatically search for the optimal hyperparameter combination using the Bayesian optimization method and tune the hyperparameters.

[0044] Among them, the main hyperparameters to be tuned may include: the number of layers and heads of the Transformer, the number of adaptive tokens, and the weight coefficients of each component in the loss function.

[0045] Step S2: Compress and quantize the spiking neuron model.

[0046] Specifically, pruning is performed on the spiking neuron model to remove unimportant connections and neurons. For example, L1 regularization is used to force the weights of some components in the loss function to be set to zero. A smaller student model is trained by knowledge distillation to mimic the large model, and the spiking neuron model is compressed. The floating-point weights in the spiking neuron model are converted to 8-bit or 16-bit integers to reduce the size and inference time of the spiking neuron model.

[0047] Step S3: Pre-train the initial eye tracking model on a large-scale eye image dataset, and fine-tune the pre-trained initial eye tracking model using a small amount of data from the target scenario, so that the pre-trained initial eye tracking model adapts to a specific application environment.

[0048] This embodiment of the present disclosure illustrates how to combine a spiking neural network and a Vision Transformer architecture to achieve high-precision eye tracking: Step A1: Prepare the input eye image.

[0049] Step A1.1: Preprocess the eye image.

[0050] Step A1.1.1: Resize the input eye image, and uniformly resize the size of the eye image to 224x224 pixels.

[0051] Step A1.1.2: Apply contrast-limited adaptive histogram equalization (CLAHE) to enhance the details of the eye image.

[0052] Step A1.1.3: Convert the eye image in RGB format to a grayscale image to reduce the computational amount.

[0053] Step A1.2: Perform spiking encoding on the eye image.

[0054] Step A1.2.1: Use rate encoding to convert the grayscale values of the grayscale image into a spike train.

[0055] For example, for an 8-bit grayscale image, the spike firing probability for a pixel value of x is set to: p(spike) = x / 255.

[0056] Step A1.2.2: Generate a spike train with T time steps in the time dimension.

[0057] Step A2: Process the eye image through a spiking neural network.

[0058] Step A2.1: Construct a spiking convolutional layer.

[0059] Step A2.1.1: Perform spiking convolution operation on the eye image using the IF neuron model.

[0060] Step A2.1.2: Design a three-layer pulse convolution structure, and use 32, 64, and 128 3x3 convolution kernels respectively.

[0061] Step A2.2: Perform temporal dimension aggregation processing on the eye image.

[0062] Step A2.2.1: Use max pooling operation in the temporal dimension to aggregate the feature maps of the eye images at T time steps into a single feature map.

[0063] Step A2.2.2: Use spatial max pooling to reduce the size of the feature map of the eye image to 14x14.

[0064] Step A3: Process the feature map of the eye image using the Vision Transformer architecture.

[0065] Step A3.1: Perform chunking and embedding processing on the feature map of the eye image.

[0066] Step A3.1.1: Divide the feature map of the eye image with a size of 14x14 into 2x2 chunks, obtaining a total of 49 chunks.

[0067] Step A3.1.2: Embed each chunk into a 768-dimensional vector space using linear projection.

[0068] Step A3.2: Fuse the tokens with the Vision Transformer architecture.

[0069] Step A3.2.1: Initialize a number of (10 are selected in the embodiments of the present disclosure) learnable tokens, and each token is a 768-dimensional vector.

[0070] Step A3.2.2: Use the attention mechanism to adjust these tokens according to the input eye image.

[0071] Step A3.2.3: Concatenate the adjusted tokens with the image chunk embeddings to form a sequence with a length of 59.

[0072] Step A3.3: Construct a Transformer encoder.

[0073] Step A3.3.1: Construct a 12-layer Transformer encoder, and each layer contains 8-head self-attention mechanism.

[0074] Step A3.3.2: Perform normalization processing and residual connection processing in the application layer behind each layer of the Transformer encoder.

[0075] Step A4: Use the initial eye-tracking model for data output processing and prediction.

[0076] Step A4.1: Extract features from the eye image.

[0077] Step A4.1.1: Use the output of the adaptive token as the global feature of the eye image.

[0078] Step A4.1.2: Map the global feature of the eye image to the required output dimension through a fully connected layer.

[0079] Step A4.2: Perform multi-task output on the global feature of the eye image.

[0080] Step A4.2.1: Use the MSE loss function to predict the coordinates of eye feature points (such as the corners of the eyes, the center of the pupil, etc.).

[0081] Step A4.2.2: Use the L1 loss function to predict the line-of-sight direction (horizontal and vertical angles).

[0082] Step A4.2.3: Use the relative error loss function to estimate the pupil size.

[0083] Step A5: Adapt the eye tracking to different lighting conditions Step A5.1: Encode the lighting conditions.

[0084] Step A5.1.1: Encode the lighting conditions using the global statistics of the eye image (such as average brightness, contrast, etc.).

[0085] Step A5.1.2: Inject the lighting condition information into the adaptive token.

[0086] Step A5.2: Data augmentation processing of the eye image.

[0087] Step A5.2.1: Randomly adjust the brightness, contrast, and noise level of the eye image during the training of the initial eye-tracking model.

[0088] Step A5.2.2: Simulate different light source directions and intensities to enhance the generalization ability of the initial eye-tracking model.

[0089] Step A5.3: Perform domain adaptation training on the initial eye-tracking model.

[0090] Step A5.3.1: Introduce a domain discriminator to distinguish eye features under different lighting conditions.

[0091] Step A5.3.2: Through adversarial training, make the eye features invariant under different lighting conditions.

[0092] Based on the above steps, embodiments of the present disclosure construct an eye tracking model that combines a spiking neural network and a Vision Transformer architecture, which can effectively process time series information, capture the dynamic characteristics of eye movements, and at the same time utilize the powerful representation ability of the Transformer encoder of the Vision Transformer architecture to adapt to light conditions in different environments. Moreover, through the adaptive token mechanism, the adaptability of the eye tracking model to environmental changes is further enhanced, enabling high-precision eye tracking performance to be maintained under various light conditions.

[0093] Embodiments of the present disclosure illustrate how to integrate various advanced deep learning technologies into an eye tracking system, laying a foundation for realizing robust and high-precision eye movement analysis. In practical applications, embodiments of the present disclosure can be further optimized and adjusted according to specific hardware limitations and precision requirements, and embodiments of the present disclosure do not limit this.

[0094] In some embodiments, the method further includes: Extracting several key eye features from the eye image samples, where the key eye features include pupil center coordinates, pupil size, eye corner positions, eyelid positions, iris edge points, eye rotation angles, eye movement speeds, or pupil light reflection positions; Using a trained random forest model to calculate the importance of each of the key eye features; Based on the importance of the key eye features, selecting target eye features from the several key eye features; Constructing a decision tree and selecting the target eye feature that can minimize impurity or error to be the best splitting feature; Based on the best splitting feature, determining the best splitting threshold; Constructing subtrees through the recursive method, and iteratively executing the steps of selecting the target eye feature that can minimize impurity or error to be the best splitting feature and, based on the best splitting feature, determining the best splitting threshold through the subtrees until the iteration stop condition is reached; Optimizing the decision tree through the post-pruning method to obtain an optimal decision tree; Performing ensemble learning on the optimal decision tree using a random forest model and a gradient boosting tree model; Using k-fold cross-validation to verify and evaluate the initial eye tracking model, and determining the trained eye tracking model according to the verification and evaluation results.

[0095] As Figure 3 shown, Figure 3Schematic flowchart of the implementation process of the optimal decision tree provided by the embodiments of the present disclosure. The implementation process of the optimal decision tree provided by the embodiments of the present disclosure includes the following steps: Step a1: Extract key eye features.

[0096] In the embodiments of the present disclosure, several key eye features are extracted and calculated from the output of the initial eye tracking model: pupil center coordinates (x, y), pupil size (diameter or area), corner positions of the eye (coordinates of the inner and outer corners), eyelid positions (contour points of the upper and lower eyelids), iris edge points (at least 8 evenly distributed points), eye rotation angles (horizontal and vertical directions), eye movement speed (based on position changes in consecutive frames), and pupil light reflex positions.

[0097] Step a2: Calculate the importance of key eye features.

[0098] The feature importance evaluation method of the random forest model is used to determine the importance of each key eye feature. The specific process is as follows: Step a2.1: Train a random forest model to predict the accurate line of sight direction using the above-extracted key eye features.

[0099] Step a2.2: Calculate the importance scores of each key eye feature through the Gini impurity or feature permutation method.

[0100] Step a2.3: Sort the key eye features according to the importance scores.

[0101] Step a3: Based on the feature importance of the key eye features, perform feature selection and dimensionality reduction.

[0102] Step a3.1: Select the key eye features with the top 80% importance scores according to the sorting of the importance scores.

[0103] Step a3.2: Apply principal component analysis (PCA) to process the selected key eye features and retain the principal components with an explained variance of 95%.

[0104] Step a3.2.3: If there is a high correlation between the key eye features, use the Lasso regression model for further feature selection.

[0105] Step a4: Select the best splitting feature and construct a decision tree using the CART (Classification and Regression Trees) algorithm.

[0106] Optionally, for each node, calculate the Gini impurity or mean squared error of all possible feature splits, and select the feature that can minimize the impurity or error to be the best splitting feature.

[0107] Step a5: Determine the optimal splitting threshold for the selected optimal splitting feature.

[0108] Optionally, sort the feature values of the optimal splitting feature, try all possible splitting points, and select the threshold that maximizes the information gain as the optimal splitting threshold. Use a dynamic programming algorithm to optimize the search process and improve the computational efficiency.

[0109] Step a6: Recursively repeat the above steps a4 (select the optimal splitting feature) and use the CART (Classification and Regression Trees) algorithm to construct a decision tree for the left and right subtrees (i.e., each child node); and the above step a5 (determine the optimal splitting threshold for the selected optimal splitting feature) until the iteration stop condition is reached.

[0110] Optionally, set the iteration stop condition to: stop the iteration when there are situations such as the maximum depth, minimum number of samples, and minimum reduction in impurity during the iteration process.

[0111] Step a7: Optimize the decision tree through post-pruning to obtain the optimal decision tree.

[0112] Generate a complete decision tree, evaluate each non-leaf node from bottom to top of the decision tree. If pruning the subtree can improve the performance of the validation set, then prune the subtree, and use the cost complexity parameter α to balance the size and performance of the decision tree.

[0113] Step a8: Use the random forest model and the gradient boosting tree model to perform ensemble learning on the optimal decision tree to further improve the performance of the optimal decision tree.

[0114] Optionally, using the random forest model to perform ensemble learning on the optimal decision tree includes the following steps: Generate multiple training sets using Bootstrap sampling; Construct a decision tree for each training set, and only consider a subset of features at each split; Use the average of all subtrees or the vote of the majority of subtrees for the final prediction of the output result.

[0115] Optionally, using the gradient boosting tree model (e.g., the optimized distributed gradient boosting library XGBoost) to perform ensemble learning on the optimal decision tree includes the following steps: Initialize a simple model; iteratively train new decision trees to fit the residuals; use the step size (learning rate) to control the contribution of each decision tree.

[0116] Step a9: Use k-fold cross-validation to validate and evaluate the initial eye tracking model, and select the best initial eye tracking model.

[0117] Optionally, divide the training set into k parts (for example, k can be set to 5 or 10); configure the parameters of the initial eye-tracking model (for example, configure parameters such as tree depth, minimum number of samples in leaf nodes, etc.), and perform k times of training and validation on the initial eye-tracking model; calculate the average performance metrics (for example, metrics such as MAE, RMSE, etc.) and standard deviation; select the model configuration with the best performance and the smallest variance to obtain the best initial eye-tracking model.

[0118] The embodiments of the present disclosure will elaborate in detail on how to use the optimal decision tree method to optimize the performance of the eye-tracking system under different light conditions to improve the eye-tracking accuracy.

[0119] As Figure 4 shown, Figure 4 is a schematic flowchart of using the optimal decision tree method to optimize the performance of the eye-tracking system under different light conditions provided by the embodiments of the present disclosure. The embodiments of the present disclosure for using the optimal decision tree method to optimize the performance of the eye-tracking system under different light conditions include the following steps: Step A1: Extract and preprocess key eye features.

[0120] Step A1.1: Extract key eye features from the output of the initial eye-tracking model.

[0121] Step A1.1.1: Extract the pupil center coordinates (x pupil , y pupil ); Step A1.1.2: Calculate the pupil diameter d pupil ; Step A1.1.3: Extract the inner corner coordinates (x inner , y inner ), and extract the outer corner coordinates (x outer , y outer ); Step A1.1.4: Calculate the eye width w eye = x outer - x inner ; Step A1.1.5: Calculate the relative pupil position r x = (x pupil - x inner ) / w eye , r y = (y pupil - y inner ) / w eye .

[0122] Step A1.2: Calculate light-related features.

[0123] Step A1.2.1: Calculate the average image brightness b avg ; Step A1.2.2: Calculate the image contrast contrast; Step A1.2.3: Detect and quantify the glare spots in the eye region intensity 。

[0124] Step A1.3: Standardize the light-related features.

[0125] Step A1.3.1: Perform Z-score standardization on all light-related features.

[0126] Step A1.3.2: According to the outliers in the results of Z-score standardization, truncate the outliers that exceed 3 standard deviations.

[0127] Step A2: Select the key eye features.

[0128] Step A2.1: Use the random forest model to evaluate the importance of each key eye feature; Step A2.1.1: Train the random forest model, with the target variable being the gaze direction (gaze x , gaze y ); Step A2.1.2: Calculate the importance score of each key eye feature; Step A2.1.3: Select the top 80% of the key eye features with the highest importance scores; Step A2.2: Apply the principal component analysis PCA to perform dimensionality reduction on the selected key eye features, and select the principal components that explain 95% of the variance.

[0129] Step A3: Build a decision tree.

[0130] Step A3.1: Initialize the decision tree; Step A3.1.1: Set the maximum depth of the decision tree to 10; Step A3.1.2: Set the minimum number of samples in the leaf nodes to 20; Step A3.1.3: Use the mean squared error as the splitting criterion; Step A3.2: Recursively build subtrees; Step A3.2.1: For each node, calculate the best split point for all key eye features; Step A3.2.2: Select the key eye feature and split point that can minimize the mean squared error to obtain the best classification feature; Step A3.2.3: Create left and right child nodes and recursively construct the sub-tree; Step A4: Optimize the decision tree.

[0131] Step A4.1: Optimize the decision tree through post-pruning to obtain the optimal decision tree.

[0132] Step A4.1.1: Evaluate each non-leaf node from bottom to top in the decision tree; Step A4.1.2: If pruning the sub-tree can reduce the validation set error, then prune the sub-tree; Step A4.1.3: Use the cost complexity parameter with α = 0.01 to control the pruning degree to balance the size and performance of the decision tree.

[0133] Step A4.2: Construct a random forest model.

[0134] Step A4.2.1: Create 100 decision trees; Step A4.2.2: Each decision tree adopts a dataset sampled by Bootstrap; Step A4.2.3: Randomly select a feature subset for each split, where the subset size is the square root of the total number of features; Step A4.3: Construct a gradient boosting tree (for example, the gradient boosting tree is XGBoost); Among them, set the maximum number of trees to 1000, the early stopping step to 50; set the learning rate to 0.01; set the maximum depth of each tree to 6.

[0135] Step A5: Validate, evaluate, and select the initial eye tracking model.

[0136] Step A5.1: Perform 5-fold cross-validation on the initial eye tracking model; Among them, perform cross-validation on the single decision tree, random forest model, and XGBoost respectively; calculate the average MAE and RMSE and their standard deviations of each initial eye tracking model.

[0137] Step A5.2: Test the initial eye tracking model under different lighting conditions; Step A5.2.1: Prepare test sets under different lighting conditions (for example, normal lighting, low light, strong light, side light, etc.); Step A5.2.2: Evaluate the performance of each initial eye tracking model under various lighting conditions.

[0138] Comprehensively considering the average performance, stability, and model performance under different lighting conditions, select the model with the best performance, the smallest variance, and the most stability as the final model to obtain the best initial eye tracking model.

[0139] Optionally, the embodiments of the present disclosure optimize the selected best eye tracking model. For example, for the finally selected eye tracking model, analyze the importance of eye features and identify the eye features that play a key role under different light conditions.

[0140] Optionally, the embodiments of the present disclosure analyze the decision paths of typical training samples in a decision tree or random forest model and identify key decision points and thresholds.

[0141] Optionally, based on the analysis results of the importance of eye features or the analysis results of the decision paths of typical training samples, the embodiments of the present disclosure fine-tune the selected best eye tracking model. For example, adjust the process of extracting key eye features, optimize model parameters (such as the depth of the decision tree, the number of samples in the smallest leaf node, etc.), and consider adding a dedicated model for specific light conditions, etc.

[0142] In summary, the embodiments of the present disclosure use the optimal decision tree method to optimize the performance of the eye tracking system under different light conditions. This method combines the powerful feature extraction ability of the deep learning model and the interpretability and adaptability of the decision tree model. Through key eye feature extraction, model integration, and targeted optimization, the embodiments of the present disclosure can achieve high-precision eye tracking under various light conditions, not only capture complex non-linear relationships, provide interpretability of the model decision-making process, but also quickly adapt to different light conditions, with high computational efficiency and being suitable for real-time applications.

[0143] In some embodiments, based on the motion trajectory, analyze the eye data and reading behavior of the target user to obtain analysis results, including: Based on the motion trajectory, calculate the eye fixation points and eye fixation durations of the target user; Identify the saccade behavior and regression behavior of the target user; Based on the saccade behavior and regression behavior, generate a heat map and a reading path; Based on the heat map and the reading path, extract the reading behavior characteristics of the target user; According to the reading behavior characteristics, perform cluster analysis to identify the reading pattern of the target user; Based on the reading pattern, establish a personalized reading analysis system.

[0144] Optionally, the reading behavior characteristics of the target user include reading speed, concentration, etc.

[0145] Optionally, an intelligent content recommendation can be provided for the target user through the established personalized reading analysis system, and the reading experience of the target user can be improved through reading assistance tools (such as automatically adjusting the font size, line spacing, etc.).

[0146] The embodiments of the present disclosure can not only maintain high-precision eye tracking in different light environments, but also analyze the reading behavior of the reader in real time, providing strong support for personalized reading experience and content recommendation.

[0147] Optionally, the hardware integration method of the embodiments of the present disclosure includes: selecting a suitable eye tracker and processor, designing a low-latency data transmission scheme, and implementing hardware acceleration (such as GPU or FPGA).

[0148] Optionally, the software development of the embodiments of the present disclosure includes: implementing a real-time image acquisition and preprocessing module, integrating a deep learning model inference engine, and developing a decision tree real-time prediction module.

[0149] Optionally, the model optimization of the embodiments of the present disclosure is achieved through parallel computing and multi-thread optimization, and also includes a caching mechanism, and also includes adaptive sampling rate adjustment and other methods for optimization.

[0150] The embodiments of the present disclosure elaborate in detail on how to extract eye features using the Vision Transformer adaptive tokens on the spiking neural network to adapt to different light conditions.

[0151] Step A1: Construct a spiking neural network.

[0152] Step A1.1: Define a spiking neuron model.

[0153] Step A1.1.1: Implement the Leaky Integrate-and-Fire (LIF) neuron model; Step A1.1.2: Set the threshold and reset mechanism of the LIF neuron model; Step A1.1.3: Define a spiking encoding function to convert the input eye image into a spiking sequence.

[0154] Step A1.2: Construct spiking neural network layers, implement spiking convolutional layers, implement spiking pooling layers, and implement spiking fully connected layers.

[0155] Step A2: Design a Vision Transformer architecture.

[0156] Step A2.1: Segment and embed the eye image. Specifically, divide the input eye image into blocks of a fixed size, perform linear projection on each image block to convert it into a D-dimensional vector, and perform embedding processing on the image block; add position encoding information to the image block.

[0157] Step A2.2: Design a Transformer encoder to implement the multi-head self-attention mechanism and the feed-forward neural network, and perform normalization processing and residual connection processing in the application layer behind each layer of the Transformer encoder.

[0158] Step A3: Implement the adaptive token mechanism.

[0159] Step A3.1: Initialize the adaptive token; Step A3.1.1: Randomly initialize a set of learnable tokens; Step A3.1.2: Design a token update strategy; Step A3.1.3: Implement the fusion mechanism between the token and the input embedding; Step A3.2: Implement the adaption of the light condition.

[0160] Step A3.2.1: Design a light condition encoder; Step A3.2.2: Inject the light condition information into the adaptive token; Step A3.2.3: Implement the dynamic adjustment of the adaptive token for different light conditions.

[0161] Step A4: Extract eye features. Specifically, design a feature extraction head to map the output of the Transformer encoder to the eye feature space, define key eye features (such as pupil center, eye corner position, etc.), and implement a feature point localization algorithm.

[0162] Step 5: Design a loss function. Specifically, define the eye feature point localization loss, add a regularization term to prevent overfitting; design the light adaptability loss to encourage the eye tracking model to remain stable under different light conditions.

[0163] Step 6: Optimize the training strategy. Specifically, implement progressive learning to gradually increase the complexity of light changes; adopt a contrastive learning strategy to improve the discriminative ability of key eye features; use knowledge distillation technology to extract key knowledge from large models.

[0164] The embodiments of the present disclosure construct an eyeball feature extraction model capable of adapting to different light conditions. This model utilizes the temporal dynamic characteristics of spiking neural networks and the powerful representation ability of the Vision Transformer architecture, combined with an adaptive token mechanism, to achieve precise capture and extraction of eyeball features, laying a foundation for how to apply advanced deep learning technologies to the field of eye tracking and for subsequent high-precision eye movement analysis.

[0165] It should be noted that in practical applications, the embodiments of the present disclosure can further optimize and adjust the model according to specific requirements, and the embodiments of the present disclosure do not limit this.

[0166] As Figure 5 shown, Figure 5 FIG. 10 is a schematic structural diagram of an adaptive eye tracking device based on deep learning provided by the embodiments of the present disclosure. The embodiments of the present disclosure also provide an adaptive eye tracking device based on deep learning, including: An acquisition unit 21, configured to acquire eyeball images under various light conditions; A preprocessing unit 22, configured to preprocess the eyeball images to obtain preprocessed eyeball images; A model training unit 23, configured to construct an initial eye tracking model based on a spiking neural network and train the initial eye tracking model to obtain a trained eye tracking model; An input / output unit 24, configured to input the preprocessed eyeball images into the trained eye tracking model and obtain the movement trajectory of the eyeball in the eyeball images output by the trained eye tracking model; An analysis unit 25, configured to analyze the eyeball data and reading behavior of the target user based on the movement trajectory to obtain an analysis result.

[0167] The computer device according to the embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0168] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the computer device executes all or part of the steps of a deep learning-based adaptive eye tracking method according to the foregoing embodiments of the present disclosure.

[0169] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included in the protection scope of the present disclosure.

[0170] Such as Figure 6 It is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. It shows a schematic structural diagram of a computer device suitable for implementing the computer device in the embodiment of the present disclosure. Figure 6 The shown computer device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0171] Such as Figure 6 As shown, the computer device may include a processor (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0172] Generally, the following devices may be connected to the I / O interface: an input device including, for example, a sensor or a visual information acquisition device; an output device including, for example, a display screen; a storage device including, for example, a magnetic tape, a hard disk, etc.; and a communication device. The communication device may allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or wiredly to exchange data. Although Figure 6 a computer device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.

[0173] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of an adaptive eye tracking method based on deep learning according to an embodiment of the present disclosure are performed.

[0174] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated here.

[0175] A computer-readable storage medium according to an embodiment of the present disclosure stores non-temporary computer-readable instructions. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of an adaptive eye tracking method based on deep learning according to the foregoing embodiments of the present disclosure are performed.

[0176] The above computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROMs and DVDs), magneto-optical storage media (e.g., MOs), magnetic storage media (e.g., magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (e.g., memory cards), and media with built-in ROMs (e.g., ROM cartridges).

[0177] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated here.

[0178] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above specific details are only for illustrative and facilitating understanding purposes, rather than limitations, and the above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0179] In this disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.

[0180] In addition, as used herein, "or" in a listing of items beginning with "at least one" indicates a disjunctive listing, so that for example a listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Further, the term "exemplary" does not mean that the examples described are preferred or better than other examples.

[0181] It should also be noted that in the systems and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0182] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings defined by the appended claims. In addition, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Accordingly, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.

[0183] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0184] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those of skill in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. An adaptive eye tracking method based on deep learning, characterized in that Including: Obtain eye images under various lighting conditions; Preprocess the eye images to obtain preprocessed eye images; Construct an initial eye tracking model based on a spiking neural network and train the initial eye tracking model to obtain a trained eye tracking model; Input the preprocessed eye images into the trained eye tracking model and obtain the movement trajectory of the eyes in the eye images output by the trained eye tracking model; Analyze the eye data and reading behavior of the target user based on the movement trajectory to obtain an analysis result.

2. The adaptive eye tracking method based on deep learning according to claim 1, wherein Obtain eye images under various lighting conditions, including: Set an image acquisition environment with different light source intensities or different light source angles; Collect the eye movements of a number of sampling users under various lighting conditions in the image acquisition environment through a shooting device; wherein, the number of sampling users have different ages and / or eye characteristics.

3. The adaptive eye tracking method based on deep learning according to claim 1, characterized in that Preprocess the eye images to obtain preprocessed eye images, including: Denoise the eye images to obtain denoised eye images; Enhance the denoised eye images to obtain enhanced eye images; Detect the position and boundary of the eye region in the enhanced eye images and segment the eye region according to the detected position and boundary; Perform standardization processing and normalization processing on the segmented eye region to obtain preprocessed eye images.

4. The adaptive eye tracking method based on deep learning according to claim 1, wherein Construct an initial eye tracking model based on a spiking neural network, including; Define a spiking neuron model; Construct a Vision Transformer architecture based on the spiking neuron model; Initialize a number of learnable tokens, where each token is a D-dimensional vector; During each forward propagation, adaptively adjust these learnable tokens according to the lighting conditions of the input eye images; Connect the adjusted learnable tokens with the image patch embeddings of the Vision Transformer architecture; Input the connected image patch embeddings and the learnable tokens into the Transformer encoder of the Vision Transformer architecture to obtain the initial eye tracking model.

5. The adaptive eye tracking method based on deep learning according to claim 1 or 4, characterized in that Train the initial eye tracking model to obtain a trained eye tracking model, including: Obtain a dataset of eye image samples of a number of sampling users, where the dataset of eye image samples contains eye image samples under various lighting conditions; Input the eye image samples under various lighting conditions into the initial eye tracking model and obtain the movement trajectory samples of the eye samples output by the initial eye tracking model in the eye image samples; Calculate the loss between the movement trajectory samples and the target movement trajectory through a defined loss function to obtain a loss value; Based on the loss value, the optimizer optimizes and adjusts the relevant model parameters of the initial eye tracking model until the loss value is less than or equal to a preset threshold, and then outputs the trained eye tracking model; Among them, the loss function is represented by the following formula: L total = α * L position + β * L gaze + γ * L adaptive + λ * L reg ; Where, L position represents the MSE loss for the feature point localization of the eyeball, and the feature points of the eyeball include the eye corners and the pupil center; L gaze represents the angular error of the gaze direction prediction; L adaptive represents the consistency loss of the adaptive token; L reg represents the regularization term to prevent overfitting; α is the weight of L position ; β is the weight of L gaze ; γ is the weight of L adaptive ; λ is the weight of L reg .

6. The adaptive eye tracking method based on deep learning according to claim 5, wherein The method further includes: Extracting several key eye features from the eye image samples, where the key eye features include pupil center coordinates, pupil size, corner of the eye position, eyelid position, iris edge points, eye rotation angle, eye movement speed, or pupil light reflection position; Calculating the importance of each of the key eye features using the trained random forest model; Based on the importance of the key eye features, selecting target eye features from the several key eye features; Constructing a decision tree and selecting the target eye feature that can minimize the impurity or error to the greatest extent as the best splitting feature; Based on the best splitting feature, determining the best splitting threshold; Constructing subtrees through the recursive method, and iteratively executing the steps of selecting the target eye feature that can minimize the impurity or error to the greatest extent as the best splitting feature and determining the best splitting threshold based on the best splitting feature through the subtrees until the iteration stop condition is reached; Optimizing the decision tree through the post - pruning method to obtain the optimal decision tree; Performing ensemble learning on the optimal decision tree using the random forest model and the gradient boosting tree model; Using k - fold cross - validation to verify and evaluate the initial eye tracking model, and determining the trained eye tracking model according to the verification and evaluation results.

7. The adaptive eye tracking method based on deep learning according to claim 1, wherein Based on the motion trajectory, analyzing the eye data and reading behavior of the target user to obtain analysis results, including: Based on the motion trajectory, calculating the eye fixation points and eye fixation duration of the target user; Identifying the saccade behavior and regression behavior of the target user; Generating a heat map and a reading path based on the saccade behavior and regression behavior; Extracting the reading behavior features of the target user based on the heat map and the reading path; Clustering and analyzing to identify the reading pattern of the target user according to the reading behavior features; Based on the reading pattern, establishing a personalized reading analysis system.

8. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; where, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the adaptive eye tracking method based on deep learning according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, This computer - readable storage medium stores computer instructions for causing a computer to execute the adaptive eye tracking method based on deep learning according to any one of claims 1 to 7.

10. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Eye-ball position trac method, device, terminal and computer-readable storage medium

    CN109359512A

  • High-speed clear pupil eye movement detection and tracking method and system based on event camera

    CN116030527A

  • Eye movement tracking method and device, display equipment, storage medium and program product

    CN118965095A

  • Eyeball tracking method and device and electronic equipment

    CN119810898A

  • Eyeball detection method, apparatus and device, and storage medium

    WO2021197466A1

Cited By

  • Hardware system for deploying sparse pulse Transform model

    CN121031688A