Human eye attention localization method and device based on pre-trained neural network

Through the human eye attention positioning method based on pre-trained neural network, using eye tracking equipment and attention mechanism models, the problems of insufficient accuracy and slow response speed of human eye attention positioning in the prior art are solved, and high-precision and fast gaze area positioning in complex scenarios are achieved.

CN119323824BActive Publication Date: 2025-06-06UNIVERSAL UBIQUITOUS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411876135.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-06-06
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

The prior art has problems in the positioning of human eye attention, slow response speed and poor adaptability to complex scenes, especially when dealing with dynamic changes in human eye movement and interference from complex backgrounds, it is difficult to effectively locate the area of ​​human eye gaze.

Method used

The human eye attention positioning method based on pretrained neural network is adopted to obtain human eye movement data in real time through eye movement tracking devices, and a pretrained neural network model containing attention mechanism is used to extract feature information related to the human eye gaze area. The model includes a convolutional neural network and attention mechanism module, which is used to weight processing of high-dimensional features and calculate the spatial coordinates of the human eye's gaze area through the regression model.

Benefits of technology

It significantly improves the gaze tracking accuracy and response speed in complex scenarios, and can more accurately locate the human eye gaze area and adapt to various complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119323824B_ABST
    Figure CN119323824B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a method and device for locating human eye attention based on a pre-trained neural network, the method comprising: obtaining human eye movement data in real time through an eye tracking device, extracting feature information related to the human eye gaze area through a pre-trained neural network model including an attention mechanism based on the human eye movement data, inputting the extracted feature information into a regression model, and calculating the spatial coordinates of the human eye gaze area according to the feature information through the regression model; the present application can achieve accurate positioning of the human eye gaze area through the pre-trained neural network model, and significantly improve the gaze tracking accuracy and response speed in complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of visual positioning, and in particular to a method and device for positioning human eye attention based on a pre-trained neural network. Background Art

[0002] Traditional methods of locating human eye attention are based on dual-camera, 3D structured light, TOF and other methods. However, they suffer from problems such as insufficient accuracy, slow response speed and poor adaptability to complex scenes, which limits their application in fields such as virtual reality, intelligent monitoring and human-computer interaction.

[0003] Current methods mainly rely on simple geometric models and fixed feature matching, which cannot effectively handle the dynamic changes of human eye movement and the interference of complex background.

[0004] In addition, due to the lack of an effective feature extraction mechanism, existing technologies cannot take into account the comprehensive analysis of spatial and temporal information when predicting the area where the human eye is looking, making it difficult to improve positioning accuracy. With the rapid development of deep learning technology, feature extraction and attention mechanisms based on neural networks provide new solutions for eye attention positioning.

[0005] However, when existing neural network models are applied to locate human eye attention, they usually have limited training data, and the model's generalization ability in different application scenarios is insufficient, making it difficult to ensure accurate positioning in various complex scenarios. Summary of the invention

[0006] In response to the problems in the prior art, the present application provides a method and device for locating human eye attention based on a pre-trained neural network, which can accurately locate the human eye gaze area through a pre-trained neural network model, and significantly improve the gaze tracking accuracy and response speed in complex scenes.

[0007] In order to solve at least one of the above problems, the present application provides the following technical solutions:

[0008] According to a first aspect of an embodiment of the present application, the present application provides a method for locating human eye attention based on a pre-trained neural network, comprising:

[0009] Acquiring human eye movement data in real time through an eye tracking device, wherein the human eye movement data includes the position of the human eye pupil and the direction of the line of sight, wherein the eye tracking device includes at least one camera and a computing unit for image processing, wherein the camera is used to capture the movement of the human eye, and the computing unit is used to analyze and extract characteristic parameters of the human eye movement;

[0010] Based on the human eye movement data, feature information related to the human eye gaze area is extracted through a pre-trained neural network model including an attention mechanism, wherein the neural network model includes a convolutional neural network and an attention mechanism module, wherein the convolutional neural network is used to extract high-dimensional features related to human eye gaze, and the attention mechanism module is used to perform weighted processing on the high-dimensional features;

[0011] The extracted feature information is input into a regression model, and the spatial coordinates of the human eye gaze area are calculated according to the feature information by the regression model. The regression model includes a multi-layer perceptron, which generates precise coordinate values ​​for determining the position of the gaze area by combining the extracted high-dimensional features and context information.

[0012] According to any embodiment of the present application, the training process of the neural network model includes:

[0013] Collecting a plurality of annotated human eye movement data, wherein the annotated human eye movement data includes the real position coordinates of the gaze target;

[0014] Build a neural network model that includes a convolutional neural network, an attention mechanism module, and a recurrent neural network;

[0015] Offline training is performed on the neural network model to minimize the error between the predicted position and the actual position of the gaze area by inputting the human eye movement data;

[0016] The trained neural network model is verified and tested to improve the positioning accuracy of the neural network model in different application scenarios to the preset standard.

[0017] According to any embodiment of the present application, the step of obtaining the pupil position and sight direction of a human eye in real time by using an eye tracking device includes:

[0018] The pupil position and sight direction of a human eye are obtained in real time by an eye tracking device, wherein the eye tracking device includes a plurality of cameras at different shooting positions;

[0019] Fusion of multiple camera view data to improve the robustness of data acquisition to a preset standard.

[0020] According to any embodiment of the present application, the extracting feature information related to the human eye gaze area based on the human eye movement data using a pre-trained neural network model including an attention mechanism includes:

[0021] The features extracted by the convolutional neural network are weighted based on the channel attention mechanism to enhance the attention to specific feature channels to the preset standard;

[0022] The feature map is weighted through the spatial attention mechanism to improve the positioning accuracy of the human eye's gaze area to the preset standard.

[0023] According to any embodiment of the present application, the step of inputting the extracted feature information into a regression model, and calculating the spatial coordinates of the eye gaze area according to the feature information through the regression model, includes:

[0024] Inputting the extracted feature information into a regression model so that the regression model combines the high-dimensional features and the time series information, updates the coordinate prediction through the output of the recursive neural network, and thus determines the spatial coordinates of the gaze area;

[0025] Through the nonlinear mapping capability of the multi-layer perceptron, the high-dimensional features extracted by the convolutional neural network and the temporal features of the recurrent neural network are integrated to output precise coordinates of the gaze area.

[0026] According to any embodiment of the present application, after calculating the spatial coordinates of the area where the human eye is looking, the method further includes:

[0027] Based on the movement trajectory and timing changes of the human eye, a recursive neural network is used to analyze the collected human eye movement data in the time dimension to obtain the stability results of the visual target.

[0028] According to a second aspect of the embodiments of the present application, the present application provides a human eye attention positioning device based on a pre-trained neural network, comprising:

[0029] A data acquisition module, used to: obtain human eye movement data in real time through an eye tracking device, wherein the human eye movement data includes the position of the human eye pupil and the direction of the line of sight, wherein the eye tracking device includes at least one camera and a computing unit for image processing, wherein the camera is used to capture the movement of the human eye, and the computing unit is used to analyze and extract characteristic parameters of the human eye movement;

[0030] A feature extraction module, used to: extract feature information related to the human eye gaze area based on the human eye movement data through a pre-trained neural network model including an attention mechanism, wherein the neural network model includes a convolutional neural network and an attention mechanism module, wherein the convolutional neural network is used to extract high-dimensional features related to human eye gaze, and the attention mechanism module is used to perform weighted processing on the high-dimensional features;

[0031] The coordinate calculation module is used to: input the extracted feature information into a regression model, and calculate the spatial coordinates of the human eye gaze area based on the feature information through the regression model, wherein the regression model includes a multi-layer perceptron, and the multi-layer perceptron generates precise coordinate values ​​for determining the position of the gaze area by combining the extracted high-dimensional features and context information.

[0032] According to any embodiment of the present application, the training process of the neural network model includes:

[0033] The annotation acquisition module is used to: collect a plurality of annotated human eye movement data, wherein the annotated human eye movement data includes the real position coordinates of the gaze target;

[0034] Model building module, used to: build a neural network model including convolutional neural network, attention mechanism module and recurrent neural network;

[0035] A model training module, used to: perform offline training on the neural network model, by inputting the human eye movement data to minimize the error between the predicted position and the actual position of the gaze area;

[0036] The model verification module is used to verify and test the trained neural network model to improve the positioning accuracy of the neural network model in different application scenarios to a preset standard.

[0037] According to any embodiment of the present application, the step of obtaining the pupil position and sight direction of a human eye in real time by using an eye tracking device includes:

[0038] A data acquisition module is used to: acquire the pupil position and sight direction of a human eye in real time through an eye tracking device, wherein the eye tracking device includes multiple cameras to capture the movement characteristics of the human eye at different angles;

[0039] The data fusion module is used to fuse the viewing angle data of multiple cameras to improve the robustness of data acquisition to a preset standard.

[0040] According to any embodiment of the present application, the extracting feature information related to the human eye gaze area based on the human eye movement data using a pre-trained neural network model including an attention mechanism includes:

[0041] The channel attention module is used to: perform weighted processing on the features extracted by the convolutional neural network based on the channel attention mechanism to enhance the attention to specific feature channels to a preset standard;

[0042] The spatial attention module is used to: perform weighted processing on the feature map through the spatial attention mechanism to improve the positioning accuracy of the human eye's gaze area to a preset standard.

[0043] According to any embodiment of the present application, the step of inputting the extracted feature information into a regression model, and calculating the spatial coordinates of the eye gaze area according to the feature information through the regression model, includes:

[0044] A feature input module, used to: input the extracted feature information into a regression model, so that the regression model combines high-dimensional features and time series information, updates the coordinate prediction through the output of the recursive neural network, and thus determines the spatial coordinates of the gaze area;

[0045] The nonlinear mapping module is used to: integrate the high-dimensional features extracted by the convolutional neural network and the temporal features of the recurrent neural network through the nonlinear mapping capability of the multi-layer perceptron to output the precise coordinates of the gaze area.

[0046] According to any embodiment of the present application, a motion analysis module is further included, which is used to:

[0047] Based on the movement trajectory and timing changes of the human eye, a recursive neural network is used to analyze the collected human eye movement data in the time dimension to obtain the stability results of the visual target.

[0048] According to the third aspect of the embodiments of the present application, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for locating human eye attention based on a pre-trained neural network are implemented.

[0049] According to a fourth aspect of an embodiment of the present application, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for locating human eye attention based on a pre-trained neural network.

[0050] According to the fifth aspect of the embodiments of the present application, the present application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the method for locating human eye attention based on a pre-trained neural network.

[0051] It can be seen from the above technical scheme that the present application provides a method and device for locating human eye attention based on a pre-trained neural network, which obtains human eye movement data in real time through an eye tracking device, and extracts feature information related to the human eye gaze area based on the human eye movement data through a pre-trained neural network model including an attention mechanism, and inputs the extracted feature information into a regression model, and calculates the spatial coordinates of the human eye gaze area according to the feature information through the regression model. The pre-trained neural network model can be used to accurately locate the human eye gaze area, significantly improving the gaze tracking accuracy and response speed in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0053] Figure 1 This is one of the flow charts of the method for locating human eye attention based on a pre-trained neural network in an embodiment of the present application;

[0054] Figure 2 This is a second flow chart of a method for locating human eye attention based on a pre-trained neural network in an embodiment of the present application;

[0055] Figure 3 This is a flowchart of the method for locating human eye attention based on a pre-trained neural network in an embodiment of the present application;

[0056] Figure 4 This is a fourth flow chart of a method for locating human eye attention based on a pre-trained neural network in an embodiment of the present application;

[0057] Figure 5 This is a fifth flow chart of the method for locating human eye attention based on a pre-trained neural network in an embodiment of the present application;

[0058] Figure 6 This is a sixth flow chart of the method for locating human eye attention based on a pre-trained neural network in an embodiment of the present application;

[0059] Figure 7 This is a structural diagram of a human eye attention positioning device based on a pre-trained neural network in an embodiment of the present application;

[0060] Figure 8 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0062] The acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.

[0063] The present application provides a method and device for locating human eye attention based on a pre-trained neural network, which achieves accurate positioning of the human eye's gaze area through a pre-trained neural network model, significantly improving the gaze tracking accuracy and response speed in complex scenes.

[0064] In order to accurately locate the human eye gaze area through a pre-trained neural network model and significantly improve the gaze tracking accuracy and response speed in complex scenes, the present application provides an embodiment of a human eye attention location method based on a pre-trained neural network, see Figure 1 The human eye attention localization method based on the pre-trained neural network specifically includes the following contents:

[0065] Step S101: acquiring human eye movement data in real time through an eye tracking device, wherein the human eye movement data includes a human eye pupil position and a line of sight direction, wherein the eye tracking device includes at least one camera and a computing unit for image processing, wherein the camera is used to capture the movement of the human eye, and the computing unit is used to analyze and extract characteristic parameters of the human eye movement;

[0066] Optionally, in this embodiment, the eye tracking device is constructed using a high-performance infrared camera and a dedicated image processing unit. The infrared camera has a sampling frequency of 240Hz, a resolution of 1280x720 pixels, and is equipped with an adaptive exposure control system, which can stably capture pupil images under different lighting conditions. The camera is fixed under the monitor through an adjustable bracket to maintain the best observation angle with the user's eyes while avoiding interference with the user's normal vision.

[0067] The image processing unit adopts an embedded system design, integrating a dedicated image processing chip and a neural network accelerator. The system first pre-processes the acquired image, including denoising, histogram equalization, and regional enhancement operations to improve the clarity of the pupil edge. Then, the improved Hough transform algorithm is used to detect the pupil contour in real time, and the Kalman filter is combined to track the dynamic changes of the pupil center position, effectively suppressing the detection jitter.

[0068] Based on the pupil position, the system calculates the sight direction through the perspective geometry model. First, based on the pre-calibrated camera parameters, the mapping relationship between the image coordinate system and the world coordinate system is established. Then, combined with the physiological characteristic parameters of the eyeball, the center position of the eyeball is estimated through the 3D reconstruction algorithm. Finally, based on the relative position of the pupil center and the eyeball center, the 3D direction of the sight vector is calculated.

[0069] To improve the reliability of data collection, the system implements a multiple verification mechanism. By analyzing the eccentricity of the pupil shape, abnormal data caused by factors such as blinking are filtered out. At the same time, sudden changes in pupil area are monitored to identify and compensate for measurement errors caused by head movement. The system also maintains a real-time updated reference model to detect and correct measurement deviations caused by changes in ambient light.

[0070] The extraction of characteristic parameters adopts a multi-scale analysis method. The system not only records the instantaneous values ​​of pupil position and line of sight direction, but also calculates the time derivatives of these parameters to characterize the dynamic characteristics of eye movement. The frequency characteristics of eye movement are analyzed by wavelet transform to identify different types of eye movement patterns such as fixation, saccade and eye saccade. At the same time, physiological characteristic parameters such as pupil size change and blinking frequency are extracted to provide multi-dimensional input data for subsequent attention analysis.

[0071] The data transmission adopts a high-speed cache structure to ensure real-time performance while providing data preprocessing capabilities. The system caches the collected raw data and extracted feature parameters separately, supporting data backtracking analysis within 500ms. The double buffer mechanism realizes seamless data collection and processing, ensuring a stable sampling frequency under high load conditions.

[0072] This embodiment solves the technical problems of unstable data collection and incomplete feature extraction in traditional eye tracking. Through multiple optimization strategies, the system achieves sub-pixel precision positioning of the pupil position under natural lighting conditions, and the line of sight direction measurement error is controlled within 0.5 degrees. At the same time, through a rich feature extraction mechanism, a reliable data basis is provided for subsequent attention analysis. In practical applications, the system shows excellent environmental adaptability and user-friendliness, and can meet the needs of various human-computer interaction scenarios.

[0073] Step S102: based on the eye movement data, extracting feature information related to the eye gaze area through a pre-trained neural network model including an attention mechanism, wherein the neural network model includes a convolutional neural network and an attention mechanism module, wherein the convolutional neural network is used to extract high-dimensional features related to the eye gaze, and the attention mechanism module is used to perform weighted processing on the high-dimensional features;

[0074] Optionally, this embodiment uses a deep learning framework to build a multi-level feature extraction system. The backbone network of the pre-trained neural network model uses an improved ResNet-50 architecture and is pre-trained through transfer learning to fully utilize the visual feature knowledge in large-scale data sets. The network structure is optimized, the fully connected layer of the original model is removed, the convolutional layer is retained for feature extraction, and the attention mechanism module is added at key positions.

[0075] The convolutional neural network adopts a multi-scale feature extraction strategy. The network contains 5 convolution blocks, each of which uses convolution kernels of different sizes for feature extraction. The first layer uses a 7×7 convolution kernel to capture large-scale features, and the subsequent layers gradually reduce the convolution kernel size to 3×3 to achieve fine-grained feature extraction. To enhance the expressiveness of features, a dilated convolution is introduced in each convolution block, and the receptive field is expanded by adjusting the dilation rate while maintaining computational efficiency.

[0076] The attention mechanism module adopts a dual attention structure, including two sub-modules: channel attention and spatial attention. Channel attention obtains the statistical features of the channel dimension through global average pooling and maximum pooling, and generates channel weights through a shared multi-layer perceptron. Spatial attention generates an attention map of the spatial dimension through convolution operations to highlight the key areas in the visual features. The outputs of the two attention mechanisms are combined through an adaptive fusion strategy to generate the final feature weights.

[0077] In order to improve the environmental adaptability of the model, a feature enhancement module is implemented. This module first normalizes the input features to reduce the impact of data distribution offset. Then, the original feature information is retained through residual connections, and a feature pyramid structure is introduced to extract and fuse features at multiple scales. At the same time, a feature denoising mechanism is implemented to filter out irrelevant background interference through the self-attention mechanism.

[0078] The system also includes a temporal feature extraction module for analyzing the dynamic patterns of human eye movements. This module uses a bidirectional LSTM network to process continuous eye movement data and capture long-term dependencies. The gating mechanism automatically adjusts the importance of features at different time scales to effectively identify eye movement patterns such as fixations and saccades.

[0079] Feature fusion uses an adaptive weight mechanism. The system dynamically adjusts the fusion weight according to the reliability of different features to ensure stable feature representation in various scenarios. At the same time, a feature selection mechanism is implemented to automatically select the most discriminative feature subsets through sparse constraints.

[0080] This embodiment solves the problem of insufficient expression ability and poor environmental adaptability of traditional feature extraction methods. Through multi-level feature extraction and attention enhancement, the system can accurately capture the key features of human eye gaze behavior. In practical applications, this solution demonstrates excellent feature extraction capabilities, and the discriminability and robustness of features are significantly improved, providing reliable feature support for subsequent gaze area positioning.

[0081] The test results of the model in different scenarios show that the accuracy of feature extraction reaches more than 95%, and it has good anti-interference ability to interference factors such as lighting changes and head movement. At the same time, through the introduction of the attention mechanism, the system can automatically focus on the most relevant feature areas, significantly improving the expression efficiency and computing performance of the features.

[0082] Step S103: input the extracted feature information into a regression model, and calculate the spatial coordinates of the gaze area of ​​the human eye according to the feature information through the regression model, wherein the regression model includes a multi-layer perceptron, and the multi-layer perceptron generates precise coordinate values ​​for determining the position of the gaze area by combining the extracted high-dimensional features and context information.

[0083] Optionally, this embodiment uses a multi-level regression architecture to convert feature information into accurate spatial coordinates of the fixation area. The core of the regression model is an optimized multi-layer perceptron, which contains four hidden layers, each of which uses 256, 128, 64, and 32 neurons. The LeakyReLU activation function is used between the layers to effectively avoid the gradient vanishing problem. At the same time, a Dropout layer is added after each layer, and the ratio is set to 0.3 to enhance the generalization ability of the model.

[0084] To improve the accuracy of coordinate prediction, the system implements a feature fusion mechanism. First, the high-dimensional features extracted by the convolutional neural network are reduced in dimension through adaptive pooling to generate feature vectors of fixed dimension. These features are then concat-operated with the temporal features extracted from the eye movement data to form fused features. At the same time, attention-weighted contextual information is introduced, including historical gaze point distribution and scene content features, and the weights of different information sources are dynamically adjusted through a gating mechanism.

[0085] The regression process adopts a multi-task learning framework. The main task is to predict the absolute coordinates of the gaze point, and the auxiliary tasks include predicting the relative displacement of the gaze point and the duration of the gaze. Through shared feature representation and task-specific output layers, the system can learn more robust feature representations. At the same time, an adaptive loss weighting mechanism is implemented to dynamically adjust the weight of the loss function according to the learning difficulty of different tasks.

[0086] To improve the temporal continuity of predictions, the system integrates a time series smoothing module. This module uses a Kalman filter to post-process the original prediction results, taking into account the physical constraints of human eye movement and filtering out unreasonable jumps. At the same time, a prediction confidence estimation mechanism is implemented to compensate or interpolate low-confidence prediction results.

[0087] The coordinate mapping adopts a nonlinear transformation strategy. The system first predicts in the standardized coordinate space, and then transforms the prediction result to the actual display coordinate system through reverse mapping. The mapping process takes into account the geometric characteristics of the display and the viewing distance, and achieves accurate coordinate transformation through piecewise polynomial functions.

[0088] This embodiment solves the problem of insufficient accuracy and poor stability of traditional regression methods when processing high-dimensional features. Through a multi-level regression architecture and feature fusion mechanism, the system can accurately predict the spatial coordinates of the gaze area. In actual application tests, the average error of coordinate prediction is controlled within 1% of the display resolution, and the time delay is less than 20 milliseconds.

[0089] The system also implements an adaptive calibration mechanism, which compensates for system deviations caused by user fatigue or environmental changes through periodic online fine-tuning. At the same time, the stability of the prediction results is significantly improved, and the jitter amplitude between consecutive predictions is reduced by more than 80%, providing a smooth user experience for gaze-based human-computer interaction.

[0090] In actual application scenarios, the regression model shows excellent environmental adaptability and user universality. Whether in office environments or mobile scenarios, the system can maintain stable prediction performance. In particular, it shows obvious advantages when dealing with rapid eye movements and gaze point jumps, providing reliable technical support for various gaze control applications.

[0091] From the above description, it can be seen that the human eye attention positioning method based on the pre-trained neural network provided in the embodiment of the present application can accurately locate the human eye gaze area through the pre-trained neural network model, significantly improving the gaze tracking accuracy and response speed in complex scenes.

[0092] In one embodiment of the human eye attention localization method based on the pre-trained neural network of the present application, see Figure 2 , and can also include the following:

[0093] Step S201: collecting a plurality of annotated human eye movement data, wherein the annotated human eye movement data includes the real position coordinates of the gaze target;

[0094] Step S202: construct a neural network model including a convolutional neural network, an attention mechanism module and a recursive neural network; perform offline training on the neural network model by inputting the human eye movement data to minimize the error between the predicted position and the actual position of the gaze area; verify and test the trained neural network model to improve the positioning accuracy of the neural network model in different application scenarios to a preset standard.

[0095] Optionally, in the data collection stage, this embodiment designs a systematic data collection scheme. First, 100 volunteers of different ages and genders were recruited, covering the age range of 18-60 years old, to ensure data diversity. The collection environment includes two conditions: standard office lighting (500 lux) and natural lighting. The volunteers maintain a standard viewing distance of 60-70 cm from the display. The display uses a 27-inch 4K resolution screen to ensure sufficient spatial resolution.

[0096] The experimental design includes a variety of visual stimulation tasks. The first is a fixed-point gaze task, in which visual targets are randomly presented at different positions on the screen for 2-3 seconds. Then there is a dynamic tracking task, in which the target moves on the screen at different speeds and trajectories. Finally, there is a free browsing task, including natural scenes such as text reading, picture browsing, and video watching. Each task lasts for 15 minutes, and the eye movement data and the actual position of the target are recorded synchronously by a professional eye tracker.

[0097] The neural network model adopts a modular design. The convolutional neural network uses the improved EfficientNet-B3 as the backbone network, and initializes the model parameters through transfer learning. The attention mechanism module adopts a multi-head self-attention structure, with the number of heads set to 8 and the dimension of each attention head being 64, which effectively captures the long-distance dependency between features. The recursive neural network uses a bidirectional GRU structure with a hidden layer dimension of 256 to model temporal dependency.

[0098] The model training adopts a phased strategy. In the first phase, the pre-trained weights are fixed and only the newly added layers are trained; in the second phase, all layers are unfrozen for end-to-end fine-tuning. The training process uses the Adam optimizer, the initial learning rate is set to 0.0001, and the cosine annealing strategy is used for dynamic adjustment. The loss function comprehensively considers the mean square error and smooth L1 loss, and introduces a regularization term to suppress overfitting. The batch size is set to 64, and the training lasts for 100 rounds until convergence.

[0099] Verification and testing are divided into three stages. First, the model performance is evaluated on the validation set, and the optimal hyperparameters are determined through cross-validation. Then, the performance is evaluated on an independent test set, which contains different environmental conditions and usage scenarios. Finally, actual scene testing is carried out, including testing under different lighting conditions, different usage distances, and different user groups.

[0100] In order to improve the generalization ability of the model, a data enhancement mechanism is implemented. The training data is expanded by adding Gaussian noise, random occlusion, and perspective transformation. At the same time, online hard sample mining is implemented to give priority to learning samples that are difficult to predict. The system also includes a model integration mechanism to fuse the prediction results of multiple models through voting or weighted averaging.

[0101] This embodiment solves the problems of insufficient generalization ability and poor environmental adaptability of traditional eye tracking systems. Through systematic data collection and deep learning model training, stable and reliable gaze point positioning in different scenarios is achieved. Test results show that the positioning accuracy of the model under standard conditions is better than 0.5 degrees of viewing angle, and it can still maintain an accuracy of less than 1 degree in complex environments, meeting most practical application needs.

[0102] In particular, through multi-stage training and verification strategies, the model shows excellent environmental adaptability. Under different lighting conditions, the fluctuation of positioning accuracy is controlled within 15%; for different user groups, stable tracking effects can be achieved without personalized calibration. These features enable the system to be widely used in various human-computer interaction scenarios, providing reliable technical support for improving the interactive experience.

[0103] In one embodiment of the human eye attention localization method based on the pre-trained neural network of the present application, see Figure 3 , and can also include the following:

[0104] Step S301: acquiring the pupil position and sight direction of a human eye in real time through an eye tracking device, wherein the eye tracking device includes a plurality of cameras at different shooting positions;

[0105] Step S302: Fusing the viewing angle data of multiple cameras to improve the robustness of data acquisition to a preset standard.

[0106] Optionally, this embodiment adopts a multi-camera collaborative acquisition strategy and is equipped with three industrial-grade infrared cameras, which are located at the lower left, lower right and directly below the display. Each camera has a sampling frequency of 120Hz, a resolution of 1920x1080 pixels, and is equipped with an 850nm narrow-band infrared filter and an autofocus system. The camera is fixed by a high-precision pan-tilt bracket, which has a six-degree-of-freedom fine-tuning mechanism that can accurately adjust the spatial position and observation angle of the camera.

[0107] Before data collection begins, the system first performs geometric calibration of multiple cameras. Using the improved Zhang calibration method, high-precision calibration plates are used to collect calibration images at different positions. Through corner point detection and nonlinear optimization, the intrinsic parameter matrix and distortion coefficient of each camera are obtained. Then, the extrinsic parameters between cameras are calibrated to establish a unified world coordinate system to ensure the spatial consistency of data from different perspectives.

[0108] During the real-time acquisition process, the system implements a multi-channel data synchronization mechanism based on timestamps. Each camera is equipped with an independent image processing unit, which realizes parallel data processing through high-speed cache. Image preprocessing includes dynamic threshold segmentation, morphological operations and regional enhancement to improve the detection accuracy of pupil edges.

[0109] Multi-view data fusion adopts a hierarchical strategy. First, pupil detection is performed on each view image separately, and the pupil contour is extracted using an improved ellipse fitting algorithm. Then, the two-dimensional features of different view angles are projected into three-dimensional space through a three-dimensional reconstruction algorithm to obtain the spatial coordinates of the pupil center. The system implements an adaptive weighting mechanism based on confidence, dynamically adjusting the fusion weight according to the detection quality of each view angle.

[0110] To improve the robustness of the system, multiple fault detection and recovery mechanisms are implemented. Data reliability is evaluated in real time by analyzing the image quality indicators of each view, including clarity, contrast, and signal-to-noise ratio. When the data quality of a certain view decreases, the system automatically adjusts the fusion strategy to increase the weight of other views. At the same time, a frame loss compensation mechanism is implemented to predict missing data through a Kalman filter.

[0111] The line of sight direction is calculated using a multi-model fusion method. First, a geometric model is established based on the physiological characteristics of the eyeball to calculate the relative position of the pupil center and the eyeball center. Then, the line of sight vector is estimated by combining the observation data from multiple perspectives through least squares optimization. The system also implements a learning-based line of sight estimation model that directly predicts the line of sight direction from multi-perspective images through a neural network.

[0112] This embodiment solves the problem that the single-camera acquisition system is easily blocked and has unstable accuracy. Through multi-camera collaborative acquisition and data fusion, the system achieves more stable and accurate eye tracking. The experimental results show that under standard test conditions, the measurement error of the pupil position is less than 0.1mm, and the measurement error of the line of sight direction is controlled within 0.3 degrees.

[0113] In particular, the system exhibits excellent environmental adaptability. In the presence of partial occlusion, as long as at least two cameras can obtain effective observation, the system can still maintain stable tracking effects. When processing rapid eye movements, the complementary effect of multi-view data significantly reduces the impact of motion blur. This solution provides reliable technical support for various application scenarios that require accurate eye tracking.

[0114] In one embodiment of the human eye attention localization method based on the pre-trained neural network of the present application, see Figure 4 , and can also include the following:

[0115] Step S401: weighting the features extracted by the convolutional neural network based on the channel attention mechanism to enhance the attention to specific feature channels to a preset standard;

[0116] Step S402: weighting the feature map through the spatial attention mechanism to improve the positioning accuracy of the human eye gaze area to a preset standard.

[0117] Optionally, this embodiment designs a dual attention mechanism to enhance the expressiveness of the model through feature enhancement in the channel dimension and the spatial dimension. In the channel attention processing, the convolution feature map is first globally compressed, and two channel descriptors are generated by global average pooling and maximum pooling respectively. These descriptors are nonlinearly transformed through a shared multi-layer perceptron. The perceptron contains two fully connected layers, the number of neurons in the middle layer is 1 / 16 of the number of channels, and the ReLU activation function is used.

[0118] The weight generation of channel attention adopts an adaptive fusion strategy. The system performs element-wise weighted summation on the results of average pooling and maximum pooling, and the weight coefficient is dynamically adjusted through learnable parameters. In order to enhance the discrimination ability of weights, a temperature parameter is introduced to adjust the smoothness of the softmax function. At the same time, a channel group attention mechanism is implemented to group related feature channels and improve the focusing effect of attention.

[0119] The spatial attention module adopts a multi-scale feature aggregation strategy. First, average pooling and maximum pooling operations are performed in the spatial dimension to generate two two-dimensional feature maps. These feature maps are extracted through a 3×3 convolution layer, and then a spatial weight map is generated through a sigmoid function. To enhance the expression of local features, a dilated convolution is introduced before the convolution operation to capture multi-scale contextual information through different dilation rates.

[0120] The integration of the two attention mechanisms adopts a cascade structure. First, channel attention processing is performed to enhance the feature response of key channels; then spatial attention is applied to the enhanced feature map to highlight important spatial regions. To avoid information loss, the system adds a residual connection after each attention module to retain the original feature information. At the same time, a feature recalibration mechanism is implemented to maintain the stability of feature distribution through a batch normalization layer.

[0121] In order to improve the environmental adaptability of the attention mechanism, a dynamic threshold adjustment mechanism is implemented. The system adaptively adjusts the threshold of the attention weight according to the statistical characteristics of the input features to maintain a stable feature enhancement effect in different scenarios. At the same time, an attention regularization term is introduced to prevent the attention weight from being overly concentrated in local areas.

[0122] This embodiment solves the problem of scattered focus and insufficient feature expression in traditional feature extraction methods. Through the synergy of the dual attention mechanism, the system can accurately locate and enhance key features related to human eye gaze. Experimental results show that on the standard test set, the discriminability of features is improved by 40%, and the positioning accuracy of the model is improved by 25%.

[0123] In particular, by introducing the attention mechanism, the system exhibits excellent noise suppression capabilities. In complex backgrounds, the selective enhancement of channels and spatial dimensions effectively suppresses the influence of background interference, and the stability of gaze point positioning is significantly improved. This solution has shown obvious advantages in practical applications, especially when dealing with rapid eye movements and complex scenes.

[0124] The computational efficiency of the system has also been optimized. Through lightweight attention module design and efficient feature fusion strategy, the performance is improved while only increasing the computational overhead by about 5%. This enables the system to meet the requirements of real-time processing and provides a practical solution for various eye tracking applications.

[0125] In one embodiment of the human eye attention localization method based on the pre-trained neural network of the present application, see Figure 5 , and can also include the following:

[0126] Step S501: inputting the extracted feature information into a regression model, so that the regression model combines high-dimensional features and time series information, updates the coordinate prediction through the output of the recursive neural network, and thus determines the spatial coordinates of the gaze area;

[0127] Step S502: Through the nonlinear mapping capability of the multi-layer perceptron, the high-dimensional features extracted by the convolutional neural network and the temporal features of the recurrent neural network are integrated to output accurate gaze area coordinates.

[0128] Optionally, this embodiment designs a multi-level feature fusion regression architecture to achieve deep integration of spatial features and temporal features. The system first constructs a feature fusion channel to compress the high-dimensional features extracted by the convolutional neural network to a fixed dimension through an adaptive pooling layer. These features pass through the feature selection module and use a gating mechanism to dynamically select important features. At the same time, a feature standardization layer is introduced to ensure the scale consistency of features from different sources.

[0129] The recursive neural network adopts a bidirectional LSTM structure with a hidden layer dimension of 512. In order to improve the time series modeling capability, multi-scale time series feature extraction is implemented. The system maintains three time windows of different lengths to capture short-term (250ms), medium-term (500ms) and long-term (1000ms) time series dependencies. The features of each time window are weighted and fused through the attention mechanism, realizing the adaptive selection of information at different time scales.

[0130] The feature integration stage adopts a progressive fusion strategy. First, the feature alignment module is used to ensure the consistency of spatial features and temporal features at the semantic level. Then, the cross-attention mechanism is used to establish the correlation mapping between the two features. The system implements a feature complementation mechanism to enhance the weight of temporal features when spatial features are insufficient, and vice versa. At the same time, residual connections are introduced to ensure the complete transmission of information.

[0131] The multi-layer perceptron adopts a deep separation architecture, which includes a feature extraction layer and a coordinate mapping layer. The feature extraction layer uses three hidden layers with dimensions of 1024, 512, and 256, respectively, and uses the ELU activation function to improve nonlinear expression capabilities. The coordinate mapping layer passes through two 32-dimensional hidden layers and finally outputs a two-dimensional coordinate value. To improve generalization capabilities, batch normalization and dropout layers are added after each hidden layer.

[0132] The system implements an adaptive learning rate adjustment mechanism. By analyzing the changing trend of the prediction error, the learning rate of different feature channels is dynamically adjusted. At the same time, gradient clipping is introduced to prevent the gradient explosion problem during training. In order to improve the stability of the prediction, a prediction smoothing mechanism is implemented to filter abnormal prediction values ​​through exponential sliding average.

[0133] This embodiment solves the problem of the integration difficulty of traditional regression methods when processing high-dimensional features and time series information. Through multi-level feature fusion and nonlinear mapping, the system can accurately predict the spatial coordinates of the gaze area. Experimental results show that under standard test conditions, the average error of coordinate prediction is reduced to less than 0.5 degrees, and the time delay is controlled within 10 milliseconds.

[0134] In particular, the system exhibits excellent temporal consistency. When processing continuous eye movement sequences, the temporal smoothness of the prediction results is significantly improved, and the coordinate jumps between adjacent frames are reduced by 85%. At the same time, through the feature complementation mechanism, the system maintains stable prediction performance when processing rapid eye movements and gaze point jumps.

[0135] In practical applications, the solution has shown wide adaptability. Whether in scenarios such as text reading, image browsing or video watching, the system can provide accurate gaze point positioning. Especially when dealing with complex dynamic scenes, the prediction accuracy and stability are significantly improved through the effective use of temporal features. These features provide reliable technical support for the development of more natural and intuitive human-computer interaction interfaces.

[0136] In one embodiment of the human eye attention localization method based on the pre-trained neural network of the present application, see Figure 6 , and can also include the following:

[0137] Step S601: integrating the movement trajectory and time sequence changes of the human eye;

[0138] Step S602: Use a recursive neural network to perform time dimension analysis on the collected human eye movement data to obtain a stability result of the visual target.

[0139] Optionally, this embodiment designs a complete eye movement trajectory analysis system, and realizes accurate modeling of human eye movement characteristics through deep learning methods. The system first performs time series preprocessing on the original eye movement data, and uses a sliding window method to divide the continuous eye movement data into fixed-length sequence segments. The window length is set to 500ms and the step length is 50ms to ensure that the complete eye movement characteristics are captured.

[0140] In the trajectory analysis stage, the system implements a multi-feature extraction mechanism. First, the basic kinematic features such as eye velocity and acceleration are calculated, and the basic eye movement types such as fixation, saccade and blink are identified through the adaptive threshold method. At the same time, trajectory curvature analysis is introduced to describe the smoothness of the eye movement trajectory by calculating the local curvature change. The system also implements a trajectory segmentation mechanism, which decomposes complex trajectories into basic motion units through a dynamic programming algorithm.

[0141] The time series analysis uses a hierarchical LSTM network structure. The first layer of LSTM is responsible for feature extraction, with a hidden dimension of 256, and captures forward and backward time series dependencies through bidirectional processing. The second layer of LSTM focuses on state modeling, with a hidden dimension of 128, and selectively retains important state information through a gating mechanism. The system implements attention-enhanced sequence modeling and captures long-distance time series associations through a self-attention mechanism.

[0142] To improve the accuracy of time series analysis, the system introduces a multi-scale feature fusion mechanism. Through parallel temporal convolutional networks, features are extracted at different time scales, including microscale (50ms), mesoscale (200ms) and macroscale (1000ms). These features are fused through an adaptive weight mechanism to achieve a comprehensive analysis of patterns at different time scales.

[0143] The stability assessment adopts a multi-index fusion method. First, the spatial distribution characteristics of the gaze point are calculated, including statistical indicators such as divergence and aggregation. Then the fluctuation characteristics of the time dimension are analyzed, and the periodic pattern is identified through spectrum analysis. The system implements an entropy-based stability quantification method to evaluate the concentration of visual attention by calculating the spatiotemporal entropy of the trajectory.

[0144] This embodiment solves the limitations of traditional eye movement analysis methods in dealing with complex time series patterns. Through deep learning and multi-dimensional analysis, the system can accurately identify and evaluate the stability of visual attention. Experimental results show that in standard test scenarios, the system's recognition accuracy of gaze state reaches 95%, and the consistency of stability assessment reaches 90%.

[0145] In particular, the system exhibits excellent noise robustness. Through time series filtering and anomaly detection, high-frequency noise and transient disturbances in eye movement data are effectively suppressed. When processing long time series, the system is able to maintain stable performance, with drift error controlled within 0.1 degrees.

[0146] In practical applications, the solution provides reliable analysis support for a variety of scenarios. In reading comprehension tests, the system can accurately identify reading difficulties and changes in cognitive load. In human-computer interaction interfaces, real-time stability assessment provides a basis for adaptive interface adjustments. In the field of medical diagnosis, the system can assist in identifying neurological abnormalities related to eye movements.

[0147] The analysis results of the system are well interpretable. Through the visualization analysis tool, the spatiotemporal characteristics and stability changes of the eye movement trajectory can be intuitively displayed. These characteristics provide a powerful tool for in-depth understanding of the human visual cognitive process, and lay the foundation for the development of more intelligent human-computer interaction systems.

[0148] In order to accurately locate the human eye gaze area through a pre-trained neural network model and significantly improve the gaze tracking accuracy and response speed in complex scenes, the present application provides an embodiment of a human eye attention positioning device based on a pre-trained neural network for implementing all or part of the human eye attention positioning method based on a pre-trained neural network, see Figure 7 The human eye attention positioning device based on the pre-trained neural network specifically includes the following contents:

[0149] The data acquisition module 10 is used to obtain human eye movement data in real time through an eye tracking device, wherein the human eye movement data includes the position of the human eye pupil and the direction of the line of sight, wherein the eye tracking device includes at least one camera and a computing unit for image processing, wherein the camera is used to capture the movement of the human eye, and the computing unit is used to analyze and extract characteristic parameters of the human eye movement;

[0150] A feature extraction module 20 is used to: extract feature information related to the human eye gaze area based on the human eye movement data through a pre-trained neural network model including an attention mechanism, wherein the neural network model includes a convolutional neural network and an attention mechanism module, wherein the convolutional neural network is used to extract high-dimensional features related to human eye gaze, and the attention mechanism module is used to perform weighted processing on the high-dimensional features;

[0151] The coordinate calculation module 30 is used to: input the extracted feature information into a regression model, and calculate the spatial coordinates of the human eye gaze area based on the feature information through the regression model, wherein the regression model includes a multi-layer perceptron, and the multi-layer perceptron generates precise coordinate values ​​for determining the position of the gaze area by combining the extracted high-dimensional features and context information.

[0152] From the above description, it can be seen that the human eye attention positioning device based on a pre-trained neural network provided in the embodiment of the present application can accurately position the human eye gaze area through a pre-trained neural network model, significantly improving the gaze tracking accuracy and response speed in complex scenes.

[0153] From the hardware level, in order to accurately locate the human eye gaze area through the pre-trained neural network model and significantly improve the gaze tracking accuracy and response speed in complex scenes, the present application provides an embodiment of an electronic device for implementing all or part of the human eye attention location method based on the pre-trained neural network, and the electronic device specifically includes the following contents:

[0154] Processor, memory, communication interface and bus; wherein the processor, memory and communication interface communicate with each other through the bus; the communication interface is used to realize information transmission between the human eye attention positioning device based on the pre-trained neural network and the core business system, user terminal and related database and other related devices; the logic controller can be a desktop computer, a tablet computer and a mobile terminal, etc., but the present embodiment is not limited thereto. In the present embodiment, the logic controller can be implemented with reference to the embodiment of the human eye attention positioning method based on the pre-trained neural network and the embodiment of the human eye attention positioning device based on the pre-trained neural network in the embodiment, and the contents thereof are incorporated herein, and the repeated parts are not repeated.

[0155] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0156] In practical applications, part of the human eye attention positioning method based on the pre-trained neural network can be executed on the electronic device side as described above, or all operations can be completed in the client device. The specific selection can be based on the processing capability of the client device and the limitations of the user's usage scenario. This application does not limit this. If all operations are completed in the client device, the client device may also include a processor.

[0157] The client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0158] Figure 8 FIG. 9 is a schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Figure 8 As shown, the electronic device 9600 may include a central processor 9100 and a memory 9140; the memory 9140 is coupled to the central processor 9100. It is worth noting that Figure 8 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0159] In one embodiment, the human eye attention localization method function based on the pre-trained neural network can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control:

[0160] Step S101: acquiring human eye movement data in real time through an eye tracking device, wherein the human eye movement data includes a human eye pupil position and a line of sight direction, wherein the eye tracking device includes at least one camera and a computing unit for image processing, wherein the camera is used to capture the movement of the human eye, and the computing unit is used to analyze and extract characteristic parameters of the human eye movement;

[0161] Step S102: based on the eye movement data, extracting feature information related to the eye gaze area through a pre-trained neural network model including an attention mechanism, wherein the neural network model includes a convolutional neural network and an attention mechanism module, wherein the convolutional neural network is used to extract high-dimensional features related to the eye gaze, and the attention mechanism module is used to perform weighted processing on the high-dimensional features;

[0162] Step S103: input the extracted feature information into a regression model, and calculate the spatial coordinates of the gaze area of ​​the human eye according to the feature information through the regression model, wherein the regression model includes a multi-layer perceptron, and the multi-layer perceptron generates precise coordinate values ​​for determining the position of the gaze area by combining the extracted high-dimensional features and context information.

[0163] From the above description, it can be seen that the electronic device provided in the embodiment of the present application can achieve accurate positioning of the human eye gaze area through a pre-trained neural network model, thereby significantly improving the gaze tracking accuracy and response speed in complex scenes.

[0164] In another embodiment, the human eye attention positioning device based on a pre-trained neural network can be configured separately from the central processing unit 9100. For example, the human eye attention positioning device based on a pre-trained neural network can be configured as a chip connected to the central processing unit 9100, and the function of the human eye attention positioning method based on the pre-trained neural network can be realized through the control of the central processing unit.

[0165] like Figure 8 As shown, the electronic device 9600 may also include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Figure 8 In addition, the electronic device 9600 may also include Figure 8 For components not shown, reference may be made to the prior art.

[0166] like Figure 8 As shown, the central processing unit 9100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.

[0167] The memory 9140 may be, for example, one or more of a cache, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory or other suitable devices. The above-mentioned information related to the failure may be stored, and a program for executing the relevant information may also be stored. The CPU 9100 may execute the program stored in the memory 9140 to implement information storage or processing, etc.

[0168] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.

[0169] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be a memory that saves information even when the power is off, can be selectively erased, and is provided with more data, examples of which are sometimes referred to as EPROMs, etc. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142, which is used to store application programs and function programs or processes for executing the operation of the electronic device 9600 through the central processor 9100.

[0170] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0171] The communication module 9110 is a transmitter / receiver 9110 that sends and receives signals via an antenna 9111. The communication module (transmitter / receiver) 9110 is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.

[0172] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module and / or a wireless LAN module, etc. The communication module (transmitter / receiver) 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby realizing a common telecommunication function. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 9130 is also coupled to the central processor 9100, so that recording can be performed on the local machine through the microphone 9132, and the sound stored on the local machine can be played through the speaker 9131.

[0173] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all the steps of the method for locating human eye attention based on a pre-trained neural network in the above-mentioned embodiments, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, all the steps of the method for locating human eye attention based on a pre-trained neural network in the above-mentioned embodiments are implemented. For example, when the processor executes the computer program, the following steps are implemented:

[0174] Step S101: acquiring human eye movement data in real time through an eye tracking device, wherein the human eye movement data includes a human eye pupil position and a line of sight direction, wherein the eye tracking device includes at least one camera and a computing unit for image processing, wherein the camera is used to capture the movement of the human eye, and the computing unit is used to analyze and extract characteristic parameters of the human eye movement;

[0175] Step S102: based on the eye movement data, extracting feature information related to the eye gaze area through a pre-trained neural network model including an attention mechanism, wherein the neural network model includes a convolutional neural network and an attention mechanism module, wherein the convolutional neural network is used to extract high-dimensional features related to the eye gaze, and the attention mechanism module is used to perform weighted processing on the high-dimensional features;

[0176] Step S103: input the extracted feature information into a regression model, and calculate the spatial coordinates of the gaze area of ​​the human eye according to the feature information through the regression model, wherein the regression model includes a multi-layer perceptron, and the multi-layer perceptron generates precise coordinate values ​​for determining the position of the gaze area by combining the extracted high-dimensional features and context information.

[0177] From the above description, it can be seen that the computer-readable storage medium provided in the embodiment of the present application can achieve accurate positioning of the human eye gaze area through a pre-trained neural network model, thereby significantly improving the gaze tracking accuracy and response speed in complex scenes.

[0178] The embodiments of the present application also provide a computer program product capable of implementing all the steps of the human eye attention localization method based on a pre-trained neural network in the above-mentioned embodiments, where the execution subject is a server or a client. When the computer program / instruction is executed by a processor, the steps of the human eye attention localization method based on a pre-trained neural network are implemented. For example, the computer program / instruction implements the following steps:

[0179] Step S101: acquiring human eye movement data in real time through an eye tracking device, wherein the human eye movement data includes a human eye pupil position and a line of sight direction, wherein the eye tracking device includes at least one camera and a computing unit for image processing, wherein the camera is used to capture the movement of the human eye, and the computing unit is used to analyze and extract characteristic parameters of the human eye movement;

[0180] Step S102: based on the eye movement data, extracting feature information related to the eye gaze area through a pre-trained neural network model including an attention mechanism, wherein the neural network model includes a convolutional neural network and an attention mechanism module, wherein the convolutional neural network is used to extract high-dimensional features related to the eye gaze, and the attention mechanism module is used to perform weighted processing on the high-dimensional features;

[0181] Step S103: input the extracted feature information into a regression model, and calculate the spatial coordinates of the gaze area of ​​the human eye according to the feature information through the regression model, wherein the regression model includes a multi-layer perceptron, and the multi-layer perceptron generates precise coordinate values ​​for determining the position of the gaze area by combining the extracted high-dimensional features and context information.

[0182] From the above description, it can be seen that the computer program product provided in the embodiment of the present application can achieve accurate positioning of the human eye gaze area through a pre-trained neural network model, thereby significantly improving the gaze tracking accuracy and response speed in complex scenes.

[0183] It should be understood by those skilled in the art that embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0184] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0185] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0186] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0187] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A method for localizing human eye attention based on a pre-trained neural network, characterized in that: The method comprises: Acquiring human eye movement data in real time through an eye tracking device, wherein the human eye movement data includes the position of the pupil and the direction of sight, and the eye tracking device includes at least one camera and a computing unit for image processing, wherein the camera is used to capture the movement of the human eye, and the computing unit is used to analyze and extract characteristic parameters of the human eye movement, wherein the characteristic parameters include the instantaneous value of the pupil position, the instantaneous value of the sight direction, the time derivative corresponding to the instantaneous value of the pupil position, the time derivative corresponding to the instantaneous value of the sight direction, the frequency characteristics of eye movement, pupil size change, and blinking frequency; Based on the human eye movement data, feature information related to the human eye gaze area is extracted through a pre-trained neural network model including an attention mechanism, wherein the neural network model includes a convolutional neural network and an attention mechanism module, wherein the convolutional neural network is used to extract high-dimensional features related to human eye gaze, and the attention mechanism module is used to perform weighted processing on the high-dimensional features; based on the channel attention mechanism, the features extracted by the convolutional neural network are weighted to enhance the attention to specific feature channels to a preset standard, and the feature map is weighted to enhance the positioning accuracy of the human eye gaze area to a preset standard through the spatial attention mechanism; The extracted feature information is input into a regression model, and the spatial coordinates of the human eye gaze area are calculated according to the feature information by the regression model. The regression model includes a multi-layer perceptron and an adaptive calibration mechanism. The multi-layer perceptron generates precise coordinate values ​​for determining the position of the gaze area by combining the extracted high-dimensional features and context information. The adaptive calibration mechanism compensates for deviations through periodic online fine-tuning.

2. The method for locating human eye attention based on a pre-trained neural network according to claim 1, characterized in that: The training process of the neural network model includes: Collecting a plurality of annotated human eye movement data, wherein the annotated human eye movement data includes the real position coordinates of the gaze target; Build a neural network model that includes a convolutional neural network, an attention mechanism module, and a recurrent neural network; Offline training is performed on the neural network model to minimize the error between the predicted position and the actual position of the gaze area by inputting the human eye movement data; The trained neural network model is verified and tested to improve the positioning accuracy of the neural network model in different application scenarios to the preset standard.

3. The method for locating human eye attention based on a pre-trained neural network according to claim 2, characterized in that: The method of obtaining the pupil position and sight direction of a human eye in real time by using an eye tracking device includes: The pupil position and sight direction of a human eye are obtained in real time by an eye tracking device, wherein the eye tracking device includes a plurality of cameras at different shooting positions; Fusion of multiple camera view data to improve the robustness of data acquisition to a preset standard.

4. The method for locating human eye attention based on a pre-trained neural network according to claim 1, characterized in that: The step of inputting the extracted feature information into a regression model and calculating the spatial coordinates of the eye gaze area according to the feature information through the regression model includes: Inputting the extracted feature information into a regression model so that the regression model combines the high-dimensional features and the time series information, updates the coordinate prediction through the output of the recursive neural network, and thus determines the spatial coordinates of the gaze area; Through the nonlinear mapping capability of the multi-layer perceptron, the high-dimensional features extracted by the convolutional neural network and the temporal features of the recurrent neural network are integrated to output precise coordinates of the gaze area.

5. The method for locating human eye attention based on a pre-trained neural network according to claim 1, characterized in that: After calculating the spatial coordinates of the area where the human eye is looking, it also includes: Based on the movement trajectory and timing changes of the human eye, a recursive neural network is used to analyze the collected human eye movement data in the time dimension to obtain the stability results of the visual target.

6. A human eye attention positioning device based on a pre-trained neural network, characterized in that: The device comprises: A data acquisition module, used for: acquiring human eye movement data in real time through an eye tracking device, wherein the human eye movement data includes the position of the pupil and the direction of sight, wherein the eye tracking device includes at least one camera and a computing unit for image processing, wherein the camera is used to capture the movement of the human eye, and the computing unit is used to analyze and extract characteristic parameters of the human eye movement, wherein the characteristic parameters include the instantaneous value of the pupil position, the instantaneous value of the sight direction, the time derivative corresponding to the instantaneous value of the pupil position, the time derivative corresponding to the instantaneous value of the sight direction, the frequency characteristics of eye movement, pupil size change, and blinking frequency; A feature extraction module, used to: extract feature information related to the human eye gaze area based on the human eye movement data through a pre-trained neural network model including an attention mechanism, wherein the neural network model includes a convolutional neural network and an attention mechanism module, wherein the convolutional neural network is used to extract high-dimensional features related to human eye gaze, and the attention mechanism module is used to perform weighted processing on the high-dimensional features; perform weighted processing on the features extracted by the convolutional neural network based on the channel attention mechanism to enhance the attention to specific feature channels to a preset standard, and perform weighted processing on the feature map through the spatial attention mechanism to improve the positioning accuracy of the human eye gaze area to a preset standard; The coordinate calculation module is used to: input the extracted feature information into a regression model, and calculate the spatial coordinates of the human eye gaze area according to the feature information through the regression model, the regression model includes a multi-layer perceptron and an adaptive calibration mechanism, the multi-layer perceptron generates accurate coordinate values ​​for determining the position of the gaze area by combining the extracted high-dimensional features and context information, and the adaptive calibration mechanism compensates for deviations through periodic online fine-tuning.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method for locating human eye attention based on a pre-trained neural network as described in any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for locating human eye attention based on a pre-trained neural network as described in any one of claims 1 to 5 are implemented.

9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the method for locating human eye attention based on a pre-trained neural network as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • System and method for realizing sight line estimation and attention analysis based on recursive convolutional neural network

    CN114387679A

  • Sight line estimation method and device based on improved face feature extraction

    CN117809353A