An adaptive line-of-sight estimation method, system, electronic device and storage medium
Through the adaptive gaze estimation method of multi-scale feature fusion, combined with the face and eye feature extraction network, the facial image feature extraction is dynamically guided, which solves the problem of insufficient fusion of the relationship between facial images and eye images in gaze estimation and achieves higher-precision gaze estimation.
Patent Information
- Application Number
- CN202211471537.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-11-23
AI Technical Summary
Existing appearance-based gaze estimation methods fail to effectively integrate the intrinsic relationship between facial images and eye images, resulting in poor gaze estimation results.
An adaptive gaze estimation method based on multi-scale feature fusion is adopted. By combining the facial feature extraction network and the eye guidance network, a multi-scale attention module is used to dynamically guide facial image feature extraction, fully exploring the intrinsic feature relationship between the face and the eyes.
Improved the accuracy and precision of gaze estimation, especially achieving higher gaze estimation performance on mobile devices, with an error range of 2.68cm to 3.8cm.
Smart Images

Figure CN115862095B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of gaze estimation, and in particular to a multi-scale feature fusion-based adaptive gaze estimation method and system, an electronic device and a storage medium. BACKGROUND
[0002] Gaze estimation is widely used in human-computer interaction, psychology, disease diagnosis and other fields. Gaze, as the main means for humans to obtain external information, can reveal human cognitive processing and cognitive processing defects. Scholars study various psychological disorders such as depression and autism through gaze estimation.
[0003] In the past few decades, many gaze estimation methods have been proposed. For example: 3D human eye model-based gaze estimation methods rely on some special devices such as infrared cameras, depth cameras, and high-resolution cameras, and many wearable eye movement tracking devices have been developed; appearance-based gaze estimation methods only need a webcam to capture images and directly learn the mapping relationship from images to gaze directions. Because of the low requirement for hardware devices, and the gradual popularity of mobile devices such as smartphones and tablets with webcams, appearance-based gaze estimation methods have more application prospects.
[0004] Appearance-based gaze estimation methods estimate gaze from face images or eye images. Convolutional neural networks (CNN) have been applied in various fields of computer vision due to their powerful feature extraction capabilities. Some appearance-based gaze estimation methods use CNN to estimate gaze from a single eye image or two eye images, some appearance-based gaze estimation methods estimate gaze from face images, and some appearance-based gaze estimation methods estimate gaze from both face images and eye images. However, these methods use simple techniques to fuse information from face images and eye images, such as simple concatenation or fully connected layers. Since gaze estimation is essentially a challenging task, simple feature splicing operations are not conducive to modeling the interaction between face images and eye images, and ignore the intrinsic relationship between faces and eyes. SUMMARY
[0005] The purpose of the present application is to provide an adaptive gaze estimation method, system, electronic device and storage medium that fully utilizes the intrinsic feature relationship between the face and the eye to achieve the purpose of adaptive gaze estimation.
[0006] To achieve the above purpose, the present application provides the following solutions:
[0007] In a first aspect, the present application provides a multi-scale feature fusion-based adaptive gaze estimation method, comprising:
[0008] obtaining a face image of a target person;
[0009] processing the face image of the target person to obtain eye-face position information of the target person; the eye-face position information comprises face boundary box information, left eye boundary box information and right eye boundary box information;
[0010] inputting target person information into an adaptive gaze estimation model to obtain a gaze line estimation result of the target person; the target person information comprises the face image and the eye-face position information of the target person;
[0011] the adaptive gaze estimation model is a model obtained by updating translation parameters and scaling parameters in a multi-scale feature fusion-based adaptive gaze estimation network using an eye guide network after training the multi-scale feature fusion-based adaptive gaze estimation network using first sample input data and sample measurement results corresponding to the first sample input data;
[0012] the eye guide network is configured to process second sample input data corresponding to the first sample input data using a deep learning algorithm to obtain the translation parameters and the scaling parameters;
[0013] the first sample input data is a face image and eye-face position information required for model training; the second sample input data is a left eye image, a right eye image and eye-face position information required for model training; and the sample measurement result is a gaze line measurement result required for model training.
[0014] Optionally, the training process of the adaptive gaze estimation model comprises:
[0015] constructing a sample data set; the sample data set comprises a plurality of sample data; the sample data comprises first sample input data, corresponding second sample input data and sample measurement results;
[0016] inputting the first sample input data into the multi-scale feature fusion-based adaptive gaze estimation network to obtain a sample prediction result;
[0017] calculating a network loss value using the sample prediction result and the sample measurement result;
[0018] updating network parameters of the multi-scale feature fusion-based adaptive gaze estimation network using the network loss value, updating translation parameters and scaling parameters in the updated network parameters using the eye guide network to obtain an updated multi-scale feature fusion-based adaptive gaze estimation network, iteratively optimizing until the number of iterations reaches a maximum number of iterations or the network loss value is less than a set threshold, and determining the last updated multi-scale feature fusion-based adaptive gaze estimation network as the adaptive gaze estimation model.
[0019] Optionally, the adaptive gaze estimation network based on multi-scale feature fusion comprises a convolutional layer, a global average pooling layer without translation parameters and scaling parameters, a first translation scaling layer, a channel dimension concatenation layer, a first multi-scale attention module, a second translation scaling layer, a second multi-scale attention module and a fully connected layer connected in turn; wherein the translation parameters and the scaling parameters in the first translation scaling layer are determined by the eye guide network; the translation parameters and the scaling parameters in the second translation scaling layer are determined by the eye guide network; the fully connected layer comprises a first fully connected block and a second fully connected block and a third fully connected block; the first fully connected block is used for inputting eye and face position information, and the second fully connected block is used for inputting the features output by the second multi-scale attention module; the input end of the third fully connected block is connected with the output end of the first fully connected block and the output end of the second fully connected block respectively; the convolutional layer is used for inputting a face image.
[0020] Optionally, the eye guide network comprises a first branch network, a second branch network, a third branch network, and a fully connected layer module connected with the output end of the first branch network, the output end of the second branch network and the output end of the third branch network; the first branch network comprises a first convolutional block, a channel dimension concatenation layer and a fully connected layer connected in turn; the first convolutional block is used for inputting a right eye image; the second branch network comprises a second convolutional block, a cross-view pooling layer and a fully connected layer connected in turn; the second convolutional block is used for inputting a left eye image; the third branch network is used for inputting eye and face position information; the fully connected layer module is used for outputting translation parameters and scaling parameters.
[0021] Optionally, the structure of the first multi-scale attention module is the same as that of the second multi-scale attention module.
[0022] The first multi-scale attention module comprises an SPC module, an SE module, a spatial attention map acquisition module and a summary module.
[0023] The input end of the SPC module is used for inputting the features output by the channel dimension concatenation layer, the output end of the SPC module is connected with the input end of the SE module, the first output end of the SE module is connected with the first input end of the summary module, the second output end of the SE module is connected with the input end of the spatial attention map acquisition module, the output end of the spatial attention map acquisition module is connected with the second input end of the summary module, and the third input end of the summary module is used for inputting the features output by the channel dimension concatenation layer; the summary module is used for outputting feature maps with multi-scale information under different receptive fields.
[0024] Optionally, the spatial attention map acquisition module is used for:
[0025] performing convolution, global average pooling and dimension transformation on the feature map output by the SE module to obtain a first feature subgraph and a second feature subgraph;
[0026] performing a normalization operation on the first feature subgraph and multiplying the second feature subgraph to obtain a two-dimensional feature;
[0027] performing dimension transformation and activation operation on the two-dimensional feature to obtain a spatial attention map.
[0028] Optionally, the face image of the target person is processed to obtain eye-face position information of the target person, specifically comprising:
[0029] processing the face image of the target person to obtain a left eye image and a right eye image of the target person;
[0030] obtaining eye-face position information of the target person according to the left eye image, the right eye image and the face image of the target person.
[0031] In a second aspect, the present application provides a self-adaptive gaze estimation system based on multi-scale feature fusion, comprising:
[0032] a face image acquisition module for acquiring a face image of a target person;
[0033] an eye-face position information calculation module for processing the face image of the target person to obtain eye-face position information of the target person; the eye-face position information includes face bounding box information, left eye bounding box information and right eye bounding box information;
[0034] a gaze line estimation result prediction module for inputting target person information into a self-adaptive gaze estimation model to obtain a gaze line estimation result of the target person; the target person information includes a face image and eye-face position information of the target person;
[0035] The self-adaptive gaze estimation model is a model obtained by updating translation parameters and scaling parameters in the self-adaptive gaze estimation network based on multi-scale feature fusion by using first sample input data and sample measured results corresponding to the first sample input data, and using an eye guide network to update the translation parameters and the scaling parameters in the self-adaptive gaze estimation network based on multi-scale feature fusion;
[0036] The eye guide network is used to process second sample input data corresponding to the first sample input data by using a deep learning algorithm to obtain translation parameters and scaling parameters;
[0037] The first sample input data is a face image and eye face position information required during model training; the second sample input data is a left eye image, a right eye image and eye face position information required during model training; and the sample measured result is a gaze line measured result required during model training.
[0038] In a third aspect, the present application provides an electronic device comprising a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to execute the computer program to enable the electronic device to perform the adaptive gaze estimation method according to the first aspect.
[0039] In a fourth aspect, the present application provides a computer readable storage medium storing a computer program, wherein the computer program is configured to be executed by a processor to implement the adaptive gaze estimation method according to the first aspect.
[0040] According to the embodiments of the present application, the following technical effects are achieved:
[0041] The present application outputs a gaze estimation result through an adaptive gaze estimation network based on multi-scale feature fusion, and better mines global features of a face image. An eye guidance network fuses features of double eye images, extracts feature parameters more focused on a gaze point, and dynamically guides feature extraction of the face image, so as to realize adaptive gaze estimation. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained based on these drawings without creative labor.
[0043] Figure 1 A flowchart of the adaptive gaze estimation method based on multi-scale feature fusion of the present application is shown in the figure.
[0044] Figure 2 A structure diagram of the adaptive gaze estimation network based on multi-scale feature fusion of the present application is shown in the figure.
[0045] Figure 3 A structure diagram of the eye guidance network of the present application is shown in the figure.
[0046] Figure 4 A structure diagram of the multi-scale attention module of the present application is shown in the figure.
[0047] Figure 5 A structure diagram of the SPC module of the present application is shown in the figure.
[0048] Figure 6A structure diagram of an adaptive gaze estimation system based on multi-scale feature fusion of the present application. DETAILED DESCRIPTION
[0049] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0050] Since a face image has more global information than an eye image, and the eye image is more focused on the gaze landing point, in order to fully utilize the inherent feature relationship between the face and the eye, the present application provides an adaptive gaze estimation method, system, electronic device and storage medium based on multi-scale feature fusion.
[0051] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0052] Embodiment one
[0053] The embodiment provides an adaptive gaze estimation method based on multi-scale feature fusion, and main points of the embodiment are as follows:
[0054] 1. The face feature extraction network is used as a backbone network to output a gaze estimation result, and a multi-scale attention mechanism is introduced to better mine global features of a face image.
[0055] 2. The eye feature extraction network is used as a guidance network to fuse features of binocular images, extract feature parameters more focused on a gaze point, and dynamically guide feature extraction of the face image, so as to realize adaptive gaze estimation.
[0056] As shown in Figure 1 The embodiment provides an adaptive gaze estimation method based on multi-scale feature fusion, and specifically includes the following steps.
[0057] Step 100: Obtain a face image of a target person.
[0058] Step 200: Process the face image of the target person to obtain eye-face position information of the target person; the eye-face position information includes face boundary box information, left eye boundary box information and right eye boundary box information.
[0059] Step 300: Input the target person information into an adaptive gaze estimation model to obtain a gaze line estimation result of the target person; the target person information includes the face image and the eye-face position information of the target person.
[0060] The adaptive gaze estimation model is a model obtained by training a multi-scale feature fusion-based adaptive gaze estimation network using first sample input data and sample measured results corresponding to the first sample input data, and updating translation parameters and scaling parameters in the multi-scale feature fusion-based adaptive gaze estimation network using an eye guide network.
[0061] The eye guide network is configured to process second sample input data corresponding to the first sample input data using a deep learning algorithm to obtain the translation parameters and the scaling parameters.
[0062] The first sample input data is face images and eye-face position information required for model training; the second sample input data is left eye images, right eye images and eye-face position information required for model training; and the sample measured results are gaze line measured results required for model training.
[0063] The training process of the adaptive gaze estimation model includes the following steps.
[0064] (1) Constructing a sample data set; the sample data set includes a plurality of sample data; the sample data includes first sample input data, corresponding second sample input data and sample measured results.
[0065] (2) Inputting the first sample input data into the multi-scale feature fusion-based adaptive gaze estimation network to obtain sample predicted results.
[0066] (3) Calculating a network loss value using the sample predicted results and the sample measured results.
[0067] (4) Updating network parameters of the multi-scale feature fusion-based adaptive gaze estimation network using the network loss value, updating translation parameters and scaling parameters in the updated network parameters using the eye guide network, obtaining an updated multi-scale feature fusion-based adaptive gaze estimation network, iteratively optimizing until the number of iterations reaches a maximum number of iterations or the network loss value is less than a set threshold, and determining the last updated multi-scale feature fusion-based adaptive gaze estimation network as the adaptive gaze estimation model.
[0068] An example is that the training process of the adaptive gaze estimation model is determined by the following steps.
[0069] Step 1: The GazeCapture dataset and the MPIIFaceGaze dataset are selected for training and testing the model, respectively. From the GazeCapture dataset with more than 1400 samples (a total of more than 240,000 face images), a 7:2:1 ratio is randomly divided into a training set, a validation set, and a test set. The training set is used to learn the mapping relationship between the image and the gaze point, the validation set is used to optimize the model during training, and the test set is used to evaluate the performance of the model in predicting the gaze point. The images in the training set, the validation set, and the test set have no overlap. From the MPIIFaceGaze dataset with 15 samples (a total of 37667 face images), 13 samples are selected as the training set and 2 samples are selected as the test set. Because the number of samples in this dataset is small, in order to better verify the performance of the model, cross-validation is performed on the MPIIFaceGaze dataset (all samples are tested as a test set), and the average value of 8 experiments is taken.
[0070] Step 2: According to the eye frame coordinates and face frame coordinates provided by the dataset, the face images in the GazeCapture dataset and the MPIIFaceGaze dataset are preprocessed, the left and right eye images are cropped, and the position information of the left and right eye images relative to the face image is obtained. The size of the face image is processed to 224*224*3, and the size of the left and right eye images is processed to 112*112*3 (representing the length, width, and RGB three-channel number of the image, respectively), and the pixel value is normalized from [0, 255] to [0, 1] interval.
[0071] Step 3: The preprocessed images are input into the adaptive gaze estimation network based on multi-scale feature fusion for training. The structure of the adaptive gaze estimation network based on multi-scale feature fusion (i.e. the face feature extraction network) is as shown in Figure 2 The structure of the eye guide network is as shown in Figure 3
[0072] The input of the face feature extraction network includes the face image, the face boundary box information, and the left and right eye boundary box information; the input of the eye guide network includes the left eye image, the right eye image, the face boundary box information, and the left and right eye boundary box information.
[0073] The face feature extraction network serves as the backbone network, the input is the face image, the face features are extracted by multiple convolutional layers, and are adaptively adjusted by translation parameters (Pa parameters) and scaling parameters (Pm parameters), then are input into the multi-scale attention module, and finally are concatenated with the eye-face position information (coordinate information of the left and right eye image boundary box and coordinate information of the face image boundary box) by the fully connected layer, and the gaze point coordinates are output; the eye guide network inputs the left and right eye images and the eye-face position information, extracts features by convolutional layers, fuses feature blocks, and finally outputs the Pa and Pm parameters by the fully connected layer to dynamically guide the face feature extraction.
[0074] The adaptive gaze estimation network based on multi-scale feature fusion comprises a convolutional layer, a global average pooling layer without translation parameters and scaling parameters, a first translation scaling layer, a channel dimension splicing layer, a first multi-scale attention module, a second translation scaling layer, a second multi-scale attention module and a fully connected layer connected in turn; wherein the translation parameters and scaling parameters in the first translation scaling layer are determined by the eye guide network; the translation parameters and scaling parameters in the second translation scaling layer are determined by the eye guide network; the fully connected layer comprises a first fully connected block and a second fully connected block and a third fully connected block; the first fully connected block is used for inputting eye face position information, and the second fully connected block is used for inputting the features output by the second multi-scale attention module; the input end of the third fully connected block is respectively connected with the output end of the first fully connected block and the output end of the second fully connected block; the convolutional layer is used for inputting a face image.
[0075] After the face image is extracted by the convolutional layer, the normalization layer containing scaling operation and translation operation in the low-dimensional feature and the high-dimensional feature is replaced by the global average pooling layer without scaling parameters and translation parameters, and the two parameters generated by the eye guide network are used to replace the translation parameters and scaling parameters, so as to realize adaptive re-extraction of face features. Then the adjusted low-dimensional features and high-dimensional features are spliced in the channel dimension and input into the multi-scale attention module, so as to effectively utilize the spatial information of different scale features and establish the dependency relationship between the channels of the features, and better capture the global information of the face image. Finally, the gaze point coordinates are output by splicing the eye face position information and the fully connected layer after adaptive adjustment and multi-scale attention module. Wherein, Figure 2 GN(.) in the formula (1) represents the global average pooling layer without translation parameters and scaling parameters, Stack represents the channel dimension splicing layer, and FC represents the fully connected layer.
[0076] The eye guide network comprises a first branch network, a second branch network, a third branch network and a fully connected layer module connected with the output end of the first branch network, the output end of the second branch network and the output end of the third branch network; the first branch network comprises a first convolutional block, a channel dimension splicing layer and a fully connected layer connected in turn; the first convolutional block is used for inputting a right eye image; the second branch network comprises a second convolutional block, a cross-view pooling layer and a fully connected layer connected in turn; the second convolutional block is used for inputting a left eye image; the third branch network is used for inputting eye face position information; the fully connected layer module is used for outputting translation parameters and scaling parameters.
[0077] Eye guidance network: Because the shape and structure of the left and right eyes are similar, the features of the left and right eye images are fused. The right eye image and the left eye image after horizontal flip are input into the network model, and the low-dimensional features and high-dimensional features are spliced in the channel dimension. The feature maps extracted from the lower layer retain more spatial information, while the feature maps extracted from the higher layer have stronger representation ability. Then, the left and right eye fused features are spliced in the channel dimension and cross-view pooling operation respectively, and finally the face eye position information is output through the fully connected layer. Pa and Pm two parameters, dynamically guide and adjust the face feature extraction. The cv-pool in FIG. 3 is a cross-view pooling layer.
[0078] The structure of the first multi-scale attention module and the structure of the second multi-scale attention module are the same. Taking the first multi-scale attention module as an example for description.
[0079] The first multi-scale attention module includes an SPC module, an SE module, a spatial attention map acquisition module, and a summary module; the input end of the SPC module is used for inputting the features output by the channel dimension splicing layer, the output end of the SPC module is connected with the input end of the SE module, the first output end of the SE module is connected with the first input end of the summary module, the second output end of the SE module is connected with the input end of the spatial attention map acquisition module, the output end of the spatial attention map acquisition module is connected with the second input end of the summary module, and the third input end of the summary module is used for inputting the features output by the channel dimension splicing layer; the summary module is used for outputting feature maps with multi-scale information under different receptive fields.
[0080] The spatial attention map acquisition module is configured to:
[0081] The feature maps output by the SE module are convolved, globally averaged pooled, and dimensionally transformed to obtain first feature subgraphs and second feature subgraphs;
[0082] After the first feature subgraphs are normalized, they are multiplied by the second feature subgraphs to obtain two-dimensional features.
[0083] The two-dimensional features are dimensionally transformed and activated to obtain a spatial attention map.
[0084] As Figure 4 and Figure 5As shown, the SPC module performs convolution on the input feature X (size: CxHxW) with different kernel sizes (kernel sizes: 3x3, 5x5, 7x7, 9x9 respectively. Feature size after convolution: C / 4xHxW) to obtain receptive fields of different scales and extract information of different scales, and then splicing in the channel dimension. Then the SE module (global average pooling operation in the channel dimension, compressed into a Cx1x1 vector) extracts the weighting value of each group of channels, and finally the softmax normalization is performed after the channel dimension multiplication with the feature, so as to weight the channels. The rescaled feature map Xc focuses on useful channels, but the pixels in the same channel still share the same weight. Therefore, further calculation is performed in the spatial dimension based on Xc.
[0085] As shown, the feature Xc is sent into a 1x1 convolution layer, global average pooling and dimension transformation are performed, and feature maps Q (1xC / 2) and V (C / 2xHW) are obtained respectively. After the softmax normalization of Q, the two-dimensional tensor multiplication with V is performed, and a two-dimensional feature of 1xHW is obtained, and then dimension transformation and sigmoid activation function are performed to obtain a spatial attention map As (1xHxW). Finally, the matrix multiplication of As and Xc in the spatial dimension is performed, and the original input feature X is added, and finally the feature map with multi-scale information under different receptive fields is output.
[0086] Experimental setup: The experiment is completed on a high-performance computing platform: system windows10, CPU AMD5800x, GPU RTX3080, memory 32g.
[0087] Adaptive gaze estimation method
[0088] The method fully utilizes the feature relationship between the face and the eyes, the face feature extraction network outputs the gaze estimation result as the backbone network, and the eye feature extraction network is used as the guidance network to fuse the features of the two eye images and extract feature parameters more focused on the gaze point, dynamically guide the feature extraction of the face, thereby realizing adaptive gaze estimation.
[0089] Multi-scale attention mechanism
[0090] A multi-scale attention module is designed to effectively utilize the spatial information of different scale features and establish the dependency relationship between the channels of the features, so as to better capture the global information of the face image.
[0091] Table 1: Comparison experiment table
[0092]
[0093] Experimental results: comparative experiments were carried out on the MPIIFaceGaze dataset and the GazeCapture dataset. Compared with the mainstream appearance-based gaze estimation method, the adaptive gaze estimation method proposed by us achieved the optimal performance. On the MPIIFaceGaze dataset, the method proposed by us achieved an error of 3.8 cm; on the GazeCapture dataset, the device is divided into mobile phones and tablets, and the method proposed by us achieved an error of 2.68 cm and 3.14 cm, respectively.
[0094] Embodiment two
[0095] In order to perform the method corresponding to the above-mentioned embodiment one, to realize the corresponding functions and technical effects, the following provides an adaptive gaze estimation system based on multi-scale feature fusion.
[0096] As shown in Figure 6 , the adaptive gaze estimation system based on multi-scale feature fusion provided by the embodiment comprises:
[0097] A face image acquisition module 1 is configured to acquire a face image of a target person.
[0098] An eye-face position information calculation module 2 is configured to process the face image of the target person to obtain eye-face position information of the target person; the eye-face position information comprises face boundary box information, left eye boundary box information and right eye boundary box information.
[0099] A gaze line estimation result prediction module 3 is configured to input target person information into an adaptive gaze estimation model to obtain a gaze line estimation result of the target person; the target person information comprises a face image and eye-face position information of the target person.
[0100] The adaptive gaze estimation model is a model obtained by updating translation parameters and scaling parameters in the adaptive gaze estimation network based on multi-scale feature fusion by using first sample input data and sample measured results corresponding to the first sample input data.
[0101] The eye guide network is configured to process second sample input data corresponding to the first sample input data by using a deep learning algorithm to obtain translation parameters and scaling parameters.
[0102] The first sample input data is a face image and eye-face position information required during model training; the second sample input data is a left eye image, a right eye image and eye-face position information required during model training; and the sample measured result is a gaze line measured result required during model training.
[0103] Embodiment three
[0104] The electronic device comprises a memory and a processor. The memory is configured to store a computer program. The processor is configured to execute the computer program to enable the electronic device to perform the adaptive line-of-sight estimation method of Embodiment One.
[0105] Optionally, the electronic device can be a server.
[0106] In addition, the embodiments of the present application further provide a computer readable storage medium storing a computer program. The computer program is executed by a processor to implement the adaptive line-of-sight estimation method of Embodiment One.
[0107] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts of each embodiment can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.
[0108] The principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges can be changed. In conclusion, the content of the specification should not be understood as a limitation of the present application.
Claims
1. An adaptive line of sight estimation method based on multi-scale feature fusion, characterized in that: The adaptive line of sight estimation method comprises: Obtaining a facial image of the target person; Processing the facial image of the target person to obtain eye and face position information of the target person; the eye and face position information includes facial bounding box information, left eye bounding box information, and right eye bounding box information; Inputting target person information into an adaptive line of sight estimation model to obtain a gaze line estimation result of the target person; the target person information includes a facial image and eyelid position information of the target person; The adaptive gaze estimation model is a model obtained by training an adaptive gaze estimation network based on multi-scale feature fusion using first sample input data and sample measurement results corresponding to the first sample input data, and updating translation parameters and scaling parameters in the adaptive gaze estimation network based on multi-scale feature fusion using an eye guidance network; The eye guidance network is used to process second sample input data corresponding to the first sample input data using a deep learning algorithm to obtain a translation parameter and a scaling parameter; The first sample input data is the facial image and eyelid position information required for model training; the second sample input data is the left eye image, right eye image and eyelid position information required for model training; the sample measured result is the gaze line measured result required for model training; The adaptive gaze estimation network based on multi-scale feature fusion includes a convolutional layer, a global average pooling layer without translation parameters and scaling parameters, a first translation scaling layer, a channel dimension splicing layer, a first multi-scale attention module, a second translation scaling layer, a second multi-scale attention module and a fully connected layer connected in sequence; wherein, the translation parameters and scaling parameters in the first translation scaling layer are determined by the eye guidance network; the translation parameters and scaling parameters in the second translation scaling layer are determined by the eye guidance network; the fully connected layer includes a first fully connected block, a second fully connected block and a third fully connected block; the first fully connected block is used to input eye and face position information, and the second fully connected block is used to input features output by the second multi-scale attention module; the input end of the third fully connected block is respectively connected to the output end of the first fully connected block and the output end of the second fully connected block; the convolutional layer is used to input a facial image; The eye guidance network includes a first branch network, a second branch network, a third branch network, and a fully connected layer module connected to the output end of the first branch network, the output end of the second branch network, and the output end of the third branch network; the first branch network includes a first convolution block, a channel dimension splicing layer, and a fully connected layer connected in sequence; the first convolution block is used to input the right eye image; the second branch network includes a second convolution block, a cross-view pooling layer, and a fully connected layer connected in sequence; the second convolution block is used to input the left eye image; the third branch network is used to input eye face position information; the fully connected layer module is used to output translation parameters and scaling parameters.
2. The adaptive line of sight estimation method based on multi-scale feature fusion according to claim 1, characterized in that: The training process of the adaptive line of sight estimation model is as follows: Constructing a sample data set; the sample data set includes a plurality of sample data; the sample data includes first sample input data and corresponding second sample input data and sample measurement results; Inputting the first sample input data into an adaptive line of sight estimation network based on multi-scale feature fusion to obtain a sample prediction result; Calculate the network loss value using the sample prediction results and sample measured results; The network parameters of the adaptive gaze estimation network based on multi-scale feature fusion are updated using the network loss value, and the translation parameters and scaling parameters in the updated network parameters are updated using the eye guidance network to obtain an updated multi-scale feature fusion adaptive gaze estimation network. The iterative optimization is performed until the number of iterations reaches a maximum number of iterations or the network loss value is less than a set threshold, and the last updated multi-scale feature fusion adaptive gaze estimation network is determined as the adaptive gaze estimation model.
3. The adaptive line of sight estimation method based on multi-scale feature fusion according to claim 1, characterized in that: The structure of the first multi-scale attention module is the same as the structure of the second multi-scale attention module; The first multi-scale attention module includes an SPC module, an SE module, a spatial attention map acquisition module and a summary module; The input end of the SPC module is used to input the features output by the channel dimension splicing layer, the output end of the SPC module is connected to the input end of the SE module, the first output end of the SE module is connected to the first input end of the summary module, the second output end of the SE module is connected to the input end of the spatial attention map acquisition module, the output end of the spatial attention map acquisition module is connected to the second input end of the summary module, and the third input end of the summary module is used to input the features output by the channel dimension splicing layer; The aggregation module is used to output feature maps with multi-scale information under different receptive fields.
4. The adaptive line of sight estimation method based on multi-scale feature fusion according to claim 3, characterized in that: The spatial attention map acquisition module is used to: Performing convolution, global average pooling, and dimensionality transformation on the feature map output by the SE module to obtain a first feature submap and a second feature submap; After normalizing the first feature subgraph, perform a two-dimensional tensor multiplication on the second feature subgraph to obtain a two-dimensional feature; Perform dimension transformation and activation operations on the two-dimensional features to obtain the spatial attention map.
5. The adaptive line of sight estimation method based on multi-scale feature fusion according to claim 1, characterized in that: The processing of the facial image of the target person to obtain the eye and face position information of the target person specifically includes: Processing the facial image of the target person to obtain a left-eye image and a right-eye image of the target person; The eye and face position information of the target person is obtained according to the left eye image, right eye image and face image of the target person.
6. An adaptive sight line estimation system based on multi-scale feature fusion, used to implement the adaptive sight line estimation method based on multi-scale feature fusion according to any one of claims 1 to 5, characterized in that: include: A facial image acquisition module is used to acquire the facial image of the target person; An eye and face position information calculation module is used to process the facial image of the target person to obtain the eye and face position information of the target person; the eye and face position information includes facial bounding box information, left eye bounding box information, and right eye bounding box information; The gaze estimation result prediction module is used to input the target person information into the adaptive gaze estimation model to obtain the target person's gaze line estimation result; The target person information includes the target person's facial image and eye and face position information; The adaptive gaze estimation model is a model obtained by training an adaptive gaze estimation network based on multi-scale feature fusion using first sample input data and sample measurement results corresponding to the first sample input data, and updating translation parameters and scaling parameters in the adaptive gaze estimation network based on multi-scale feature fusion using an eye guidance network; The eye guidance network is used to process second sample input data corresponding to the first sample input data using a deep learning algorithm to obtain a translation parameter and a scaling parameter; The first sample input data is the facial image and eyelid position information required for model training; the second sample input data is the left eye image, right eye image and eyelid position information required for model training; the sample measured results are the gaze line measured results required for model training.
7. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the adaptive sight line estimation method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that The device stores a computer program, which, when executed by a processor, implements the adaptive sight line estimation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Monocular fixation point estimation method and system based on mixed attention mechanism
CN114582009A
Line-of-sight estimation method based on cooperation network
CN114898453A