3D Gaze Estimation Method, Device, Equipment and Storage Medium Based on MLP

By building a UM-Net network based on MLP, combining feature extraction and feature splicing modules, the problem of slow line of sight estimation speed in situations with high real-time requirements is solved, and efficient and high-precision three-dimensional line of sight detection is achieved.

CN115951775BActive Publication Date: 2025-07-18CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211621733.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-16
Publication Date
2025-07-18
Estimated Expiration
2042-12-16

AI Technical Summary

Technical Problem

The existing three-dimensional line of sight estimation method based on deep convolutional neural networks (CNNs) has problems such as complex network structure and insufficient model loading speed in situations where real-time requirements are high, making it difficult to achieve efficient and high-precision line of sight detection.

Method used

UM-Net network based on MLP, including left eye, right eye and face feature extraction branch, feature dimensionality reduction and regression are performed through feature splicing module and full connection layer, feature refinement is used by Mixer Layer module, combining jump connection and layer normalization, and efficient estimation of three-dimensional line of sight direction is achieved.

Benefits of technology

It realizes efficient and high-precision three-dimensional line of sight estimation, fast prediction speed, and close to CNNs-based networks, and is suitable for application scenarios with high real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115951775B_ABST
    Figure CN115951775B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional gaze estimation method, device, equipment and storage medium based on MLP. The method includes: constructing a UM-Net network based on MLP, which includes three branches, namely a left-eye feature extraction branch, a right-eye feature extraction branch and a face feature extraction branch; a feature splicing module connected to all three branches, and two fully connected layers connected to the feature splicing module in sequence; obtaining a dataset to be measured, preprocessing it and then inputting it into the UM-Net network; respectively extracting left-eye image features, right-eye image features and face image features through the three branches, after feature splicing, performing feature dimension reduction through the first fully connected layer, and regressing the three-dimensional gaze direction through the second fully connected layer. The present invention uses a network based on MLP for gaze estimation. The network structure is simple, has a large throughput, a fast prediction speed, and the estimation accuracy is comparable to that of a network based on CNNs, having the advantages of high efficiency, high accuracy and real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gaze estimation, and particularly to a three-dimensional gaze estimation method, device, equipment and storage medium based on MLP. Background Art

[0002] Gaze is one of the most important non-verbal communication cues, which contains rich human intention information and enables researchers to deeply understand human cognition and behavior. It is widely used in medical treatment, assisted driving, marketing, human-computer interaction and other fields. High-precision gaze estimation methods are crucial for its applications. With the rise of deep convolutional neural networks (CNNs) in the field of computer vision and the public release of a large number of datasets, researchers have begun to use CNNs for appearance-based three-dimensional gaze estimation methods. Researchers such as Chen Z proposed the Dilated-Net, a dilated convolutional network, which uses dilated convolution to extract features of the face and both eyes. The accuracy of appearance-based three-dimensional gaze estimation is improved by using a deep neural network to extract higher-resolution features from eye images. Researchers such as Cheng Y proposed a plug-and-play self-contradictory framework to simplify gaze features in order to reduce the interference of gaze-independent factors and reduce the influence of illumination, personal appearance and even facial expressions on the learning of gaze estimation. However, due to reasons such as the complex structure of CNNs and the slow model loading speed, such methods still need to be further improved in occasions with high real-time requirements. Therefore, it is of great significance to design an efficient and high-precision three-dimensional gaze estimation network. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide an efficient and high-precision three-dimensional gaze estimation network to meet the requirements of efficient and high-precision detection of three-dimensional gaze in occasions with high real-time requirements. To solve this technical problem, the technical solution adopted by the present invention is: a three-dimensional gaze estimation method, device, equipment and storage medium based on MLP.

[0004] According to the first aspect of the present invention, a three-dimensional gaze estimation method based on MLP includes the following steps:

[0005] Construct a UM-Net network (Use-MLP Network) based on MLP, where the UM-Net network includes three branches, namely a left-eye feature extraction branch, a right-eye feature extraction branch and a face feature extraction branch; and a feature splicing module connected to all three branches, and a fully connected layer FC1 and a fully connected layer FC2 connected to the feature splicing module in sequence;

[0006] Obtain a dataset to be measured, including left-eye images, right-eye images and face images, and perform preprocessing on them respectively;

[0007] Input the preprocessed image into the UM-Net network; extract the left-eye image features, right-eye image features, and face image features through three branches respectively. After splicing through the feature splicing module, perform feature dimensionality reduction through the fully connected layer FC1, and regress the three-dimensional gaze direction through the fully connected layer FC2.

[0008] Further, the feature extraction branch includes a feature extraction module, N Mixer Layer modules, a global average pooling layer GAP, and a fully connected layer FC connected in sequence;

[0009] First, the feature extraction module splits the input image into image patches; then projects each image patch into a 512-dimensional space through a fully connected layer, and after projection, obtains a sequence of image feature patches;

[0010] Then send the sequence of image feature patches into N Mixer Layer modules to perform feature refinement in the column direction and feature refinement in the row direction on the sequence of image feature patches, and repeatedly pass the sequence of image feature patches through N Mixer Layer modules to refine the image feature information;

[0011] Next, the global average pooling layer GAP regularizes the entire network model in structure to prevent overfitting;

[0012] Finally, use the fully connected layer FC to regress the required image features respectively.

[0013] Further, the Mixer Layer module includes a token-mixing MLP module and a channel-mixing MLP module;

[0014] The token-mixing MLP module and the channel-mixing MLP module are alternately stacked to perform feature refinement in the column direction and feature refinement in the row direction on the sequence of image feature patches.

[0015] Further, the token-mixing MLP module contains an MLP1 module, and the channel-mixing MLP module contains an MLP2 module;

[0016] The token-mixing MLP module first processes the sequence of image feature patches X ∈ R 16×512After transposing, the MLP1 module is applied to each column of the sequence of image feature blocks to enable communication between different spatial positions of the sequence of image feature blocks, and the parameters of the MLP1 module are shared among all columns. The obtained output is transposed again, and then in the channel-mixing MLP module, the MLP2 module is applied to each row of the sequence of image feature blocks to enable communication between different channels of the sequence of image feature blocks, and the parameters of the MLP2 module are shared among all rows;

[0017] The Mixer Layer module also uses skip connections and layer normalization.

[0018] Further, for the input sequence of image feature blocks X ∈ R 16×512 , the operation process of the MixerLayer module is expressed by the following formula:

[0019] U *,i = M1(LayerNorm(X) *,i ), i ∈ [1, 512]

[0020] Y j,* = M2(LayerNorm(U) j,* ), j ∈ [1, 16]

[0021] M1 and M2 represent the MLP1 module and the MLP2 module respectively. LayerNorm(X) *,i represents the i-th column of the sequence of image feature blocks after layer normalization. LayerNorm(U) j,* represents the j-th row of the sequence of image feature blocks after layer normalization. U *,i represents the i-th column of the sequence of image feature blocks after the action of the MLP1 module. Y j,* represents the j-th row of the sequence of image feature blocks after the action of the MLP2 module.

[0022] Further, each MLP1 module or MLP2 module contains two fully connected layers and a non-linear activation function; for the input of the MLP1 module or MLP2 module, the operation process is expressed by the following formula:

[0023]

[0024] Φ represents the non-linear activation function acting on the input elements. W1 and W2 represent the two fully connected layers in the MLP1 module or MLP2 module. σ represents the output after the action of the input through the MLP1 module or MLP2 module.

[0025] Further, the three-dimensional line-of-sight direction is represented by the pitch angle in the vertical direction and the yaw angle in the horizontal direction:

[0026]

[0027] f, l, r respectively represent the input face image, left-eye image, and right-eye image of the model, represents the feature extraction module of the network, C represents the connection of the extracted left-eye image features, right-eye image features, and face image features, and δ represents the regression of the three-dimensional line-of-sight direction using a fully connected layer;

[0028] Calculate the three-dimensional vector representing the line-of-sight direction according to the pitch angle and yaw angle The calculation formula is as follows:

[0029] x = cos(pitch)cos(yaw)

[0030] y = cos(pitch)sin(yaw)

[0031] z = sin(pitch)

[0032] Three-dimensional vector And the true direction vector The included angle between them is the evaluation index of three-dimensional line-of-sight estimation, that is, the line-of-sight angle error θ. The loss function adopts the mean square error function MSE. The total number of predicted three-dimensional line-of-sight vectors is n. The calculation formulas are as follows:

[0033]

[0034]

[0035] According to the second aspect of the present invention, a three-dimensional line-of-sight estimation device based on MLP for implementing the method includes the following modules:

[0036] A construction module for constructing a UM-Net network based on MLP. The UM-Net network includes three branches, namely a left-eye feature extraction branch, a right-eye feature extraction branch, and a face feature extraction branch; and a feature splicing module connected to all three branches, and a fully connected layer FC1 and a fully connected layer FC2 connected to the feature splicing module in sequence;

[0037] A preprocessing module for obtaining a test data set, including left-eye images, right-eye images, and face images, and performing preprocessing on them respectively;

[0038] An estimation module for inputting the preprocessed image into the UM-Net network; extracting left-eye image features, right-eye image features, and face image features through three branches respectively, performing splicing through a feature splicing module, reducing the features through a fully connected layer FC1, and regressing the three-dimensional line-of-sight direction through a fully connected layer FC2.

[0039] According to the third aspect of the present invention, an electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the three-dimensional line-of-sight estimation method are implemented.

[0040] According to the fourth aspect of the present invention, a storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the three-dimensional line-of-sight estimation method are implemented.

[0041] The technical solution provided by the present invention has the following beneficial effects:

[0042] 1. The method of the present invention does not require CNNs and uses an MLP-based network for line-of-sight estimation, so the network structure is simple.

[0043] 2. Due to the simple structure and large throughput of the MLP-based network, the prediction speed of the three-dimensional line-of-sight direction is fast, and the line-of-sight estimation accuracy is comparable to that of the CNNs-based network. The efficient and high-precision line-of-sight estimation model has good application prospects in the fields that require real-time line-of-sight estimation. Description of the Drawings

[0044] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0045] Figure 1 is the overall flowchart of a three-dimensional line-of-sight estimation method based on MLP of the present invention;

[0046] Figure 2 is the structure diagram of the UM-Net network of the present invention;

[0047] Figure 3 is the structure diagram of the Mixer Layer module of the present invention;

[0048] Figure 4 is the structure diagram of the MLP module of the present invention;

[0049] Figure 5 is the accuracy comparison of the method of the present invention and several advanced line-of-sight estimation methods on two datasets;

[0050] Figure 6 is the average angle error comparison of the method of the present invention and the CNNs-based line-of-sight estimation method on the MPIIFaceGaze dataset;

[0051] Figure 7 Comparison of the average angular error between the method of the present invention and the CNNs-based gaze estimation method on the EyeDiap dataset;

[0052] Figure 8 Comparison of the prediction time between the method of the present invention and the CNNs-based gaze estimation method on the MPIIFaceGaze dataset;

[0053] Figure 9 Comparison of the prediction time between the method of the present invention and the CNNs-based gaze estimation method on the EyeDiap dataset;

[0054] Figure 10 Schematic structural diagram of a three-dimensional gaze estimation device based on MLP according to the present invention;

[0055] Figure 11 Schematic structural diagram of an electronic device according to the present invention. Detailed implementation manners

[0056] For a clearer understanding of the technical features, objectives, and effects of the present invention, the detailed implementation manners of the present invention will now be described in detail with reference to the accompanying drawings.

[0057] Currently, appearance-based gaze estimation faces many challenges, such as head movement and subject differences, especially in an unconstrained environment. These factors have a great impact on eye appearance and complicate eye appearance. Traditional appearance-based gaze estimation methods have weak fitting capabilities and cannot well cope with these challenges. Neural networks have shown good performance in gaze estimation. Different from other networks that use CNNs, refer to Figure 1 , the present invention provides a three-dimensional gaze estimation method based on MLP, which mainly includes the following steps:

[0058] S1: Construct a UM-Net network based on MLP, as Figure 2 shown. This UM-Net network includes three branches, namely the left-eye feature extraction branch, the right-eye feature extraction branch, and the face feature extraction branch; and a feature splicing module connected to all three branches, a fully connected layer FC1 and a fully connected layer FC2 connected to the feature splicing module in sequence;

[0059] S2: Obtain a dataset to be measured, including left-eye images, right-eye images, and face images, and perform preprocessing on them respectively;

[0060] S3: Input the preprocessed images into the UM-Net network; extract left-eye image features, right-eye image features, and face image features through the three branches respectively. After splicing through the feature splicing module, perform feature dimensionality reduction through the fully connected layer FC1, and regress the three-dimensional gaze direction through the fully connected layer FC2.

[0061] As Figure 2 shown, each feature extraction branch includes a feature extraction module, N Mixer Layer modules, a global average pooling layer GAP, and a fully connected layer FC connected in sequence. Next, the network and method provided by the present invention will be described in detail from the following aspects:

[0062] 1. Feature extraction module

[0063] Feature extraction is crucial for most learning-based tasks. Due to the complexity of eye appearance, effectively extracting features from eye appearance is a challenge. The quality of the extracted features determines the accuracy of gaze estimation.

[0064] The feature extraction module of UM-Net first splits the input image into image patches (without overlap between each image patch) to facilitate information exchange and feature integration in the image. Assuming the input image resolution is (J, K) and the resolution of the split image patches is (P, P), then the number H of image patches is:

[0065]

[0066] The original image resolution input to UM-Net is (64, 64). The present invention first splits the input image into 16 image patches with a resolution of (16, 16), and then projects each image patch into a 512-dimensional space through a fully connected layer. All image patches use the same projection matrix for linear projection. After projection, the image feature patch sequence X ∈ R 16×512 . The fully connected operation projects non-overlapping image patches into a higher hidden dimension, not only retaining the key features of the image but also facilitating information fusion in subsequent local regions.

[0067] Next, the image feature patch sequence is fed into N Mixer Layer modules. The Mixer Layer module does not use convolution or self-attention, but only uses a simple-structured MLP and applies it repeatedly to spatial positions or feature channels.

[0068] UM-Net uses the token-mixing MLP module and the channel-mixing MLP module in the Mixer Layer module to perform feature refinement along the column direction and feature refinement along the row direction on the image feature patch sequence respectively. They are stacked alternately, which helps to support the communication of the two input dimensions. The network repeatedly passes the image feature patch sequence through N MixerLayers to refine the image feature information. Then UM-Net uses the global average pooling GAP to regularize the entire network model structurally to prevent overfitting, and finally uses a fully connected layer to regress the required image features respectively.

[0069] 2. Three branches of the network

[0070] The above-mentioned feature extraction module is used to extract image features from binocular and face images.

[0071] Binocular feature extraction branch: The line-of-sight direction is related to the apparent height of the eyes, and any perturbation of the line-of-sight direction will cause changes in the eye appearance. For example, the rotation of the eyeball will change the position of the iris and the shape of the eyelids, resulting in a change in the line-of-sight direction. This relationship enables the estimation of the line-of-sight from the eye appearance. However, with the change of the environment, the eye image features will also be interfered by redundant information. Using the MLP model can directly extract deep features from the eye images, which is more robust to environmental changes. Therefore, UM-Net uses the feature extraction module to extract 256-dimensional features from the left-eye image and the right-eye image respectively.

[0072] Face feature extraction branch: The three-dimensional line-of-sight direction depends not only on the eye appearance (such as iris position, eye opening degree, etc.), but also on the head pose. The face image contains head pose information, so UM-Net uses the feature extraction module to extract 32-dimensional features from the face image to supplement more abundant information. The parameters of the three feature extraction branches are shared.

[0073] After UM-Net uses the three feature extraction branches to regress the left-eye image features, the right-eye image features and the face image features, the extracted features are concatenated to combine the features from different input images. Then the network sends the 544-dimensional features to the first fully connected layer FC1 to reduce them to 256 dimensions, and then uses the second fully connected layer FC2 to regress the three-dimensional line-of-sight direction. This three-dimensional line-of-sight direction is represented by the pitch angle in the vertical direction and the yaw angle in the horizontal direction:

[0074]

[0075] f, l, r represent the face image, the left-eye image, and the right-eye image input to the model respectively, represents the feature extraction module of the network, C represents the concatenation of the extracted left-eye image features, right-eye image features and face image features, and δ represents the regression of the three-dimensional line-of-sight direction using the fully connected layer.

[0076] UM-Net (no face): Since the three-dimensional line-of-sight is highly correlated with the binocular image information. To improve the speed of line-of-sight estimation, the face feature extraction branch can be removed, and only the left-eye image and the right-eye image are used as inputs. After feature extraction, the features are concatenated and the three-dimensional line-of-sight direction is regressed:

[0077]

[0078] On the other hand, the present invention believes that the facial feature extraction branch helps to provide richer information such as head pose, and removing the facial feature extraction branch will reduce the accuracy of gaze estimation. Therefore, in the experimental part, the present invention compares the gaze estimation accuracy and speed of the network before and after removing the facial feature extraction branch to evaluate the effectiveness of the facial feature extraction branch in the network of the present invention.

[0079] After the UM-Net estimates the pitch angle and yaw angle, a three-dimensional vector representing the gaze direction can be calculated As shown in formulas (4), (5), and (6):

[0080] x = cos(pitch)cos(yaw) (4)

[0081] y = cos(pitch)sin(yaw) (5)

[0082] z = sin(pitch) (6)

[0083] The angle between this vector and the true direction vector is the commonly used evaluation index in the field of three-dimensional gaze estimation, that is, the gaze angle error θ, as shown in formula (7). The loss function uses the mean square loss function (MSE Loss), as shown in formula (8):

[0084]

[0085]

[0086] 3. Mixer Layer Module

[0087] In order to achieve image feature fusion, the current deep learning models mainly act on images in three ways: fusing between different channels; fusing at different spatial positions; fusing both different channels and spaces. Different models have different acting methods. In CNNs, 1×1 convolutions are used for fusing different channels, and for fusing different spatial positions, S×S (S > 1) convolutions or pooling are used, and larger convolution kernels are used for fusing the above two kinds of features. In attention models such as VisionTransformer, self-attention layers can be used for fusing different channels and different spatial positions; while MLP can only fuse different channels. The main idea of the Mixer Layer module is to use multiple MLPs to achieve the above two kinds of feature fusions and the acting processes are separated.

[0088] The structure of the Mixer Layer module is as Figure 3As shown, the token-mixing MLP module is within the left dashed box, and the channel-mixing MLP module is within the right dashed box. The token-mixing MLP module first transposes the sequence of image feature blocks \(X\in\mathbb{R}\) 16 ×512 and then applies the MLP1 module to each column of the sequence of image feature blocks, enabling communication among different spatial positions of the sequence of image feature blocks and sharing the parameters of the MLP1 module for all columns. The obtained output is transposed again, and then in the channel-mixing MLP module, the MLP2 module is applied to each row of the sequence of image feature blocks, enabling communication among different channels of the sequence of image feature blocks and sharing the parameters of the MLP2 module for all rows.

[0089] The Mixer Layer module also uses skip-connections and layer normalization. The model can alleviate the problem of vanishing gradients using skip-connections, and layer normalization can improve the training speed and accuracy of the model, making the model more robust. For the input sequence of image feature blocks \(X\in\mathbb{R}\) 16×512 , the operation process of the Mixer Layer module can be expressed by the following formula:

[0090] U *,i = M1(LayerNorm(X) *,i ), \(i\in[1,512]\) (9)

[0091] Y j,* = M2(LayerNorm(U) j,* ), \(j\in[1,16]\) (10)

[0092] M1 and M2 represent the MLP1 module and the MLP2 module. LayerNorm(X) *,i represents the \(i\)-th column of the sequence of image feature blocks after layer normalization, and LayerNorm(U) j,* represents the \(j\)-th row of the sequence of image feature blocks after layer normalization. U *,i represents the \(i\)-th column of the sequence of image feature blocks after the action of the MLP1 module, and Y j,* represents the \(j\)-th row of the sequence of image feature blocks after the action of the MLP2 module.

[0093] 4. MLP Module (MLP1 Module or MLP2 Module)

[0094] Each MLP module in UM-Net contains two fully connected layers and a non-linear activation function (GELU), as Figure 4 shown. For the input of the MLP module The operation process can be expressed by the following formula:

[0095]

[0096] Φ represents the non-linear activation function acting on the input element, W1 and W2 represent the fully connected layers in the MLP module, and σ represents the input the output after the action of the MLP module.

[0097] Next, in order to verify that the three-dimensional line-of-sight estimation accuracy of the method of the present invention can be comparable to that of advanced three-dimensional line-of-sight estimation methods, and the three-dimensional line-of-sight estimation prediction speed is in the leading position. The specific implementation details are as follows:

[0098] 1. Dataset

[0099] MPIIFaceGaze dataset: The MPIIFaceGaze dataset uses the same batch of data as the MPIIGaze dataset, but adds full-face images. The MPIIFaceGaze dataset is a commonly used dataset in appearance-based three-dimensional line-of-sight estimation methods. The MPIIFaceGaze dataset contains 15 folders, and 15 subjects with significantly different appearances are selected. Each folder contains 3000 groups of data (including face images, left-eye images, and right-eye images) of one subject. The subjects were collected over several months in daily life, so the images have different lighting conditions and head postures.

[0100] EyeDiap dataset: Different from the MPIIFaceGaze dataset, the EyeDiap dataset was collected in a laboratory environment. The center points of the eyes and the positions of table tennis balls in the RGB video are annotated using a depth camera. These two positions are mapped to the three-dimensional point cloud recorded by the depth camera, thereby obtaining the corresponding three-dimensional position coordinates. After subtracting these two three-dimensional position coordinates, the three-dimensional line-of-sight direction is obtained.

[0101] The leave-one-subject-out method is adopted for the experiment, that is, 1 folder in the dataset is selected as the test set, and the remaining folders are used as the training set. Each folder is sequentially selected as the test set and tested separately, and the average value of the three-dimensional line-of-sight angle errors of each obtained test set is taken.

[0102] 2. Dataset preprocessing

[0103] The present invention first preprocesses the dataset, using the same image normalization method as the advanced gaze estimation method. First, the camera is virtually rotated and translated so that the virtual camera faces the reference point at a fixed distance and cancels the rolling angle of the head. The present invention sets the reference points of the MPIIFaceGaze dataset and the EyeDiap dataset to the center of the face and the centers of the two eyes, respectively. After normalizing the face images, the present invention crops the eye images from the face images, and the contrast of the eye images is adjusted by histogram equalization. The true values of the gaze angles are also normalized.

[0104] 3. Precision Comparison of Each Method

[0105] The present invention makes comparisons with the following several advanced gaze estimation methods on the MPIIFaceGaze dataset and the EyeDiap dataset.

[0106] Gaze360: A video-based gaze estimation model using bidirectional long short-term memory capsules (LSTM) is proposed, providing a method for modeling sequences where the output of one element depends on past and future inputs. In this paper, a sequence of 7 frames is used to predict the gaze of the central frame.

[0107] RT-Gene: One of the main challenges in appearance-based gaze estimation is to accurately estimate the gaze of subjects with natural appearances while allowing free movements. The RT-GENE proposed in the paper allows automatic annotation of the true gaze and head pose labels of subjects under free viewing conditions and large shot distances.

[0108] FullFace: A full-face gaze estimation model based on the attention mechanism is proposed. The main idea of the attention mechanism in this paper is to learn the weights of each position in the face region through a branch, and its goal is to increase the weights of the eye region and suppress the weights of other regions irrelevant to the gaze.

[0109] CA-Net: A coarse-to-fine gaze direction estimation model is proposed, which estimates the basic gaze direction from face images and improves it using the corresponding residuals in eye images. Under the guidance of this idea, the paper framework introduces a binary model to bridge the gaze residuals and the basic gaze direction, and introduces an attention component to adaptively obtain appropriate fine-grained features.

[0110] The method of the present invention is experimentally compared with several advanced gaze estimation methods, such as Figure 5 As shown, although UM-Net does not use CNNs but uses an MLP model aiming to improve the speed of gaze estimation, the gaze estimation accuracy of UM-Net is close to these advanced gaze estimation methods.

[0111] 4. Experimental Comparison between MLP Model and CNNs in Gaze Estimation

[0112] In the present invention, the accuracy and speed of UM-Net and CNNs-based networks in gaze estimation are compared on the MPIIFaceGaze dataset and the EyeDiap dataset. The average angular error and prediction time on each subject in the dataset are respectively shown.

[0113] For CNNs, the present invention selects Dilated-Convolutions and ResNet50. The Dilated-Net proposed by researchers such as Zhang X shows excellent performance of Dilated-Convolutions in gaze estimation. As a classic CNN structure, ResNet50 has a wide range of applications due to its powerful performance. As an experimental comparison, the present invention uses Dilated-Net to extract features of the face and both eyes. The present invention uses the ResNet50 model (ResNet50-Net) to replace the MLP model to extract 32, 256, and 256-dimensional features from face, left eye, and right eye images with a resolution of 64×64 respectively.

[0114] In terms of the average angular error, the experimental results are as Figure 6 and Figure 7 shown. The experiments show that different results are obtained when using the folders where different subjects are located as the test set. On the MPIIFaceGaze dataset, the comprehensive average angular error of UM-Net is 4.94°, that of Dilated-Net is 4.51°, and that of ResNet50-Net is 5.49°. On the EyeDiap dataset, UM-Net is 6.66°, Dilated-Net is 6.17°, and ResNet50-Net is 6.21°. The above results show that the average angular error of UM-Net on subjects with large appearance differences is comparable to that of CNNs-based networks, and it has an advantage in prediction accuracy for some subjects.

[0115] In terms of the prediction time of each selected test set, the experimental results are as Figure 8 shown. On the MPIIFaceGaze dataset, the comprehensive average prediction time of UM-Net is 3.74 seconds, that of Dilat-ed-Net is 5.23 seconds, and that of ResNet50-Net is 23.69 seconds. The average time for UM-Net to process 3000 groups of data in the MPIIFaceGaze dataset is 3.74 seconds, that is, it can process 800 groups of data per second on average, proving that the present invention can well meet the real-time requirements of gaze estimation.

[0116] As Figure 9As shown in the figure, on the EyeDiap dataset, the comprehensive average prediction time of UM-Net is 6.91 seconds, that of Dilated-Net is 11.12 seconds, and that of ResNet50-Net is 47.52 seconds. The experimental results show that the prediction time of UM-Net on the test set of any subject with significant appearance differences is significantly better than that of Dilated-Net and much better than that of ResNet50-Net. The above experiments indicate that UM-Net has a fast gaze estimation speed and has good prospects in application scenarios with strong real-time performance of 3D gaze estimation.

[0117] In summary, the experiments show that UM-Net uses an MLP model to extract image features. In the field of gaze estimation, the prediction accuracy can be comparable to that of CNNs-based networks, and the prediction speed is in the leading position.

[0118] 5. Verification of the effectiveness of the face feature extraction branch

[0119] UM-Net uses a branch to extract 32-dimensional features from face images to supplement more abundant information. To verify the effectiveness of the face feature extraction branch, this paper removes the face feature extraction branch and retains the remaining two eye feature extraction branches for gaze estimation. The experimental results are shown in Table 1. The average gaze angle error after removing the face feature extraction branch is 5.93°, higher than 4.94° of UM-Net, and the average prediction time is 3.13 seconds, with a small difference from 3.74 seconds of UM-Net. Therefore, adding a face feature extraction branch to extract 32-dimensional features from face images can supplement feature information other than binocular images, significantly improve the gaze estimation accuracy, but have little impact on the prediction time, verifying the effectiveness of the face feature extraction branch. On the other hand, this experiment also shows that in scenarios where the gaze estimation speed is pursued, the face feature extraction branch in UM-Net can be removed.

[0120] Table 1 Comparison of experimental results of UM-Net network after removing the face feature extraction branch

[0121]

[0122] Next, a 3D gaze estimation device based on MLP provided by the present invention will be described. The 3D gaze estimation device described below can be mutually referred to with the 3D gaze estimation method described above.

[0123] As Figure 10 shown, a 3D gaze estimation device based on MLP includes the following modules:

[0124] Building block 001 is used to build an MLP-based UM-Net network. The UM-Net network includes three branches, namely the left-eye feature extraction branch, the right-eye feature extraction branch, and the face feature extraction branch; and a feature splicing module connected to all three branches, a fully connected layer FC1 and a fully connected layer FC2 connected to the feature splicing module in sequence;

[0125] Preprocessing module 002 is used to obtain a dataset to be measured, including left-eye images, right-eye images, and face images, and perform preprocessing on them respectively;

[0126] Estimation module 003 is used to input the preprocessed images into the UM-Net network; extract left-eye image features, right-eye image features, and face image features through the three branches respectively. After splicing through the feature splicing module, perform feature dimensionality reduction through the fully connected layer FC1, and regress the three-dimensional line-of-sight direction through the fully connected layer FC2.

[0127] Based on but not limited to the above device, the feature extraction branch includes a feature extraction module, N MixerLayer modules, a global average pooling layer GAP, and a fully connected layer FC connected in sequence;

[0128] First, the feature extraction module splits the input image into image patches; then projects each image patch into a 512-dimensional space through a fully connected layer, and obtains a sequence of image feature patches after projection;

[0129] Then, send the sequence of image feature patches into N MixerLayer modules to perform feature refinement in the column direction and feature refinement in the row direction on the sequence of image feature patches, and repeatedly pass the sequence of image feature patches through N MixerLayer modules to refine the image feature information;

[0130] Next, the global average pooling layer GAP regularizes the entire network model in structure to prevent overfitting;

[0131] Finally, use the fully connected layer FC to regress the required image features respectively.

[0132] Based on but not limited to the above device, the Mixer Layer module includes a token-mixing MLP module and a channel-mixing MLP module;

[0133] The token-mixing MLP module and the channel-mixing MLP module are alternately stacked to perform feature refinement in the column direction and feature refinement in the row direction on the sequence of image feature patches.

[0134] Based on but not limited to the above device, the token-mixing MLP module includes the MLP1 module, and the channel-mixing MLP module includes the MLP2 module;

[0135] The token-mixing MLP module first transposes the image feature block sequence X ∈ R 16×512 After that, the MLP1 module is applied to each column of the image feature block sequence to enable communication between different spatial positions of the image feature block sequence, and all columns share the parameters of the MLP1 module. The obtained output is transposed again, and then in the channel-mixing MLP module, the MLP2 module is applied to each row of the image feature block sequence to enable communication between different channels of the image feature block sequence, and all rows share the parameters of the MLP2 module;

[0136] Furthermore, skip connections and layer normalization are also used in the MixerLayer module.

[0137] Based on but not limited to the above device, for the input image feature block sequence X ∈ R 16×512 , the operation process of the Mixer Layer module is expressed by the following formula:

[0138] U *,i = M1(LayerNorm(X) *,i ), i ∈ [1, 512]

[0139] Y j,* = M2(LayerNorm(U) j,* ), j ∈ [1, 16]

[0140] M1 and M2 represent the MLP1 module and the MLP2 module, LayerNorm(X) *,i represents the i-th column of the image feature block sequence after layer normalization, LayerNorm(U) j,* represents the j-th row of the image feature block sequence after layer normalization, U *,i represents the i-th column of the image feature block sequence after the action of the MLP1 module, Y j,* represents the j-th row of the image feature block sequence after the action of the MLP2 module.

[0141] Based on but not limited to the above device, each MLP module (MLP1 module and MLP2 module) includes two fully connected layers and a non-linear activation function; for the input of the MLP module The operation process is expressed by the following formula:

[0142]

[0143] Φ represents the non - linear activation function acting on the input element, W1 and W2 represent the fully - connected layers in the MLP module, and σ represents the input The output after the action of the MLP1 module or the MLP2 module.

[0144] Based on but not limited to the above device, the three - dimensional line - of - sight direction is represented by the pitch angle in the vertical direction and the yaw angle in the horizontal direction:

[0145]

[0146] f, l, r respectively represent the face image, left - eye image, and right - eye image input to the model, represents the feature extraction module of the network, C represents the connection of the extracted left - eye image features, right - eye image features, and face image features, and δ represents using the fully - connected layer to regress the three - dimensional line - of - sight direction;

[0147] Calculate the three - dimensional vector representing the line - of - sight direction according to the pitch angle and the yaw angle The calculation formula is as follows:

[0148] x = cos(pitch)cos(yaw)

[0149] y = cos(pitch)sin(yaw)

[0150] z = sin(pitch)

[0151] Three - dimensional vector And the true direction vector The included angle between them is the evaluation index of three - dimensional line - of - sight estimation, that is, the line - of - sight angle error θ, and the loss function adopts the mean - square loss function MSE. The calculation formulas are as follows respectively:

[0152]

[0153]

[0154] Such as Figure 11As shown in the figure, a schematic diagram of the physical structure of an electronic device is exemplified. The electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the steps of the above three-dimensional line-of-sight estimation method, specifically including: constructing a UM-Net network based on MLP. The UM-Net network includes three branches, namely, a left-eye feature extraction branch, a right-eye feature extraction branch, and a face feature extraction branch; and a feature splicing module connected to all three branches, a fully connected layer FC1 and a fully connected layer FC2 connected to the feature splicing module in sequence; obtaining a dataset to be measured, including a left-eye image, a right-eye image, and a face image, and performing preprocessing on them respectively; inputting the preprocessed images into the UM-Net network; extracting left-eye image features, right-eye image features, and face image features through the three branches respectively, splicing them through the feature splicing module, performing feature dimensionality reduction through the fully connected layer FC1, and regressing the three-dimensional line-of-sight direction through the fully connected layer FC2.

[0155] In addition, when the logical instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0156] In another aspect, an embodiment of the present invention further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned three-dimensional gaze estimation method are implemented, specifically including: constructing a UM-Net network based on MLP, where the UM-Net network includes three branches, namely a left-eye feature extraction branch, a right-eye feature extraction branch, and a face feature extraction branch; and a feature splicing module connected to all three branches, a fully connected layer FC1 and a fully connected layer FC2 connected to the feature splicing module in sequence; obtaining a dataset to be measured, including left-eye images, right-eye images, and face images, and performing preprocessing on each of them respectively; inputting the preprocessed images into the UM-Net network; extracting left-eye image features, right-eye image features, and face image features through the three branches respectively, splicing them through the feature splicing module, performing feature dimensionality reduction through the fully connected layer FC1, and regressing the three-dimensional gaze direction through the fully connected layer FC2.

[0157] It should be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or system including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or system. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or system including that element.

[0158] The serial numbers of the above-mentioned embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments. Among the several unit claims listing a number of devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order, and these words can be interpreted as identifiers.

[0159] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A three-dimensional line-of-sight estimation method based on MLP, characterized in that, It includes the following steps: Construct a U-M-Net network based on MLP. The U-M-Net network includes three branches, namely the left-eye feature extraction branch, the right-eye feature extraction branch, and the face feature extraction branch; and a feature splicing module connected to all three branches, a fully connected layer FC1 and a fully connected layer FC2 connected to the feature splicing module in sequence; Obtain the dataset to be measured, including left-eye images, right-eye images, and face images, and perform preprocessing on them respectively; Input the preprocessed images into the U-M-Net network; extract the left-eye image features, right-eye image features, and face image features through the three branches respectively. After splicing through the feature splicing module, perform feature dimensionality reduction through the fully connected layer FC1, and regress the three-dimensional line-of-sight direction through the fully connected layer FC2; Among them, the feature extraction branch includes a feature extraction module, N Mixer Layer modules, a global average pooling layer GAP, and a fully connected layer FC connected in sequence; First, the feature extraction module splits the input image into image patches; then projects each image patch into a 512-dimensional space through a fully connected layer, and an image feature patch sequence is obtained after projection; Then send the image feature patch sequence into N Mixer Layer modules to perform feature refinement in the column direction and feature refinement in the row direction on the image feature patch sequence, and repeatedly pass the image feature patch sequence through N Mixer Layer modules to refine the image feature information; Next, the global average pooling layer GAP regularizes the entire network model structurally to prevent overfitting; Finally, use the fully connected layer FC to regress the required image features respectively; The Mixer Layer module includes a token-mixing MLP module and a channel-mixing MLP module; The token-mixing MLP module and the channel-mixing MLP module are alternately stacked to perform feature refinement in the column direction and feature refinement in the row direction on the image feature patch sequence.

2. The three-dimensional line-of-sight estimation method according to claim 1, characterized in that The token-mixing MLP module contains an MLP1 module, and the channel-mixing MLP module contains an MLP2 module; The token-mixing MLP module first transposes the sequence of image feature blocks \(X\in\mathbb{R}\) 16×512 After that, the MLP1 module is applied to each column of the sequence of image feature blocks to enable communication between different spatial positions of the sequence of image feature blocks, and the parameters of the MLP1 module are shared among all columns. The resulting output is transposed again, and then in the channel-mixing MLP module, the MLP2 module is applied to each row of the sequence of image feature blocks to enable communication between different channels of the sequence of image feature blocks, and the parameters of the MLP2 module are shared among all rows; The Mixer Layer module also uses skip connections and layer normalization.

3. The three-dimensional line-of-sight estimation method according to claim 2, characterized in that For the input sequence of image feature blocks \(X\in\mathbb{R}\) 16×512 , the operation process of the Mixer Layer module is expressed by the following formula: U *,i = M1(LayerNorm(X) *,i ), i ∈ [1, 512] Y j,* = M2(LayerNorm(U) j,* ), j ∈ [1, 16] M1 and M2 represent the MLP1 module and the MLP2 module, LayerNorm(X) *,i represents the i-th column of the sequence of image feature blocks after layer normalization, LayerNorm(U) j,* represents the j-th row of the sequence of image feature blocks after layer normalization, U *,i represents the i-th column of the sequence of image feature blocks after the action of the MLP1 module, Y j,* represents the j-th row of the sequence of image feature blocks after the action of the MLP2 module.

4. The three-dimensional line-of-sight estimation method according to claim 2, wherein Each MLP1 module or MLP2 module contains two fully-connected layers and a non-linear activation function; for the input of the MLP1 module or MLP2 module The working process is expressed by the following formula: Φ represents a non - linear activation function acting on the input element, W1 and W2 represent two fully - connected layers in the MLP1 module or the MLP2 module, and σ represents the output after the action of the MLP1 module or the MLP2 module. The output after passing through the MLP1 module or the MLP2 module.

5. The three-dimensional line-of-sight estimation method according to claim 1, characterized in that The three-dimensional line-of-sight direction is represented by the pitch angle in the vertical direction and the yaw angle in the horizontal direction: f, l, and r respectively represent the input face image, left-eye image, and right-eye image of the model. represents the feature extraction module of the network. C represents the connection of the extracted left-eye image features, right-eye image features, and face image features. δ represents the regression of the three-dimensional line-of-sight direction using a fully connected layer. Calculate a three-dimensional vector representing the line-of-sight direction based on the pitch angle and yaw angle The calculation formula is as follows: x = cos(pitch)cos(yaw) y = cos(pitch)sin(yaw) z = sin(pitch) Three-dimensional vector and the true direction vector The included angle between them is the evaluation index of three-dimensional gaze estimation, that is, the gaze angle error θ. The loss function uses the mean square error function MSE. The total number of predicted three-dimensional gaze vectors is n, and the calculation formulas are as follows:

6. A three-dimensional line-of-sight estimation device based on MLP, characterized in that It is used to implement the three-dimensional line-of-sight estimation method according to any one of claims 1-5.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the three-dimensional line-of-sight estimation method according to any one of claims 1-5.

8. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the steps of the three-dimensional line-of-sight estimation method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Quick vehicle detection method based on improved YOLOv3

    CN111079584A

  • Training data generation model construction method for aircraft fault diagnosis and application

    CN114169396A