A Method and System for Estimating Target Gaze Direction Based on Multi-Task Learning
Through a network model based on multitask learning, combined with feature extraction and enhancement modules, the line of sight estimation error problem when facial posture and gaze direction are large, and a more accurate and efficient gaze direction estimation is achieved.
Patent Information
- Application Number
- CN202311134202.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-04
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-09-04
AI Technical Summary
In the prior art, when the facial posture and gaze direction are different, the line of sight estimation is prone to failure or error, and the gaze direction estimation based on the target human eye characteristics requires better spot and pupil imaging, and when the line of sight angle is larger, the spot and imaging are poor.
Using the target gaze direction estimation method based on multitask learning, the network model including feature extraction, feature enhancement, facial pose estimation and gaze direction estimation modules is achieved while estimating facial pose and gaze direction and improving the accuracy of estimation.
It effectively reduces the line of sight estimation error when facial posture and gaze direction are different, improves the accuracy and calculation efficiency of gaze direction estimation, and solves the spot and imaging problems of traditional methods when gaze angle is large.
Smart Images

Figure CN117218630B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target gaze direction estimation, and in particular, to a method and system for target gaze direction estimation based on multi-task learning. Background Art
[0002] Patent CN116563827A discloses a method, system and electronic device for gaze estimation, specifically discloses obtaining at least two images at a driving position collected by at least two image acquisition devices located at different positions in a vehicle; obtaining a depth image at the driving position based on the at least two images; extracting face key point information at the driving position in the at least two images; determining face pose information at the driving position based on the depth image and the face key point information; determining the gaze angle at the driving position based on the face pose information at the driving position.
[0003] Patent CN116311486A discloses a method, device, equipment and medium for gaze estimation, specifically discloses obtaining a left eye image and a right eye image when gaze estimation is required. Inputting the left eye image and the right eye image into a pre-trained gaze estimation model to obtain the monocular gaze estimation results output by the gaze estimation model, that is, the left eye gaze angle and the right eye gaze angle, and the binocular gaze estimation result, that is, the binocular combined gaze angle.
[0004] However, the above existing methods have the following defects:
[0005] (1) The gaze direction estimation based on the target eye features requires obtaining good spot and pupil imaging. When the angle of the line of sight is large, the spot and imaging are poor, resulting in gaze estimation failure or large error.
[0006] (2) The gaze estimation result based on facial features is easily affected by the facial pose result. For scenarios where the facial pose and gaze direction have large differences, the method of estimating the gaze direction based on the target facial features will cause a large error in gaze direction estimation. Summary of the Invention
[0007] Based on the technical problems existing in the background art, the present invention proposes a method and system for target gaze direction estimation based on multi-task learning, which solves the problems of gaze estimation failure and large error when the facial pose and gaze direction have large differences in the actual scenario.
[0008] A method for target gaze direction estimation based on multi-task learning proposed by the present invention includes the following steps:
[0009] Input the image to be estimated into a trained network model to output the gaze estimation of the target gaze, and the network model includes a feature extraction module, a feature enhancement module, a facial pose estimation module and a gaze direction estimation module;
[0010] The training process of the network model is as follows:
[0011] S1: Obtain a training set, where each face in the training set is annotated with the facial pose and gaze direction, and calculate the class truth value class and offset truth value offset of the face participating in the training. The facial pose includes the facial pitch angle and the facial yaw angle, the gaze direction includes the gaze pitch angle and the gaze yaw angle, the class truth value class includes the facial pose class truth value face_class_gt and the gaze direction class truth value gaze_class_gt, and the offset truth value offset includes the facial pose offset value truth value face_offset_gt and the gaze direction offset value truth value gaze_offset_gt;
[0012] S2: Extract the face features in the training set based on the feature extraction module;
[0013] S3: Perform feature enhancement on the face features based on the feature enhancement module to obtain the first enhanced face feature of the facial pose and the second enhanced face feature of the gaze direction;
[0014] S4: Perform non-linear layer embedding1 on the first enhanced face feature based on the facial pose estimation module to obtain the third face feature, and obtain the target facial pose based on the third face feature;
[0015] S5: Perform non-linear layer embedding2 on the first enhanced face feature based on the gaze direction estimation module to obtain the fourth face feature, and obtain the target gaze direction based on the fourth face feature.
[0016] Furthermore, the loss function of the network model is as follows:
[0017] loss = l face_class + l face_offset + w gaze1 * l gaze_class + w gaze2 * l gaze_offset
[0018]
[0019]
[0020] Among them, l face_class represents the facial pose classification loss function, l face_offset represents the facial pose offset value loss function, l gaze_class represents the gaze direction classification loss function, l gaze_offset represents the gaze direction offset loss function, and α and β represent hyperparameters.
[0021] Facial pose classification loss function l face_class and gaze direction classification loss function l gaze_class Adopt the cross - entropy loss function, facial pose offset value loss function l face_offset and gaze direction offset loss function l gaze_offset Adopt the mean square error loss function.
[0022] Furthermore, the value range of the offset ground truth offset is all [0, 1]. The calculation formulas of the class ground truth class and the offset ground truth offset for the facial pitch angle, facial yaw angle, gaze pitch angle, and gaze yaw angle are similar. The specific calculation formulas are as follows:
[0023]
[0024]
[0025] Among them, [a, b] is the value range of the pitch angle and yaw angle of the human face facial pose and gaze direction. n represents that the angle range interval of b - a degrees is evenly divided into n parts, and m represents the standard value of the facial pitch angle or the facial yaw angle or the gaze pitch angle or the gaze yaw angle.
[0026] Furthermore, the feature extraction module uses ResNet50 pre - trained on the large - scale image classification dataset ImageNet as the backbone network to extract face features.
[0027] Furthermore, in the process of feature enhancement of the face features based on the feature enhancement module to obtain the enhanced face feature one of the facial pose and the enhanced face feature two of the gaze direction, it specifically includes:
[0028] Use the multi - head self - attention residual structure to perform self - enhancement of the facial pose features on the obtained face features to obtain the enhanced face feature one of the facial pose;
[0029] Use the multi - head self - attention residual structure to perform self - enhancement of the gaze direction features on the obtained face features to obtain the self - enhanced feature of the gaze direction;
[0030] Use the multi - head mutual - attention residual structure to perform mutual enhancement of the features corresponding to the atrium part of the face feature one on the self - enhanced feature of the gaze direction to obtain the enhanced face feature two of the gaze direction.
[0031] Furthermore, the facial pose estimation module designs two facial pose output heads. The first facial pose output head is respectively used to output the pitch angle class value face_pitch_class1 and the yaw angle class value face_yaw_class1 of the facial pose. The second facial pose output head is used to output the pitch angle offset value face_pitch_offset1 and the yaw angle offset value face_yaw_offset1 of the facial pose. The pitch angle face_pitch of the facial pose is calculated based on the pitch angle class value face_pitch_class1 and the pitch angle offset value face_pitch_offset1 of the facial pose, and the yaw angle face_yaw is calculated based on the yaw angle class value face_yaw_class1 and the yaw angle offset value face_yaw_offset1.
[0032] The gaze direction estimation module designs two gaze output heads. The first gaze output head is respectively used to output the pitch angle class value gaze_pitch_class2 and the yaw angle class value gaze_yaw_class2 of the gaze direction. The second gaze output head is used to output the pitch angle offset value gaze_pitch_offset2 and the yaw angle offset value gaze_yaw_offset2 of the gaze direction. The pitch angle gaze_pitch of the gaze direction is calculated based on the pitch angle class value gaze_pitch_class2 and the pitch angle offset value gaze_pitch_offset2, and the yaw angle gaze_yaw of the gaze direction is calculated based on the yaw angle class value gaze_yaw_class2 and the yaw angle offset value gaze_yaw_offset2.
[0033] Furthermore, the calculation formulas for the pitch angle face_pitch and the yaw angle face_yaw of the facial pose are as follows:
[0034]
[0035]
[0036] The calculation formulas for the pitch angle gaze_pitch and the yaw angle gaze_yaw of the gaze direction are as follows:
[0037]
[0038]
[0039] Among them, [a, b] is the value range of the pitch angle and the yaw angle of the human face facial pose and the gaze direction, n represents that the angle range interval of b - a degrees is evenly divided into n parts, and * represents multiplication.
[0040] A target gaze direction estimation system based on multi-task learning inputs an image to be estimated into a trained network model to output the line-of-sight estimation of the target gaze. The network model includes a feature extraction module, a feature enhancement module, a facial pose estimation module, and a gaze direction estimation module;
[0041] The training process of the network model is as follows:
[0042] Obtain a training set. Each face in the training set is labeled with the facial pose and gaze direction of the face, and the class truth value class and offset truth value offset of the face participating in the training are calculated. The facial pose includes the facial pitch angle and the facial yaw angle, and the gaze direction includes the gaze pitch angle and the gaze yaw angle. The class truth value class includes the facial pose class truth value face_class_gt and the gaze direction class truth value gaze_class_gt. The offset truth value offset includes the facial pose offset truth value face_offset_gt and the gaze direction offset truth value gaze_offset_gt;
[0043] Extract the face features in the training set based on the feature extraction module;
[0044] Perform feature enhancement on the face features based on the feature enhancement module to obtain the first enhanced face feature of the facial pose and the second enhanced face feature of the gaze direction;
[0045] Perform nonlinear layer embedding1 on the first enhanced face feature based on the facial pose estimation module to obtain the third face feature, and obtain the target facial pose based on the third face feature;
[0046] Perform nonlinear layer embedding2 on the first enhanced face feature based on the gaze direction estimation module to obtain the fourth face feature, and obtain the target gaze direction based on the fourth face feature.
[0047] A computer-readable storage medium stores a number of programs thereon, and the number of programs is used to be called by a processor and execute the target gaze direction estimation method as described above.
[0048] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0049] The advantages of a target gaze direction estimation method and system based on multi-task learning provided by the present invention are as follows: In the structure of the present invention, a target gaze direction estimation method and system based on multi-task learning, through training the network model, enables the target gaze direction estimation method based on multi-task learning of this network model to change the traditional target gaze regression task into a classification + regression task, which can greatly reduce the estimation error of the target gaze direction; it can simultaneously obtain the target facial pose and the target gaze direction, effectively reducing the computational amount. Additionally, the feature enhancement module can simultaneously improve the accuracy of target facial pose and gaze direction estimation; the mutual enhancement unit of facial pose features to gaze direction features can achieve feature fusion with a relatively small computational amount; this embodiment solves the problems of gaze estimation failure and large errors when there are significant differences in facial pose and gaze direction in the actual scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of the present invention;
[0051] Figure 2 is a flowchart of the feature enhancement module;
[0052] Figure 3 is a flowchart of the multi-head mutual attention residual structure in the mutual enhancement unit of facial pose features to gaze direction features. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] Next, the technical solutions of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0054] As Figures 1 to 3 shown, a target gaze direction estimation method based on multi-task learning proposed by the present invention includes the following steps:
[0055] Input the image to be estimated into the trained network model to output the gaze estimation of the target. The network model includes a feature extraction module, a feature enhancement module, a facial pose estimation module, and a gaze direction estimation module;
[0056] The training process of the network model is as follows:
[0057] S1: Obtain a training set, where each face in the training set is labeled with the face facial pose and the gaze direction, and calculate the class ground truth class and the offset ground truth offset for the face to participate in training. The face facial pose includes the facial pitch angle and the facial yaw angle, and the gaze direction includes the gaze pitch angle and the gaze yaw angle. The class ground truth class includes the facial pose class ground truth face-class-gt and the class ground truth gaze-class_gt of the gaze direction. The offset ground truth offset includes the facial pose offset value ground truth face_offset_gt and the gaze direction offset value ground truth gaze_offset_gt;
[0058] In the actual scenario, the value ranges of the pitch angle and the yaw angle of the facial pose and the gaze direction are [a, b]. Therefore, the angle range interval is b - a degrees, and the b - a degrees are evenly divided into n parts. The calculation formulas for the class ground truth class and the offset ground truth offset of the facial pitch angle, the facial yaw angle, the gaze pitch angle, and the gaze yaw angle are similar. Taking the labeled value m of the facial pitch angle as an example, use the following formula to calculate its class label ground truth class and offset ground truth offset for participating in training (the value range of the offset ground truth offset is [0, 1])
[0059]
[0060]
[0061] It can be understood that the facial pitch angle corresponds to two ground truths. Similarly, the facial yaw angle, the gaze pitch angle, and the gaze yaw angle each correspond to two ground truths. That is to say, the facial pitch angle, the facial yaw angle, the gaze pitch angle, and the gaze yaw angle, these four angles correspond to eight ground truths. The calculation formulas for the four angles are similar, except that m represents the standard value of the facial pitch angle or the standard value of the facial yaw angle or the standard value of the gaze pitch angle or the standard value of the gaze yaw angle, and the specific value is taken during calculation.
[0062] S2: Extract the face features in the training set based on the feature extraction module;
[0063] The feature extraction module uses ResNet50 pre-trained on the large-scale image classification dataset ImageNet as the backbone network to extract face features.
[0064] S3: Perform feature enhancement on the face features based on the feature enhancement module to obtain the enhanced face feature one of the face facial pose and the enhanced face feature two of the gaze direction;
[0065] (a) The feature enhancement module includes a facial pose feature self-enhancement unit, a gaze direction feature self-enhancement unit, and a mutual enhancement unit of the facial pose feature to the gaze direction feature;
[0066] (a1) The facial pose feature self-enhancement unit uses a multi-head self-attention residual structure to perform self-enhancement of facial pose features on the obtained face features, and obtains the first face features after facial pose enhancement of the face;
[0067] (a2) The gaze direction feature self-enhancement unit uses a multi-head self-attention residual structure to perform self-enhancement of gaze direction features on the obtained face features, and obtains the self-enhanced features of gaze direction features;
[0068] (a3) The mutual enhancement unit of facial pose features to gaze direction features uses a multi-head mutual-attention residual structure to mutually enhance the features corresponding to the middle atrium part of the first face features to the self-enhanced features of gaze direction features, and obtains the second face features after gaze direction enhancement.
[0069] After obtaining the first face features and the self-enhanced features of gaze direction features, since the gaze direction of a person in the actual scene is easily affected by the facial pose, the mutual enhancement unit of facial pose features to gaze direction features mutually enhances the first face features to the self-enhanced features of gaze direction features. In order to reduce the computational complexity of the mutual enhancement unit, in this embodiment, only the features corresponding to the middle atrium part (the eye region) of the enhanced first face features are used to mutually enhance the enhanced gaze features, and the enhanced second face features are obtained, and feature fusion is performed with a smaller computational complexity to reduce the gaze direction estimation error.
[0070] The process of obtaining the second face features is as follows: The middle atrium part of the face in the original face image is preset as a rectangular area, and the rectangular position is [x, y, w, h], where (x, y) represents the coordinates of the upper left vertex of the rectangle, w represents the width of the rectangle, and h represents the height of the rectangle. The dimension n*d of the first face features is transformed into w0*h0*d, where n = w0*h0, w0 represents the width of the feature map, h0 represents the height of the feature map, and d represents the channels of the feature map. The present invention uses the ROI Align method in mask-rcnn to obtain the features of the middle atrium area on the first face features, and the dimension is w1*h1*d. This method can solve the quantization error problem brought by the ROI Pooling operation. The dimension of the features of the middle atrium area is transformed into n1*d, where n1 = w1*h1. The features of the middle atrium area are respectively embedded once to obtain the transformed features 1 and features 2 of the middle atrium area. In the present invention, the self-enhanced features of gaze direction features are used as the query, the features 1 of the middle atrium area are used as the key, and the features 2 of the middle atrium are used as the value, and the query, key, and value are sent into Figure 3 the multi-head mutual-attention residual structure shown in the figure to obtain the face features after gaze direction enhancement
[0071] S4: Based on the face pose estimation module, perform non-linear layer embedding1 on the first face feature to obtain the third face feature, and obtain the target face pose based on the third face feature;
[0072] The face pose estimation module designs two face pose output heads. The first face pose output head is respectively used to output the pitch angle class value face_pitch_class1 and the yaw angle class value face_yaw_class1 of the face pose. The second face pose output head is used to output the pitch angle offset value face_pitch_offset1 and the yaw angle offset value face_yaw_offset1 of the face pose. Calculate the pitch angle face_pitch of the face pose based on the pitch angle class value face_pitch_class1 and the pitch angle offset value face_pitch_offset1 of the face pose, and calculate the yaw angle face_yaw based on the yaw angle class value face_yaw_class1 and the yaw angle offset value face_yaw_offset1.
[0073] During actual inference, the calculation formulas for the pitch angle face_pitch and the yaw angle face_yaw of the face pose are as follows:
[0074]
[0075]
[0076] Among them, * represents multiplication.
[0077] S5: Based on the gaze direction estimation module, perform non-linear layer embedding2 on the first face feature to obtain the fourth face feature, and obtain the target gaze direction based on the fourth face feature.
[0078] The gaze direction estimation module designs two gaze output heads. The first gaze output head is respectively used to output the pitch angle class value gaze_pitch_class2 and the yaw angle class value gaze_yaw_class2 of the gaze direction. The second gaze output head is used to output the pitch angle offset value gaze_pitch_offset2 and the yaw angle offset value gaze_yaw_offset2 of the gaze direction. Calculate the pitch angle gaze_pitch of the gaze direction based on the pitch angle class value gaze_pitch_class2 and the pitch angle offset value gaze_pitch_offset2, and calculate the yaw angle gaze_yaw of the gaze direction based on the yaw angle class value gaze_yaw_class2 and the yaw angle offset value gaze_yaw_offset2.
[0079] During actual inference, the calculation formulas for the pitch angle gaze_pitch and yaw angle gaze_yaw of the gaze direction are as follows:
[0080]
[0081]
[0082] where * represents multiplication.
[0083] By training the network model through steps S1 to S5, the target gaze direction estimation method based on this network model for multi-task learning changes the traditional target gaze regression task into a classification + regression task, which can greatly reduce the estimation error of the target gaze direction; it can simultaneously obtain the target face pose and target gaze direction, effectively reducing the computational amount. In addition, the feature enhancement module can improve the accuracy of target face pose and gaze direction estimation at the same time; the mutual enhancement unit of the face pose feature and the gaze direction feature can achieve feature fusion with a relatively small computational amount; this embodiment solves the problems of gaze estimation failure and large error when there are large differences in face pose and gaze direction in the actual scenario.
[0084] The loss function of the network model in this embodiment is as follows:
[0085] loss = l face_class + l face_offset + w gaze1 * l gaze_class + w gaze2 * l ga z e_offset
[0086]
[0087]
[0088] where, l f a ce_class represents the face pose classification loss function, l face_offset represents the face pose offset value loss function, l gaze_class represents the gaze direction classification loss function, l gaze_offset represents the gaze direction offset loss function, and α and β represent hyperparameters. The face pose classification loss function l face_class and the gaze direction classification loss function l gaze_class adopt the cross-entropy loss function, and the face pose offset value loss function l face_offset and the gaze direction offset loss function l gaze_offset adopt the mean square error loss function.
[0089] This loss function can effectively reduce the problem of large gaze direction estimation error caused by inconsistent face pose and gaze direction.
[0090] As described above, it is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.
Claims
1. A method for estimating the target gaze direction based on multi-task learning, characterized in that, It includes the following steps: Input the image to be estimated into the trained network model to output the line-of-sight estimation of the target gaze. The network model includes a feature extraction module, a feature enhancement module, a facial pose estimation module, and a gaze direction estimation module; The training process of the network model is as follows: S1: Obtain a training set, where each face in the training set is labeled with the face facial pose and the gaze direction, and calculate the class ground truth of the face participating in the training and the offset ground truth , the face facial pose includes the facial pitch angle and the facial yaw angle, the gaze direction includes the gaze pitch angle and the gaze yaw angle, and the class ground truth includes the facial pose class ground truth , the class ground truth of the gaze direction , and the offset ground truth includes the facial pose offset value ground truth and the gaze direction offset value ground truth ; S2: Extract the face features in the training set based on the feature extraction module; S3: Use the multi-head self-attention residual structure to perform self-enhancement of facial pose features on the obtained face features to obtain the first enhanced face features of the face facial pose; use the multi-head self-attention residual structure to perform self-enhancement of gaze direction features on the obtained face features to obtain the self-enhanced features of the gaze direction features; use the multi-head mutual attention residual structure to perform mutual enhancement of the features corresponding to the middle atrium part of the first face features on the self-enhanced features of the gaze direction features to obtain the second enhanced face features of the gaze direction; S4: Perform nonlinear layer embedding1 on the first face features based on the facial pose estimation module to obtain the third face features, and obtain the target facial pose based on the third face features; S5: Perform nonlinear layer embedding2 on the second face features based on the gaze direction estimation module to obtain the fourth face features, and obtain the target gaze direction based on the fourth face features.
2. The method for estimating the target gaze direction based on multi-task learning according to claim 1, wherein The loss function of the network model is as follows: Among them, represents the facial pose classification loss function, represents the facial pose offset value loss function, represents the gaze direction classification loss function, represents the gaze direction offset loss function, and represents the hyperparameter; Facial pose classification loss function and gaze direction classification loss function Adopt cross-entropy loss function, facial pose offset value loss function and gaze direction offset loss function Adopt mean square error loss function.
3. The method for estimating the target gaze direction based on multi-task learning according to claim 1, wherein The offset true value both have a value range of [0, 1]. The category true values of the facial pitch angle, facial yaw angle, gaze pitch angle, and gaze yaw angle and the offset true value have similar calculation formulas. The specific calculation formulas are as follows: Among them, is the value range of the pitch angle and yaw angle of the face facial pose and gaze direction, represents the angle range interval degrees are evenly divided into parts, represents the standard value of the face pitch angle or the standard value of the face yaw angle or the standard value of the gaze pitch angle or the standard value of the gaze yaw angle.
4. The method for estimating the target gaze direction based on multi-task learning according to claim 1, wherein The feature extraction module uses ResNet50 pre-trained on the large-scale image classification dataset ImageNet as the backbone network to extract face features.
5. The method for estimating the target gaze direction based on multi-task learning according to claim 1, wherein The facial pose estimation module designs two facial pose output heads. The first facial pose output head is respectively used to output the pitch angle category value of the facial pose and the yaw angle category value . The second facial pose output head is used to output the pitch angle offset value of the facial pose and the yaw angle offset value . Based on the pitch angle category value of the facial pose and the pitch angle offset value of the facial pose , the pitch angle of the facial pose is calculated . Based on the yaw angle category value and the yaw angle offset value , the yaw angle is calculated ; The gaze direction estimation module designs two gaze output heads. The first gaze output head is respectively used to output the pitch angle category value of the gaze direction and the yaw angle category value . The second gaze output head is used to output the pitch angle offset value of the gaze direction and the yaw angle offset value . Based on the pitch angle category value and the pitch angle offset value , the pitch angle of the gaze direction is calculated. Based on the yaw angle category value and the yaw angle offset value , the yaw angle of the gaze direction is calculated.
6. The method for estimating the target gaze direction based on multi-task learning according to claim 5, characterized in that, Pitch angle of facial pose and yaw angle The calculation formula is as follows: Pitch angle of the gaze direction and yaw angle The calculation formulas are as follows: Among them, is the value range of the pitch angle and yaw angle of the face facial pose and gaze direction, represents the angle range interval degrees are evenly divided into parts, represents the product.
7. A target gaze direction estimation system based on multi-task learning, characterized in that, Input the image to be estimated into the trained network model to output the line-of-sight estimation of the target gaze. The network model includes a feature extraction module, a feature enhancement module, a facial pose estimation module, and a gaze direction estimation module; The training process of the network model is as follows: Obtain a training set, where each face in the training set is annotated with the face pose and the gaze direction, and calculate the class ground truth of the face participating in the training and the offset ground truth , the face pose includes the face pitch angle and the face yaw angle, the gaze direction includes the gaze pitch angle and the gaze yaw angle, and the class ground truth includes the face pose class ground truth , the class ground truth of the gaze direction , and the offset ground truth includes the face pose offset value ground truth and the gaze direction offset value ground truth ; Extract the face features in the training set based on the feature extraction module; Based on the feature enhancement module, use the multi-head self-attention residual structure to perform self-enhancement of facial pose features on the obtained face features to obtain the first enhanced face features of the face facial pose; use the multi-head self-attention residual structure to perform self-enhancement of gaze direction features on the obtained face features to obtain the self-enhanced features of the gaze direction features; use the multi-head mutual attention residual structure to perform mutual enhancement of the features corresponding to the middle atrium part of the first face features on the self-enhanced features of the gaze direction features to obtain the second enhanced face features of the gaze direction; Perform nonlinear layer embedding1 on the first face features based on the facial pose estimation module to obtain the third face features, and obtain the target facial pose based on the third face features; Perform nonlinear layer embedding2 on the second face features based on the gaze direction estimation module to obtain the fourth face features, and obtain the target gaze direction based on the fourth face features.
8. A computer-readable storage medium, characterized in that, Several programs are stored on the computer-readable storage medium, and the several programs are used to be called by the processor and execute the target gaze direction estimation method as claimed in claim 1.
Citation Information
Patent Citations
Facial expression recognition method and device based on Emo-ResNet, equipment and medium
CN115862091A
Target tracking method based on convolution Transform combination
CN116645625A