A method, device, equipment and storage medium for head pose estimation

The two-dimensional and three-dimensional key point coordinates are output through the dual-branch network model, and combined with PnP solution, the stability problem of head pose estimation in large-scale expressions is solved, and the stability and reliability of head pose estimation is achieved.

CN117011929BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211130441.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-15
Publication Date
2025-07-11
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

The existing head posture estimation methods are insufficient in stability and reliability when large expressions are large, especially the conversion method based on 2D image information is obvious when large expressions are large.

Method used

The two-branch network model is used to output two-dimensional key point coordinates and uncertainties and three-dimensional key point coordinates respectively. The head posture is calculated through PnP solution, and the real-time update of the three-dimensional key point coordinates is used to improve stability.

Benefits of technology

Maintain the stability and reliability of head posture estimation when expressing large scales, improving the accuracy of the coordinate correspondence between 2D key points and 3D key points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117011929B_ABST
    Figure CN117011929B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a head pose estimation method, device, equipment, and storage medium, which are used to ensure the stable and reliable head pose estimation. The method includes: obtaining an image to be recognized, where the image to be recognized includes a target face image; inputting the image to be recognized into a first network model to obtain a set of two-dimensional key point coordinates of the target face image in the image to be recognized, the uncertainty of each two-dimensional key point coordinate in the set of two-dimensional key point coordinates, and a set of three-dimensional key point coordinates of the target face image. The first network model includes a first branch network and a second branch network. The first branch network is used to recognize and obtain the two-dimensional key point coordinates and the uncertainty, and the second branch network is used to recognize and obtain the three-dimensional key point coordinates; and recognizing the head pose corresponding to the target face image in the image to be recognized according to the set of two-dimensional key point coordinates, the uncertainty, and the set of three-dimensional key point coordinates. The technical solution provided by the present application can be applied to the fields of artificial intelligence and computer vision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to a head posture estimation method, apparatus, device and storage medium. Background Art

[0002] In the context of computer vision, head pose estimation is most often interpreted as the ability to infer the direction of a person's head relative to the camera view. Therefore, head pose estimation is a very important part of visual motion capture. Accurate head pose can allow the virtual image to perfectly replicate the head movements of the person, making the virtual human animation more vivid, flexible, and realistic. At present, the more mainstream head pose estimation methods mostly use traditional motion sensors, and the other is to obtain the three-dimensional coordinate information of the head through a three-dimensional (3-dimension, 3D) image acquisition device for judgment. Limited by the fact that the current mainstream video image acquisition devices collect two-dimensional (2-dimension, 2D) image information, based on the coordinate information of the key points of the face, the 3D conversion of the 2D coordinate information in the world coordinate system is realized, thereby obtaining the 3D coordinate information of the person's head pose, and then according to the change of coordinates, the estimation of the head pose and the judgment of the head movement are realized.

[0003] The above method is based on solving the motion of 3D to 2D point pairs (also known as Perspective-n-Point, PnP). This method first estimates the 2D key points of the face; then calibrates the corresponding 3D points in a fixed 3D head model. Through PnP solution, the transformation posture from 3D points to 2D key points can be obtained. The above method is generally accurate and has strong interpretability, but when a person makes a large expression, the jitter will be obvious.

[0004] Therefore, there is an urgent need for a head posture estimation method that can ensure stable and reliable head posture estimation when making large facial expressions. Summary of the invention

[0005] The embodiments of the present application provide a head posture estimation method, apparatus, device and storage medium for ensuring the stability and reliability of head posture estimation.

[0006] In view of this, on the one hand, the present application provides a head pose estimation method, including: obtaining an image to be recognized, where the image to be recognized includes a target face image; inputting the image to be recognized into a first network model to obtain a set of two-dimensional key point coordinates of the target face image in the image to be recognized, the uncertainty of each two-dimensional key point coordinate in the set of two-dimensional key point coordinates, and a set of three-dimensional key point coordinates of the target face image, where the first network model includes a first branch network and a second branch network, where the first branch network is used to recognize the set of two-dimensional key point coordinates and the uncertainty, and the second branch network is used to recognize the set of three-dimensional key point coordinates; and recognizing the head pose corresponding to the target face image in the image to be recognized according to the set of two-dimensional key point coordinates, the uncertainty, and the set of three-dimensional key point coordinates.

[0007] On the other hand, the present application provides a head pose estimation device, including:

[0008] An acquisition module, configured to obtain an image to be recognized, where the image to be recognized includes a target face image;

[0009] A processing module, configured to input the image to be recognized into a first network model to obtain a set of two-dimensional key point coordinates of the target face image in the image to be recognized, the uncertainty of each two-dimensional key point coordinate in the set of two-dimensional key point coordinates, and a set of three-dimensional key point coordinates of the target face image, where the first network model includes a first branch network and a second branch network, where the first branch network is used to recognize the set of two-dimensional key point coordinates and the uncertainty, and the second branch network is used to recognize the set of three-dimensional key point coordinates;

[0010] An output module, configured to recognize the head pose corresponding to the target face image in the image to be recognized according to the set of two-dimensional key point coordinates, the uncertainty, and the set of three-dimensional key point coordinates.

[0011] In a possible design, in another implementation manner of the other aspect of the embodiments of the present application, the output module is specifically configured to sort each two-dimensional key point coordinate in the set of two-dimensional key point coordinates according to the uncertainty, and remove the two-dimensional key point coordinates with an uncertainty greater than a preset threshold from the set of two-dimensional key point coordinates to obtain an intermediate set of two-dimensional key point coordinates;

[0012] Obtaining an intermediate set of three-dimensional key point coordinates from the set of three-dimensional key point coordinates according to the intermediate set of two-dimensional key point coordinates;

[0013] Recognizing the head pose corresponding to the target face image in the image to be recognized according to the intermediate set of two-dimensional key point coordinates and the intermediate set of three-dimensional key point coordinates.

[0014] In a possible design, in another implementation of another aspect of the embodiments of the present application, the output module is specifically configured to use PnP algorithm to calculate and identify the head pose corresponding to the target face image in the image to be recognized according to the intermediate two-dimensional key point coordinate set and the intermediate three-dimensional key point coordinate set.

[0015] In a possible design, in another implementation of another aspect of the embodiments of the present application, the acquisition module is further configured to acquire a first training sample set and establish a first initial network model. The first training sample set is labeled with a real two-dimensional key point coordinate set and a real three-dimensional key point coordinate set of a face image. The first initial network model includes a feature extraction network layer, an initial first branch network, and an initial second branch network. The initial first branch network is used to identify and output two-dimensional key point coordinates and uncertainty, and the initial first branch network is used to identify and output three-dimensional key point coordinates;

[0016] The head pose estimation device further includes a training module, configured to input the first training sample set into the feature extraction network layer to obtain the feature representation of each training sample in the first training sample set;

[0017] Input the feature representation into the initial first branch network to obtain a predicted two-dimensional key point coordinate set and the predicted uncertainty corresponding to the predicted two-dimensional key point coordinates, and input the feature representation into the initial second branch network to obtain a predicted three-dimensional key point coordinate set;

[0018] Calculate a first loss value according to the predicted two-dimensional key point coordinate set, the predicted uncertainty, and the real two-dimensional key point coordinate set, and calculate a second loss value according to the predicted three-dimensional key point coordinate set and the real three-dimensional key point coordinate set;

[0019] Adjust the initial first branch network according to the first loss value to obtain the first branch network, and adjust the initial second branch network according to the second loss value to obtain the second branch network;

[0020] Obtain the first network model according to the first branch network and the second branch network.

[0021] In a possible design, in another implementation of another aspect of the embodiments of the present application, the acquisition module is specifically configured to collect a training image set through a depth camera. Each training image in the training image set includes three-dimensional point cloud data of a face image and a real head pose;

[0022] Perform pose projection on the three-dimensional point cloud data to obtain two-dimensional key point data of the face image in the training image;

[0023] Input the training image set into the image processing network to output the training sample set.

[0024] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the obtaining module is specifically configured to input the training image set into the image processing network to obtain sparse key points of the face images in the training images, where the sparse key points include facial feature points and face contour points of the face images in the training images;

[0025] Obtain the target face image from the training image according to the face contour points;

[0026] Horizontally align and scale the target face image to a target size according to the facial feature points to obtain the training samples in the training sample set.

[0027] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the sparse key points include at least five facial feature points and four face contour points.

[0028] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the feature extraction network includes a residual neural network ResNet and a pooling layer, the first branch network is a fully connected layer, and the second branch network is a fully connected layer.

[0029] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the training module is specifically configured to calculate a first loss value according to the predicted two-dimensional key point coordinate set, the predicted uncertainty, and the true two-dimensional key point coordinate set by using the Gaussian negative log-likelihood loss; calculate a second loss value according to the predicted three-dimensional key point coordinate set and the true three-dimensional key point coordinate set by using a regression loss function.

[0030] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the obtaining module is specifically configured to obtain a second training sample set and establish a second initial network model. The second training sample set is labeled with a true two-dimensional key point coordinate set, a true three-dimensional key point coordinate set, and a true head pose of a face image. The second initial network model includes a feature extraction network, an initial first branch network, an initial second branch network, and a calculation network. The initial first branch network is used to identify and output two-dimensional key point coordinates and uncertainty, the initial first branch network is used to identify and output three-dimensional key point coordinates, and the calculation network is used to solve the head pose according to the two-dimensional key point coordinates and the three-dimensional key point coordinates;

[0031] The head pose estimation device further includes a training module, configured to input the second training sample set into the feature extraction network layer to obtain feature representations of each training sample in the second training sample set;

[0032] Input the feature representation into the initial first branch network to obtain a set of predicted 2D key point coordinates and the predicted uncertainty corresponding to the predicted 2D key point coordinates, and input the feature representation into the initial second branch network to obtain a set of predicted 3D key point coordinates;

[0033] Calculate the predicted head pose from the set of predicted 2D key point coordinates and the set of predicted 3D key point coordinates;

[0034] Calculate a first loss value based on the set of predicted 2D key point coordinates, the predicted uncertainty, and the set of true 2D key point coordinates, calculate a second loss value based on the set of predicted 3D key point coordinates and the set of true 3D key point coordinates, and calculate a third loss value based on the predicted head pose and the true head pose;

[0035] Adjust the second initial network model according to the first loss value, the second loss value, and the third loss value to obtain the first network model.

[0036] In a possible design, in another implementation of another aspect of the embodiments of the present application, the feature extraction network includes a Residual Neural Network (ResNet) and a pooling layer, the first branch network is a fully connected layer, and the second branch network is a fully connected layer; the calculation network is a differentiable Perspective-n-Point (PnP) solution network.

[0037] In a possible design, in another implementation of another aspect of the embodiments of the present application, the acquisition module is specifically configured to acquire an image to be processed, where the image to be processed includes a target face image collected by a camera;

[0038] Input the image to be processed into an image preprocessing network to obtain sparse key points of the target face image, where the sparse key points include the facial feature points and the face contour points of the target face image;

[0039] Acquire the target face image from the image to be processed according to the face contour points;

[0040] Horizontally align and scale the target face image to a target size according to the facial feature points to obtain the image to be recognized.

[0041] Another aspect of the present application provides a computer device, including: a memory, a processor, and a bus system;

[0042] Wherein, the memory is used to store programs;

[0043] The processor is used to execute the programs in the memory, and the processor is used to execute the methods in the above aspects according to the instructions in the program code;

[0044] The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.

[0045] Another aspect of the present application provides a computer-readable storage medium storing instructions which, when run on a computer, cause the computer to execute the methods in the above aspects.

[0046] Another aspect of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the methods provided in the above aspects.

[0047] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages: The 2D key point coordinates and 3D key point coordinates of the target face image in the image to be recognized are respectively output by two branch networks, and then the head pose of the target face image is calculated according to the 2D key point coordinates and the 3D key point coordinates. Among them, since the 3D key point coordinates can be obtained in real time, the 3D head model can change with the change of the expression, so that the coordinate correspondence between the 2D key points and the 3D key points is more accurate, thus making the solution of the head pose stable and reliable when making large expressions. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic diagram of an architecture of an application system in an embodiment of the present application;

[0049] Figure 2 It is a schematic diagram of an architecture of a first network model in an embodiment of the present application;

[0050] Figure 3 It is another schematic diagram of an architecture of a first network model in an embodiment of the present application;

[0051] Figure 4 It is a schematic flowchart of a head pose estimation method in an embodiment of the present application;

[0052] Figure 5 It is a schematic diagram of a target face image in an image to be recognized in an embodiment of the present application;

[0053] Figure 6 It is a schematic flowchart of an image to be processed being processed by an image processing model to obtain an image to be recognized in an embodiment of the present application;

[0054] Figure 7 It is a schematic diagram of an embodiment of head pose estimation in an embodiment of the present application;

[0055] Figure 7aIt is a schematic diagram of a virtual image generated after head pose estimation for the image to be processed in the embodiment of the present application;

[0056] Figure 8 It is a schematic diagram of an embodiment of the head pose estimation device in the embodiment of the present application;

[0057] Figure 9 It is a schematic diagram of another embodiment of the head pose estimation device in the embodiment of the present application;

[0058] Figure 10 It is a schematic diagram of another embodiment of the head pose estimation device in the embodiment of the present application;

[0059] Figure 11 It is a schematic diagram of another embodiment of the head pose estimation device in the embodiment of the present application. Detailed implementation manners

[0060] The embodiment of the present application provides a head pose estimation method, device, equipment and storage medium, which are used to ensure the stability and reliability of head pose estimation.

[0061] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or equipment including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0062] In the context of computer vision, head pose estimation is most often interpreted as the ability to infer the direction of a person's head relative to the camera view. Therefore, head pose estimation is a very important part in visual motion capture. Accurate head pose can make the virtual image perfectly replicate the head movements of the person, making the virtual human animation more vivid, flexible, and realistic. At present, the more mainstream head pose estimation methods mostly use traditional motion sensors, and the other is to obtain the three-dimensional coordinate information of the head through a three-dimensional (3-dimension, 3D) image acquisition device for judgment. Limited by the current mainstream video image acquisition equipment, all two-dimensional (2-dimension, 2D) image information is collected. Therefore, based on the coordinate information of the key points of the face, the 3D conversion of the 2D coordinate information in the world coordinate system is realized, so as to obtain the 3D coordinate information of the head pose, and then the estimation of the head pose and the judgment of the head movement are realized according to the change of coordinates. The above method is based on the method of solving the motion of 3D to 2D points (also known as Perspective-n-Point, PnP). The method first estimates the 2D key points of the face; then calibrates the corresponding 3D points in a fixed 3D head model. Through PnP solution, the transformation posture from 3D point to 2D key point can be obtained. The above method is generally accurate and has strong interpretability, but when a person makes a large expression, the jitter will be very obvious. Therefore, there is an urgent need for a stable and reliable head posture estimation method that ensures head posture estimation when making large expressions. In order to solve the above solution, the present application provides the following technical solution: obtain an image to be identified, which includes a target face image; input the image to be identified into a first network model to obtain a two-dimensional key point coordinate set of the target face image in the image to be identified, the uncertainty of each two-dimensional key point coordinate in the two-dimensional key point coordinate set, and a three-dimensional key point coordinate set of the target face image, wherein the first network model includes a first branch network and a second branch network, wherein the first branch network is used to identify the two-dimensional key point coordinate set and the uncertainty, and the second branch network is used to identify the three-dimensional key point coordinate set; identify the head posture corresponding to the target face image in the image to be identified according to the two-dimensional key point coordinate set, the uncertainty and the three-dimensional key point coordinate set. In this way, since the coordinates of the 3D key points can be obtained in real time, the 3D head model can change with the change of expression, so that the coordinate correspondence between the 2D key points and the 3D key points is more accurate, so that the solution of the head posture is stable and reliable when making large expressions.

[0063] For ease of understanding, some professional terms in this application are explained below:

[0064] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, and augmented reality.

[0065] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.

[0066] Key points of facial features: Used to represent the positions of facial features on the human face, and the position of each facial feature is a key point. The facial feature key points involved in the embodiments of this application include the points corresponding to the left eye pupil, right eye pupil, tip of the nose, left corner of the mouth, and right corner of the mouth on the human face.

[0067] Eulerian angles: Refers to a set of three independent angular parameters proposed by Euler to determine the position of a rigid body undergoing fixed-axis rotation. In the embodiments of this application, a right-handed coordinate system is established based on the human face. As Figure 1 shown, the embodiments of this application take the facial pose angle as Eulerian angles as an example for illustration. In a three-dimensional rectangular coordinate system, this three-dimensional rectangular coordinate system has the center or centroid of the human head as the origin, the direction from one ear on the human face to the other ear as the X-axis direction, the direction from the top of the human head to the neck as the Y-axis, and the direction from the human face to the back of the head as the Z-axis. Eulerian angles include the following three angles:

[0068] Pitch angle: The angle of rotation around the X-axis;

[0069] Yaw angle: The angle of rotation around the Y-axis;

[0070] Roll angle: The angle of rotation around the Z-axis.

[0071] Visual motion capture: Traditional motion capture uses inertial sensors or markers attached to the performer to capture movements. Visual motion capture does not require the performer to wear any equipment and can capture the performer's facial and body movements using a single or multiple cameras.

[0072] 6DoF (Six Degrees of Freedom): DoF refers to the number of directions in which an object can move in 3D space, and there are a total of six degrees of freedom. That is, the head pose includes rotation and translation. Rotation is represented by three Euler angles, and translation is also represented by displacements in three directions, which together form the six-degree-of-freedom pose parameters.

[0073] PnP (Perspective-n-Point): PnP (Perspective-n-Point) is a method for solving the motion of 3D to 2D point pairs, aiming to solve the pose of the camera coordinate system relative to the world coordinate system. It describes how to estimate the pose of the camera (i.e., solve the rotation matrix and translation vector from the world coordinate system to the camera coordinate system) when the coordinates of known 3D points (relative to the world coordinate system) and the 2D coordinates of these points are given.

[0074] The convolutional layer (Convolutional layer, Conv) refers to the layer structure composed of several convolutional units in a convolutional neural network. A convolutional neural network (Convolutional Neural Network, CNN) is a feedforward neural network that includes at least two neural network layers. Each neural network layer contains several neurons, and the neurons are arranged in layers. Neurons in the same layer are not connected to each other, and the transmission of information between layers occurs only in one direction.

[0075] Pooling layer: Also known as the sampling layer, it is a layer structure that can extract features from the input values for the second time after the convolutional layer. The pooling layer can ensure the main features of the values in the previous layer and reduce the parameters and computational amount in the next layer. The pooling layer consists of multiple feature maps. One feature map in the convolutional layer corresponds to one feature map in the pooling layer, and the number of feature maps remains unchanged. By reducing the resolution of the feature maps, features with spatial invariance are obtained.

[0076] Fully Connected layer (Fully Connected layer, FC) refers to a layer structure in which each node is connected to all nodes in the previous layer and can be used to comprehensively process the features extracted by the neural network layer in the previous layer, playing the role of a "classifier" in the neural network model.

[0077] Backpropagation: Forward propagation refers to the feed-forward processing process of the model. Opposite to forward propagation, backpropagation means updating the weight parameters of each layer of the model according to the result output by the model. For example, if the model includes an input layer, a hidden layer, and an output layer, forward propagation means processing in the order of input layer - hidden layer - output layer, and backpropagation means updating the weight parameters of each layer in the order of output layer - hidden layer - input layer.

[0078] A head pose estimation method, device, equipment, and storage medium provided by an embodiment of the present application can ensure the stable and reliable head pose estimation. The following describes the exemplary applications of the electronic device provided by the embodiment of the present application. The electronic device provided by the embodiment of the present application can be implemented as various types of user terminals or as a server.

[0079] By running the head pose estimation solution provided by the embodiment of the present application, the electronic device can ensure the stable and reliable head pose estimation, that is, improve the stable and reliable head pose estimation of the electronic device itself, and is applicable to multiple application scenarios of head pose estimation. For example, augmented reality (AR) games, virtual reality (VR) games, assisting gaze estimation, modeling attention, fitting 3D models to videos, and performing face alignment.

[0080] See Figure 1 , Figure 1It is an optional architecture schematic diagram in an application scenario of the head pose estimation solution provided by the embodiments of the present application. To support a head pose estimation application, the terminal device 100 (exemplarily showing the terminal device 1001 and the terminal device 1002) is connected to the server 300 through the network 200, and the server 300 is connected to the database 400. The network 200 can be a wide area network, a local area network, or a combination of the two. The client for implementing the head pose estimation solution is deployed on the terminal device 100. The client can run on the terminal device 100 in the form of a browser, or in the form of an independent application (APP), etc. The specific presentation form of the client is not limited here. The server 300 involved in the present application can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device 100 can be a smart phone, a tablet computer, a notebook computer, a handheld computer, a personal computer, a smart TV, a smart watch, a vehicle-mounted device, a wearable device, etc., but is not limited thereto. The terminal device 100 and the server 300 can be directly or indirectly connected through the network 200 by wired or wireless communication methods, and the present application does not limit this here. The number of the server 300 and the terminal device 100 is also not limited. The solution provided by the present application can be completed independently by the terminal device 100, can be completed independently by the server 300, or can be completed by the cooperation of the terminal device 100 and the server 300. The present application does not make specific limitations on this. Among them, the database 400 can be regarded as an electronic filing cabinet in short - a place for storing electronic files. Users can perform operations such as adding, querying, updating, and deleting data in the files. The so - called "database" is a data set stored together in a certain way, shared by multiple users, having the smallest possible redundancy, and independent of application programs. The Database Management System (DBMS) is a computer software system designed to manage the database and generally has basic functions such as storage, interception, security guarantee, backup, etc.Database management systems can be classified according to the database models they support, such as relational, Extensible Markup Language (XML); or according to the types of computers they support, such as server clusters, mobile phones; or according to the query languages used, such as Structured Query Language (SQL), XQuery; or according to the performance impulse focus, such as maximum scale, highest running speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories. For example, they can support multiple query languages simultaneously. In this application, the database 400 can be used to store the training sample set and the image to be recognized. Of course, the storage location of the training sample set is not limited to the database. For example, it can also be stored in the terminal device 100, the blockchain, or the distributed file system of the server 300, etc.

[0081] In some embodiments, the server 300 can execute the head pose estimation method and the training method of the first network model in the head pose estimation provided by the embodiments of this application. In this embodiment, the first network model includes a first branch network and a second branch network. Among them, the first branch network is used to identify the two-dimensional key point coordinates and the uncertainty, and the second branch network is used to identify the three-dimensional key point coordinates. When executing the training method of the first network model, the specific process can be as follows: Obtain a first training sample set corresponding to the real 2D key point coordinates and real 3D key point coordinates from the terminal device 100 and / or the database 400, and establish the first initial network model. Detect the first training sample set through the first initial network model to obtain the predicted 2D key point coordinates, predicted 3D key point coordinates of each face in the first training sample set, and the uncertainty corresponding to the predicted 2D key point coordinates. According to the loss function including pre-designed loss factors (such as the interval value and the distance), determine the first loss value corresponding to the predicted 2D key point coordinates and the uncertainty corresponding to the predicted 2D key point coordinates, and determine the second loss value corresponding to the predicted 3D key point coordinates. Then, adjust the parameters of the first branch network according to the backpropagation of the first loss value; at the same time, adjust the parameters of the second branch network according to the backpropagation of the second loss value, so as to realize the training of the first initial network model to obtain the first network model. In this embodiment, the first branch network and the second branch network are trained independently and do not affect each other's parameter adjustment. This can improve the generalization ability of the first branch network and the second branch network. In this embodiment, when training the first network model, the server can calculate the first loss value using the Gaussian negative log-likelihood loss (GNLL). In an exemplary solution, the specific calculation process can adopt Formula 1:

[0082]

[0083] Among them, N is used to indicate the number of 2D key points, y is used to indicate the 2D key point coordinates, f(x) is used to represent the true 2D key point coordinates, and δ is used to represent the uncertainty of the 2D key point coordinates.

[0084] When the server calculates the second loss value, a regression loss function can be adopted. In an exemplary solution, when the server adopts L2LOSS, its specific calculation process can adopt Formula 2:

[0085]

[0086] Among them, N is used to indicate the number of 3D key points, y is used to indicate the 3D key point coordinates, and f(x) is used to represent the true 3D key point coordinates.

[0087] When the server 300 trains the first network model, the initial model architecture of the first network model may include a feature extraction network, a fully connected layer, a pooling layer, the first branch network, and the second branch network. Among them, the feature extraction network can be a CNN network such as a Residual Neural Network (ResNet), Le Net, or AlexNet, or a High-Resolution net V2P (HRNetV2P) with a feature pyramid, or a hierarchical vision self-attention model based on a mobile window (Swin Transformer), and the first branch network and the second branch network can be fully connected layers. In an exemplary solution, the first network model is illustrated below with the feature extraction network being ResNet50. As Figure 2As shown, the first network model includes the ResNet50. Among them, the ResNet50 includes 49 convolutional layers and one fully connected layer, and the fully connected layer is then connected to the pooling layer; the output of the pooling layer is respectively connected to two fully connected layers, one of the fully connected layers is the first branch network, and the other fully connected layer is the second branch network. The input of the network is 224×224×3. After the convolutional calculations of the first five parts, the output is 7×7×2048. The pooling layer will convert it into a feature vector, and finally the classifier will calculate this feature vector and output the class probability. The ResNet50 network structure can be divided into seven parts. The first part does not contain residual blocks and mainly performs calculations of convolution, regularization, activation function, and max pooling on the input. The second, third, fourth, and fifth parts of the structure all contain residual blocks, and there are structures that do not change the size of the residual blocks but only change the dimensions of the residual blocks. In the ResNet50 network structure, each residual block has three convolutional layers, so the network has a total of 1 + 3×(3 + 4 + 6 + 3) = 49 convolutional layers. Adding the final fully connected layer, there are a total of 50 layers, which is also the origin of the name ResNet50. The input of the ResNet50 network is 256x256×3. After the convolutional calculations of the first five parts, the output is a feature map of Nx2048x8x8, where N is the number of samples selected for one training (also known as batchsize). Then the bottom layer passes the Nx2048x8x8 feature map through the pooling layer to obtain a feature of Nx2048. The feature of Nx2048 will pass through two FC layers, respectively outputting 2D key point coordinates and uncertainty and 3D key point coordinates. The weight dimension of each FC layer is 2048x3660 (1220 points), where 3660 is regarded as 1220x3. For the 2D branch, 3 represents the x, y coordinates and the uncertainty δ. For the 3D branch, 3 represents the x, y, z coordinate values.

[0088] In some other embodiments, the server 300 may execute the head pose estimation method provided in the embodiments of the present application and the training method of the first network model in head pose estimation. In this embodiment, the first network model includes a first branch network, a second branch network, and a calculation network. Among them, the first branch network is used to identify the two-dimensional key point coordinates and the uncertainty, the second branch network is used to identify the three-dimensional key point coordinates, and the calculation network is used to estimate the head pose according to the 2D key point coordinates and the 3D key point coordinates. When executing the training method of the first network model, its specific process may be as follows: Obtain a second training sample set corresponding to labeled true 2D key point coordinates, true 3D key point coordinates, and true head poses from the terminal device 100 and / or the database 400, and establish the second initial network model. Detect and process the second training sample set through the second initial network model to obtain the predicted 2D key point coordinates, predicted 3D key point coordinates, and the uncertainty corresponding to the predicted 2D key point coordinates of each face in the second training sample set. According to a loss function including pre-designed loss factors (such as an interval value and a distance factor), determine a first loss value corresponding to the predicted 2D key point coordinates and the uncertainty corresponding to the predicted 2D key point coordinates, and determine a second loss value corresponding to the predicted 3D key point coordinates; then calculate the predicted head pose according to the predicted 2D key point coordinates and the predicted 3D key point coordinates, and calculate a third loss value according to the predicted head pose and the true head pose; furthermore, adjust the parameters of the first branch network and the second branch network by backpropagation according to the first loss value, the second loss value, and the third loss value, so as to realize the training of the second initial network model to obtain the first network model. In this embodiment, the first branch network and the second branch network can be jointly trained so that the two branches can influence each other and enhance the network learning ability. In this embodiment, when training the first network model, the server may calculate the first loss value by using the Gaussian negative log-likelihood loss (GaussianNLLLoss, GNLL). In an exemplary solution, the specific calculation process may adopt Formula 1:

[0089]

[0090] Among them, the N is used to indicate the number of 2D key points, the y is used to indicate the 2D key point coordinates, the f(x) is used to represent the true 2D key point coordinates, and the δ is used to represent the uncertainty of the 2D key point coordinates.

[0091] The server may calculate the second loss value by using a regression loss function. In an exemplary solution, the server adopts L2 LOSS, and its specific calculation process may adopt Formula 2:

[0092]

[0093] Among them, the N is used to indicate the number of the 3D key points, the y is used to indicate the 3D key point coordinates, and the f(x) is used to represent the true 3D key point coordinates.

[0094] When the server 300 trains the first network model, the initial model architecture of the first network model may include a feature extraction network, a fully connected layer, a pooling layer, the first branch network, the second branch network, and the calculation network. Among them, the feature extraction network may be a CNN network such as a Residual Neural Network (ResNet), Le Net, or AlexNet, or a High-Resolution net V2P (HRNetV2P) with a feature pyramid, or a hierarchical vision self-attention model based on a mobile window (Swin Transformer), and the first branch network and the second branch network may be fully connected layers. In an exemplary solution, the first network model is described below by taking the feature extraction network as ResNet50 as an example. As Figure 3As shown, the first network model includes the ResNet50, wherein the ResNet50 includes 49 convolutional layers and one fully connected layer, and the fully connected layer is then connected to the pooling layer; the output of the pooling layer is also connected to two fully connected layers, one of which is the first branch network, and the other is the second branch network. The input of the network is 224×224×3, and after the convolution calculation of the first five parts, the output is 7×7×2048, which is converted into a feature vector by the pooling layer. Finally, the classifier calculates the feature vector and outputs the category probability. The ResNet50 network structure can be divided into seven parts. The first part does not contain residual blocks, and mainly performs convolution, regularization, activation function, and maximum pooling calculations on the input. The second, third, fourth, and fifth structures all contain residual blocks, among which there is a structure that does not change the size of the residual block and is only used to change the dimension of the residual block. In the Resnet50 network structure, the residual block has three layers of convolution, so the network has a total of 1+3×(3+4+6+3)=49 convolution layers, plus the final fully connected layer, a total of 50 layers, which is also the origin of the name Resnet50. The input of the Resnet50 network is 256x256×3. After the convolution calculation of the first five parts, the output is a feature map of Nx2048x8x8, where N is the number of samples selected for one training (also known as batchsize). Then the Nx2048x8x8 feature map is passed through the pooling layer to obtain the Nx2048 feature. The Nx2048 feature will pass through two FC layers to output the 2D key point coordinates and uncertainty and 3D key point coordinates respectively. The weight dimension of each FC layer is 2048x3660 (1220 points), where 3660 is regarded as 1220x3. For the 2D branch, 3 represents the x, y coordinates and uncertainty δ. For the 3D branch, 3 represents the x, y, z coordinate values. Then, the predicted head posture is calculated based on the predicted 2D key point coordinates and uncertainty output by the first branch network and the predicted 3D key point coordinates output by the second branch network.

[0095] In this embodiment, the data of the first training sample set and the second training sample set can adopt the following technical solution: collect face images through the depth camera built into the iPhone, then use the augmented reality technology (ARKit) based on the Apple system to capture the facial 3D point cloud data and the corresponding head posture in real time, develop data collection software based on ARKit, and then collect facial data at a speed of 60 frames (FPS). In this way, facial data collection based on existing technology can reduce the difficulty of collecting training sample sets.

[0096] It can be understood that other external cameras can also be used when obtaining the first training sample set and the second training sample set. Among them, the external camera can be a depth camera or other cameras, that is, as long as face data acquisition can be achieved, the specific method is not limited here. Using an external device for face data in this way can reduce the requirements for hardware devices, thereby reducing costs.

[0097] After the first network model is trained, the server 300 can save the first network model locally, so as to provide a remote head pose estimation function for the terminal device 100. For example, the server 300 can receive the image to be recognized sent by the terminal device 100, and detect and process the image to be recognized through the first network model to obtain the head pose and the corresponding confidence probability of the target face image in the image to be recognized; finally, send the head pose to the terminal device 100, so that the terminal device 100 can display the head pose in the graphical interface 110 (exemplarily showing the graphical interface 1101 and the graphical interface 1102).

[0098] The server 300 can also train and send (deploy) the trained first network model to the terminal device 100, so as to realize head pose estimation locally on the terminal device 100. For example, the terminal device 100 can obtain the image to be recognized in real time or obtain the image to be recognized from other devices, and detect and process the image to be recognized through the first network model to obtain the head pose and the corresponding confidence probability of the target face image in the image to be recognized; finally, the terminal device 100 displays the head pose in the graphical interface 110 (exemplarily showing the graphical interface 1101 and the graphical interface 1102).

[0099] Based on the above system, specifically, please refer to Figure 4 As shown, an execution process of the head pose estimation method in this application can be as follows:

[0100] Step 1: Generate an image to be recognized for the target face. In this embodiment, the image to be recognized includes a target face image, where the target face image refers to an image that only includes a face image in the image to be recognized and no longer includes other background images. As Figure 5 shown, the Figure 5 in (a) is an image including a background image, while Figure 5 in (b) is the image to be recognized including the target face image. In this embodiment, various cameras can first collect the image to be processed including other background images, and then preprocess the image to be processed to obtain the image to be recognized. The specific process can be as Figure 6As shown: The to-be-processed image a is collected through a camera; then sparse key points in the to-be-processed image a are obtained through face detection, where the sparse key points can be face feature points and face contour points; then the target face image is cropped from the to-be-processed image according to the face contour points; then the target face image is horizontally aligned through the eye key points among the face feature points, and the image is scaled to the target size to obtain the to-be-recognized image.

[0101] Step 2: Detect the to-be-recognized image through the first network model to obtain 2D key point coordinates, the uncertainty of the 2D key point coordinates, and 3D key point coordinates.

[0102] Step 3: Screen the 2D key point coordinates and the 3D key point coordinates according to the uncertainty to obtain target 2D key point coordinates and target 3D key point coordinates with the uncertainty less than the prediction threshold.

[0103] Step 4: Use the PnP algorithm to estimate the head pose corresponding to the target face image with the target 2D key point coordinates and the target 3D key point coordinates.

[0104] It can be understood that in the specific implementation manner of this application, relevant data such as the to-be-detected image and the training sample set are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0105] Combined with the above introduction, the head pose estimation method in this application is introduced below with the terminal device as the execution subject. Please refer to Figure 7 , An embodiment of the head pose estimation method in the embodiment of this application includes:

[0106] 701. Obtain a to-be-recognized image, where the to-be-recognized image includes a target face image.

[0107] The terminal device can collect the to-be-processed image through its own camera, and then input the to-be-processed image into an image processing model for processing to obtain the to-be-recognized image. At the same time, the terminal device can also obtain the to-be-recognized image saved in the memory. The terminal device can also obtain the to-be-recognized image through an instant messaging application, where the instant messaging application refers to software that realizes online chatting and communication through instant messaging technology. The terminal device can also obtain the to-be-recognized image from the Internet. For example, obtain a video image from a video network on the Internet and extract a face image from the video image. For example, directly download a face image from the Internet, etc.

[0108] In an exemplary solution, the specific process may be as follows: Input the image to be processed into an image preprocessing network to obtain sparse key points of the target face image, where the sparse key points include facial feature points and face contour points of the target face image; obtain the target face image from the image to be processed according to the face contour points; horizontally align and scale the target face image to a target size according to the facial feature points to obtain the image to be recognized. The target face image refers to an image that only includes a face image in the image to be recognized and no longer includes other background images. As Figure 5 shown, in Figure 5 (a) in the figure is an image including a background image, while Figure 5 (b) in the figure is the image to be recognized including the target face image. In this embodiment, various cameras may first collect the image to be processed including other background images, and then preprocess the image to be processed to obtain the image to be recognized. The specific process may be as Figure 6 shown: Collect the image to be processed a through a camera; then obtain the sparse key points in the image to be processed a through face detection, where the sparse key points may be facial feature points and face contour points; then extract the target face image from the image to be processed according to the face contour points; then horizontally align the target face image through the eye key points in the facial feature points to obtain the image to be recognized.

[0109] 702. Input the image to be recognized into a first network model to obtain a set of two-dimensional key point coordinates of the target face image in the image to be recognized, the uncertainty of each two-dimensional key point in the set of two-dimensional key point coordinates, and a set of three-dimensional key point coordinates of the target face image, where the first network model includes a first branch network and a second branch network, where the first branch network is used to identify and obtain the set of two-dimensional key point coordinates and the uncertainty, and the second branch network is used to identify and obtain the set of three-dimensional key point coordinates.

[0110] The terminal device inputs the image to be recognized into the first network model, and then the feature extraction network of the first network model will perform corresponding feature extraction on the image to be recognized to obtain the final feature representation of the image to be recognized; then input the final feature representation into the first branch network and the second branch network of the first network model respectively; where the first branch network will output a set of two-dimensional key point coordinates (i.e., 2D key point coordinates) and uncertainty of the target face image in the image to be recognized; the second branch network will output a set of three-dimensional key point coordinates (i.e., 3D key point coordinates) of the target face image in the image to be recognized.

[0111] It can be understood that the training process of the first network model in this application may refer to Figures 2 to 3 shown, and details are not described here.

[0112] 703. Identify the head pose corresponding to the target face image in the image to be recognized based on the set of two-dimensional key point coordinates, the uncertainty, and the set of three-dimensional key point coordinates.

[0113] In this embodiment, the terminal device can screen the two-dimensional key points and the three-dimensional key points according to the uncertainty to obtain intermediate three-dimensional key points and intermediate two-dimensional key points, and then use the PnP algorithm to solve according to the intermediate three-dimensional key point coordinates and the intermediate two-dimensional key point coordinates to obtain the head pose corresponding to the target face image. In an exemplary solution, the terminal device sorts the uncertainties, eliminates 20% of the two-dimensional key points with the greatest uncertainty and uses the three-dimensional key points corresponding to the two-dimensional key points, and retains the other two-dimensional key points and uses the three-dimensional key points corresponding to the other two-dimensional key points for PnP calculation. In an exemplary solution, the terminal device can use the solvePnP algorithm built in opencv. The principle of this method is to iteratively solve the pose so that after the 3D key points are projected by the pose, they are as close as possible to the 2D coordinates.

[0114] It can be understood that the technical solution provided in the embodiments of the present application can be applied to virtual image construction, assisting gaze estimation, modeling attention, making 3D models suitable for videos, and performing face alignment. In an exemplary application scenario, it is illustrated by taking the construction of a game character as an example. The game device collects the facial data of the user through a camera and performs head pose estimation, so as to generate the image of the virtual object corresponding to the user, and then the virtual object interacts with other virtual objects in the game, thereby realizing game interaction. Among them, the process of the game device collecting facial data through the camera and generating a virtual object according to the facial data can be as Figure 7a shown, as Figure 7a shown, the head movement of the game character is generated by collecting facial movements. Specifically, the image to be processed collected can be as Figure 7a shown in (a), that is, the head movement shown by the target face image is tilting the head; then the head pose estimation provided by the embodiments of the present application is used to obtain as Figure 7a shown in (b), that is, the head movement of the corresponding game character is synchronously shown as tilting the head. In Figure 7a in this way, the head movement of the corresponding virtual object is generated by collecting the facial movement of the user, improving the game interaction experience, and at the same time realizing real-time synchronization of the head movement of the virtual object, improving the data processing efficiency.

[0115] In practical applications, this head pose estimation method can also be applied to live streaming or video recording. That is, when the user does not want to appear in the live video with their own appearance, the facial data of the user can be collected through a camera, and then a corresponding virtual image can be generated based on the facial data, and then the virtual image can be used for live streaming or video recording. In this way, the actions of the virtual image can be synchronized with the actions of the user, effectively realizing the interaction between the user and the users watching the video, while protecting the user's privacy.

[0116] The beneficial effects of the technical solution provided by this application will be described below with a specific example:

[0117] Obtain a dataset of about 500,000 images of 40 people to evaluate the technical indicators of our method and other methods. There are three methods participating in the evaluation: 1. Directly estimate the 6D0F parameters; 2. Adopt the PnP scheme, but do not estimate the uncertainty; 3. The technical solution provided by this application. The results are shown in Table 1:

[0118] Table 1

[0119]

[0120] As shown in Table 1, comparisons are made in all 6 dimensions of 6DoF, namely pitch, yaw, roll, tx, ty, and tz. From the results, the technical solution provided by this application is significantly better than other methods in terms of indicators in each dimension.

[0121] The head pose estimation device in this application will be described in detail below. Please refer to Figure 8 , Figure 8 is a schematic diagram of an embodiment of the head pose estimation device in an embodiment of this application. The head pose estimation device 20 includes:

[0122] An acquisition module 201, configured to acquire an image to be recognized, where the image to be recognized includes a target face image;

[0123] A processing module 202, configured to input the image to be recognized into a first network model to obtain a set of two-dimensional key point coordinates of the target face image in the image to be recognized, the uncertainty of each two-dimensional key point coordinate in the set of two-dimensional key point coordinates, and a set of three-dimensional key point coordinates of the target face image, where the first network model includes a first branch network and a second branch network, where the first branch network is used to recognize and obtain the set of two-dimensional key point coordinates and the uncertainty, and the second branch network is used to recognize and obtain the set of three-dimensional key point coordinates;

[0124] An output module 203, configured to recognize the head pose corresponding to the target face image in the image to be recognized according to the set of two-dimensional key point coordinates, the uncertainty, and the set of three-dimensional key point coordinates.

[0125] In an embodiment of the present application, a head pose estimation device is provided. By using the above device, the 2D key point coordinates and the 3D key point coordinates of the target face image in the image to be recognized are respectively output by two branch networks, and then the head pose of the target face image is calculated according to the 2D key point coordinates and the 3D key point coordinates. Among them, since the 3D key point coordinates can be obtained in real time, the 3D head model can change with the change of expressions, so that the coordinate correspondence between the 2D key points and the 3D key points is more accurate, thereby ensuring the stable and reliable head pose estimation when making large expressions.

[0126] Optionally, on the basis of the corresponding embodiment above, Figure 8 in another embodiment of the head pose estimation device 20 provided in the embodiment of the present application,

[0127] The output module 203 is specifically configured to sort each two-dimensional key point coordinate in the two-dimensional key point coordinate set according to the uncertainty, and remove the two-dimensional key point coordinates with uncertainty greater than a preset threshold from the two-dimensional key point coordinate set to obtain an intermediate two-dimensional key point coordinate set;

[0128] Obtain an intermediate three-dimensional key point coordinate set from the three-dimensional key point coordinate set according to the intermediate two-dimensional key point coordinate set;

[0129] Identify the head pose corresponding to the target face image in the image to be recognized according to the intermediate two-dimensional key point coordinate set and the intermediate three-dimensional key point coordinate set.

[0130] In an embodiment of the present application, a head pose estimation device is provided. By using the above device, the 2D key point coordinates and the 3D key point coordinates are screened according to the uncertainty corresponding to the 2D key point coordinates, so that the points with greater uncertainty are deleted during head pose estimation, making the head pose estimation more robust.

[0131] Optionally, on the basis of the corresponding embodiment above, Figure 8 in another embodiment of the head pose estimation device 20 provided in the embodiment of the present application, the output module 203 is specifically configured to use the PnP algorithm to solve and identify the head pose corresponding to the target face image in the image to be recognized according to the intermediate two-dimensional key point coordinate set and the intermediate three-dimensional key point coordinate set.

[0132] In an embodiment of the present application, a head pose estimation device is provided. By using the PnP algorithm to perform pose estimation, the head pose estimation is made more feasible.

[0133] Optionally, on the basis of the corresponding embodiment above, Figure 8Based on the corresponding embodiment, in another embodiment of the head pose estimation device 20 provided in the embodiments of the present application, as Figure 9 shown:

[0134] The acquisition module 201 is further configured to acquire a first training sample set and establish a first initial network model. The first training sample set is labeled with a set of true two-dimensional key point coordinates and a set of true three-dimensional key point coordinates of a face image. The first initial network model includes a feature extraction network layer, an initial first branch network, and an initial second branch network. The initial first branch network is used to identify and output two-dimensional key point coordinates and uncertainty, and the initial first branch network is used to identify and output three-dimensional key point coordinates;

[0135] The head pose estimation device further includes a training module 204, configured to input the first training sample set into the feature extraction network layer to obtain a feature representation of each training sample in the first training sample set;

[0136] Input the feature representation into the initial first branch network to obtain a set of predicted two-dimensional key point coordinates and the predicted uncertainty corresponding to the predicted two-dimensional key point coordinates, and input the feature representation into the initial second branch network to obtain a set of predicted three-dimensional key point coordinates;

[0137] Calculate a first loss value according to the set of predicted two-dimensional key point coordinates, the predicted uncertainty, and the set of true two-dimensional key point coordinates, and calculate a second loss value according to the set of predicted three-dimensional key point coordinates and the set of true three-dimensional key point coordinates;

[0138] Adjust the initial first branch network according to the first loss value to obtain the first branch network, and adjust the initial second branch network according to the second loss value to obtain the second branch network;

[0139] Obtain the first network model according to the first branch network and the second branch network.

[0140] In the embodiments of the present application, a head pose estimation device is provided. By using the above device, the first branch network and the second branch network are generated during the training process. The two branch networks respectively output the 2D key point coordinates and the 3D key point coordinates of the target face image in the image to be recognized, and then the head pose of the target face image is calculated according to the 2D key point coordinates and the 3D key point coordinates. Among them, since the 3D key point coordinates can be obtained in real time, the 3D head model can change with the change of the expression, so that the coordinate correspondence between the 2D key points and the 3D key points is more accurate, so that when making large expressions, the stability and reliability of the head pose estimation are ensured. At the same time, the first branch network and the second branch network are independently trained respectively, thereby increasing the generalization of the model.

[0141] Optionally, based on the corresponding embodiment above, in another embodiment of the head pose estimation device 20 provided by the embodiments of the present application, the acquisition module 201 is specifically configured to collect a training image set through a depth camera, and each training image in the training image set includes three-dimensional point cloud data of a face image and a true head pose; Figure 9 Project the three-dimensional point cloud data to obtain two-dimensional key point data of the face image in the training image;

[0142] Input the training image set into the image processing network to output the training sample set.

[0143] In the embodiments of the present application, a head pose estimation device is provided. By using the above device and collecting training images through a depth camera, it is more convenient to obtain the 3D point cloud data and the true head pose, thereby simplifying the processing flow of the training sample set.

[0144] Optionally, based on the corresponding embodiment above, in another embodiment of the head pose estimation device 20 provided by the embodiments of the present application, the acquisition module 201 is specifically configured to input the training image set into the image processing network to obtain sparse key points of the face image in the training image, and the sparse key points include facial feature points and face contour points of the face image in the training image;

[0145] Obtain the target face image from the training image according to the face contour points; Figure 9 Align the target face image horizontally and scale it to a target size according to the facial feature points to obtain the training samples in the training sample set.

[0146] In the embodiments of the present application, a head pose estimation device is provided. By using the above device, the face image is cropped and aligned through sparse key points, which can reduce the interference caused by background information in the images collected by the camera. And scaling the images to a unified size is beneficial to feature extraction of the images and reduces the training difficulty.

[0147] Optionally, based on the corresponding embodiment above, in another embodiment of the head pose estimation device 20 provided by the embodiments of the present application, the sparse key points include at least five facial feature points and four face contour points.

[0148] In the embodiments of the present application, a head pose estimation device is provided. By using the above device, the process of image preprocessing can be reduced while ensuring accurate cropping and alignment.

[0149] Optionally, based on the corresponding embodiment above, in another embodiment of the head pose estimation device 20 provided by the embodiments of the present application, the sparse key points include at least five facial feature points and four face contour points. Figure 9 In the embodiments of the present application, a head pose estimation device is provided. By using the above device, the process of image preprocessing can be reduced while ensuring accurate cropping and alignment.

[0150] In the embodiments of the present application, a head pose estimation device is provided. By using the above device, the process of image preprocessing can be reduced while ensuring accurate cropping and alignment.

[0151] Optionally, based on the above Figure 9 In another embodiment of the head pose estimation device 20 provided by the embodiments of the present application, on the basis of the corresponding embodiment, the feature extraction network includes a residual neural network ResNet and a pooling layer, the first branch network is a fully connected layer, and the second branch network is a fully connected layer.

[0152] In the embodiments of the present application, a head pose estimation device is provided. By using the above device, the feasibility of the solution can be increased.

[0153] Optionally, based on the above Figure 9 In another embodiment of the head pose estimation device 20 provided by the embodiments of the present application, on the basis of the corresponding embodiment, the training module 204 is specifically configured to calculate a first loss value according to the predicted two-dimensional key point coordinate set, the predicted uncertainty, and the true two-dimensional key point coordinate set by using the Gaussian negative log-likelihood loss; calculate a second loss value according to the predicted three-dimensional key point coordinate set and the true three-dimensional key point coordinate set by using a regression loss function.

[0154] In the embodiments of the present application, a head pose estimation device is provided. By using the above device, the feasibility of the solution can be increased.

[0155] Optionally, based on the above Figure 8 In another embodiment of the head pose estimation device 20 provided by the embodiments of the present application, on the basis of the corresponding embodiment, as Figure 9 shown,

[0156] The obtaining module 201 is specifically configured to obtain a second training sample set and establish a second initial network model. The second training sample set is labeled with a true two-dimensional key point coordinate set, a true three-dimensional key point coordinate set, and a true head pose of a face image. The second initial network model includes a feature extraction network, an initial first branch network, an initial second branch network, and a calculation network. The initial first branch network is used to identify and output two-dimensional key point coordinates and uncertainties. The initial first branch network is used to identify and output three-dimensional key point coordinates. The calculation network is used to solve the head pose according to the two-dimensional key point coordinates and the three-dimensional key point coordinates;

[0157] The head pose estimation device further includes a training module 204, configured to input the second training sample set into the feature extraction network layer to obtain a feature representation of each training sample in the second training sample set;

[0158] Input the feature representation into the initial first branch network to obtain a predicted two-dimensional key point coordinate set and the predicted uncertainty corresponding to the predicted two-dimensional key point coordinates, and input the feature representation into the initial second branch network to obtain a predicted three-dimensional key point coordinate set;

[0159] Calculate the predicted head pose based on the set of predicted 2D key point coordinates and the set of predicted 3D key point coordinates;

[0160] Calculate a first loss value based on the set of predicted 2D key point coordinates, the predicted uncertainty, and the set of true 2D key point coordinates, calculate a second loss value based on the set of predicted 3D key point coordinates and the set of true 3D key point coordinates, and calculate a third loss value based on the predicted head pose and the true head pose;

[0161] Adjust the second initial network model according to the first loss value, the second loss value, and the third loss value to obtain the first network model.

[0162] In an embodiment of the present application, a head pose estimation device is provided. By using the above device, the first branch network and the second branch network are generated during the training process, and the 2D key point coordinates and the 3D key point coordinates of the target face image in the image to be recognized are output through the two branch networks respectively, and then the head pose of the target face image is calculated according to the 2D key point coordinates and the 3D key point coordinates. Among them, since the 3D key point coordinates can be obtained in real time, the 3D head model can change with the change of the expression, so that the coordinate correspondence between the 2D key points and the 3D key points is more accurate, so that when making large expressions, the calculation of the head pose is stable and reliable. At the same time, the first branch network and the second branch network are jointly trained, thereby increasing the learnability of the model.

[0163] Optionally, on the basis of the corresponding embodiment above, Figure 9 In another embodiment of the head pose estimation device 20 provided in the embodiment of the present application, the feature extraction network includes a residual neural network ResNet and a pooling layer, the first branch network is a fully connected layer, and the second branch network is a fully connected layer; the calculation network is a differentiable PnP solution network.

[0164] In an embodiment of the present application, a head pose estimation device is provided. By using the above device, the feasibility of the solution can be increased.

[0165] Optionally, on the basis of the corresponding embodiment above, Figure 8 In another embodiment of the head pose estimation device 20 provided in the embodiment of the present application, the acquisition module 201 is specifically configured to acquire an image to be processed, and the image to be processed includes a target face image collected by a camera;

[0166] Input the image to be processed into an image preprocessing network to obtain sparse key points of the target face image, and the sparse key points include the facial feature points and the face contour points of the target face image;

[0167] Obtain the target face image from the image to be processed according to the face contour points;

[0168] Horizontally align and scale the target face image to the target size according to the facial feature points to obtain the image to be recognized. In the embodiments of the present application, a head pose estimation device is provided. By using the above device, the face image is cropped and aligned through sparse key points, which can reduce the interference caused by background information in the image collected by the camera. And scaling the image to a unified size is beneficial to feature extraction of the image.

[0169] The head pose estimation device provided by the present application can be used in a server. Please refer to Figure 10 , Figure 10 , which is a schematic structural diagram of a server provided by an embodiment of the present application. The server 300 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the server 300.

[0170] The server 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0171] The steps executed by the server in the above embodiments may be based on the Figure 10 shown server structure.

[0172] The head pose estimation device provided by the present application can be used in a terminal device. Please refer to Figure 11 , for the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. In the embodiments of the present application, a smart phone is taken as an example of the terminal device for illustration:

[0173] Figure 11 Shown is a block diagram of a partial structure of a smart phone related to the terminal device provided in an embodiment of the present application. Referring to Figure 11 , the smart phone includes components such as a radio frequency (RF) circuit 410, a memory 420, an input unit 430, a display unit 440, a sensor 450, an audio circuit 460, a wireless fidelity (WiFi) module 470, a processor 480, and a power supply 490. Those skilled in the art can understand that Figure 11 the smart phone structure shown in

[0174] does not limit the smart phone, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Figure 11 The following specifically introduces each component of the smart phone:

[0175] The RF circuit 410 can be used to receive and send signals during information reception or call processes. Specifically, after receiving the downlink information of the base station, it is given to the processor 480 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 410 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 410 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the global system of mobile communication (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), long term evolution (LTE), email, short messaging service (SMS), etc.

[0176] The memory 420 can be used to store software programs and modules. The processor 480 executes various functional applications and data processing of the smart phone by running the software programs and modules stored in the memory 420. The memory 420 may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store data created according to the use of the smart phone (such as audio data, phone book, etc.). In addition, the memory 420 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0177] The input unit 430 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the smart phone. Specifically, the input unit 430 may include a touch panel 431 and other input devices 432. The touch panel 431, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 431), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 431 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 480, and can receive and execute commands sent by the processor 480. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 431. In addition to the touch panel 431, the input unit 430 may also include other input devices 432. Specifically, the other input devices 432 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.

[0178] The display unit 440 can be used to display information input by the user or information provided to the user, as well as various menus of the smart phone. The display unit 440 may include a display panel 441. Optionally, the display panel 441 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 431 can cover the display panel 441. When the touch panel 431 detects a touch operation on or near it, it is transmitted to the processor 480 to determine the type of touch event. Subsequently, the processor 480 provides a corresponding visual output on the display panel 441 according to the type of touch event. Although in Figure 11 the touch panel 431 and the display panel 441 are implemented as two independent components to realize the input and input functions of the smart phone, in some embodiments, the touch panel 431 and the display panel 441 can be integrated to realize the input and output functions of the smart phone.

[0179] The smart phone may further include at least one sensor 450, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 441 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 441 and / or the backlight when the smart phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used in applications for identifying the posture of the smart phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the smart phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0180] The audio circuit 460, the speaker 461, and the microphone 462 can provide an audio interface between the user and the smart phone. The audio circuit 460 can transmit the electrical signal converted from the received audio data to the speaker 461, and the speaker 461 converts it into a sound signal for output; on the other hand, the microphone 462 converts the collected sound signal into an electrical signal, which is received by the audio circuit 460 and converted into audio data. After the audio data is output to the processor 480 for processing, it is sent through the RF circuit 410 to, for example, another smart phone, or the audio data is output to the memory 420 for further processing.

[0181] WiFi belongs to short - range wireless transmission technology. Through the WiFi module 470, a smart phone can help users send and receive emails, browse the web, and access streaming media, etc. It provides users with wireless broadband Internet access. Although Figure 11 the WiFi module 470 is shown, it can be understood that it does not belong to an essential component of the smart phone and can be omitted entirely within the scope of not changing the essence of the invention as needed.

[0182] The processor 480 is the control center of the smart phone, connecting various parts of the entire smart phone using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 420, and by calling data stored in the memory 420, it executes various functions of the smart phone and processes data, thereby monitoring the smart phone as a whole. Optionally, the processor 480 may include one or more processing units; optionally, the processor 480 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 480 either.

[0183] The smart phone also includes a power source 490 (such as a battery) that powers each component. Optionally, the power source can be logically connected to the processor 480 through a power management system, thereby implementing functions such as management of charging, discharging, and power consumption management through the power management system.

[0184] Although not shown, the smart phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0185] In the above - mentioned embodiments, the steps executed by the terminal device can be based on the Figure 11 shown terminal device structure.

[0186] In the embodiments of the present application, a computer - readable storage medium is also provided. A computer program is stored in the computer - readable storage medium. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0187] In the embodiments of the present application, a computer program product including a program is also provided. When it runs on a computer, it causes the computer to execute the methods described in the foregoing various embodiments.

[0188] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0189] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0190] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0191] In addition, in each embodiment of the present application, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0192] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0193] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A head pose estimation method, characterized in that, Including: Obtain an image to be recognized in real time, where the image to be recognized includes a target face image; Input the image to be recognized into a first network model to obtain a set of two-dimensional key point coordinates of the target face image in the image to be recognized, the uncertainty of each two-dimensional key point coordinate in the set of two-dimensional key point coordinates, and a set of three-dimensional key point coordinates of the target face image, where the first network model includes a first branch network and a second branch network, and where the first branch network is used to recognize and obtain the set of two-dimensional key point coordinates and the uncertainty, and the second branch network is used to recognize and obtain the set of three-dimensional key point coordinates; Sort each two-dimensional key point coordinate in the set of two-dimensional key point coordinates according to the uncertainty, and remove the two-dimensional key point coordinates with an uncertainty greater than a preset threshold from the set of two-dimensional key point coordinates to obtain an intermediate set of two-dimensional key point coordinates; Obtain an intermediate set of three-dimensional key point coordinates from the set of three-dimensional key point coordinates according to the intermediate set of two-dimensional key point coordinates; Identify the head pose corresponding to the target face image in the image to be recognized according to the intermediate set of two-dimensional key point coordinates and the intermediate set of three-dimensional key point coordinates.

2. The method according to claim 1, characterized in that, The identifying the head pose corresponding to the target face image in the image to be recognized according to the intermediate set of two-dimensional key point coordinates and the intermediate set of three-dimensional key point coordinates includes: Identify the head pose corresponding to the target face image in the image to be recognized by using PnP solution according to the intermediate set of two-dimensional key point coordinates and the intermediate set of three-dimensional key point coordinates.

3. The method according to any one of claims 1 to 2, characterized in that The method further includes: Obtain a first training sample set and establish a first initial network model. The first training sample set is labeled with a set of true two-dimensional key point coordinates and a set of true three-dimensional key point coordinates of a face image. The first initial network model includes a feature extraction network, an initial first branch network, and an initial second branch network. The initial first branch network is used to recognize and output two-dimensional key point coordinates and uncertainty, and the initial first branch network is used to recognize and output three-dimensional key point coordinates; Input the first training sample set into the feature extraction network to obtain the feature representation of each training sample in the first training sample set; Input the feature representation into the initial first branch network to obtain a set of predicted two-dimensional key point coordinates and the predicted uncertainty corresponding to the predicted two-dimensional key point coordinates, and input the feature representation into the initial second branch network to obtain a set of predicted three-dimensional key point coordinates; Calculate a first loss value according to the set of predicted two-dimensional key point coordinates, the predicted uncertainty, and the set of true two-dimensional key point coordinates, and calculate a second loss value according to the set of predicted three-dimensional key point coordinates and the set of true three-dimensional key point coordinates; Adjust the initial first branch network according to the first loss value to obtain the first branch network, and adjust the initial second branch network according to the second loss value to obtain the second branch network; Obtain the first network model according to the first branch network and the second branch network.

4. The method according to claim 3, wherein The obtaining the first training sample set includes: Collect a training image set through a depth camera, where each training image in the training image set includes 3D point cloud data of a face image and a true head pose; Perform pose projection on the 3D point cloud data to obtain 2D key point data of the face image in the training image; Input the training image set into an image processing network to output the training sample set.

5. The method according to claim 4, characterized in that, The obtaining the training sample set by passing the training image set through the image processing network includes: Input the training image set into the image processing network to obtain sparse key points of the face image in the training image, where the sparse key points include facial feature points and face contour points of the face image in the training image; Obtain the target face image from the training image according to the face contour points; Horizontally align and scale the target face image to a target size according to the facial feature points to obtain a training sample in the training sample set.

6. The method according to claim 5, wherein The sparse key points include at least five facial feature points and four face contour points.

7. The method according to claim 3, wherein The feature extraction network includes a residual neural network ResNet and a pooling layer, the first branch network is a fully connected layer, and the second branch network is a fully connected layer.

8. The method according to claim 3, wherein The calculating a first loss value according to the predicted 2D key point coordinate set, the predicted uncertainty, and the true 2D key point coordinate set includes: Calculate a first loss value according to the predicted 2D key point coordinate set, the predicted uncertainty, and the true 2D key point coordinate set by using a Gaussian negative log-likelihood loss; The calculating a second loss value according to the predicted 3D key point coordinate set and the true 3D key point coordinate set includes: Calculate a second loss value according to the predicted 3D key point coordinate set and the true 3D key point coordinate set by using a regression loss function.

9. The method according to any one of claims 1 to 2, characterized in that, The method further includes: Obtain a second training sample set and establish a second initial network model. The second training sample set is labeled with a true 2D key point coordinate set, a true 3D key point coordinate set, and a true head pose of a face image. The second initial network model includes a feature extraction network, an initial first branch network, an initial second branch network, and a calculation network. The initial first branch network is used to identify and output 2D key point coordinates and uncertainty, the initial first branch network is used to identify and output 3D key point coordinates, and the calculation network is used to solve the head pose according to the 2D key point coordinates and the 3D key point coordinates; Input the second training sample set into the feature extraction network to obtain a feature representation of each training sample in the second training sample set; Input the feature representation into the initial first branch network to obtain a predicted 2D key point coordinate set and the predicted uncertainty corresponding to the predicted 2D key point coordinates, and input the feature representation into the initial second branch network to obtain a predicted 3D key point coordinate set; Calculate a predicted head pose according to the predicted 2D key point coordinate set and the predicted 3D key point coordinate set; Calculate a first loss value based on the predicted 2D key point coordinate set, the predicted uncertainty, and the ground-truth 2D key point coordinate set, calculate a second loss value based on the predicted 3D key point coordinate set and the ground-truth 3D key point coordinate set, and calculate a third loss value based on the predicted head pose and the ground-truth head pose; Adjust the second initial network model according to the first loss value, the second loss value, and the third loss value to obtain the first network model.

10. The method according to claim 9, wherein The feature extraction network includes a Residual Neural Network (ResNet) and a pooling layer. The first branch network is a fully connected layer, and the second branch network is a fully connected layer; the calculation network is a differentiable PnP solution network.

11. The method according to any one of claims 1 to 2 or any one of claims 4 to 8 or claim 10, characterized in that, The obtaining of the image to be recognized includes: Obtain a to-be-processed image, where the to-be-processed image includes a target face image collected by a camera; Input the to-be-processed image into an image preprocessing network to obtain sparse key points of the target face image, where the sparse key points include facial feature points and face contour points of the target face image; Obtain the target face image from the to-be-processed image according to the face contour points; Horizontally align and scale the target face image to a target size according to the facial feature points to obtain the image to be recognized.

12. A head pose estimation device, characterized in that, It includes: An acquisition module, configured to acquire an image to be recognized in real time, where the image to be recognized includes a target face image; A processing module, configured to input the image to be recognized into the first network model to obtain a 2D key point coordinate set of the target face image in the image to be recognized, the uncertainty of each 2D key point coordinate in the 2D key point coordinate set, and a 3D key point coordinate set of the target face image, where the first network model includes a first branch network and a second branch network, and the first branch network is configured to recognize and obtain the 2D key point coordinate set and the uncertainty, and the second branch network is configured to recognize and obtain the 3D key point coordinate set; An output module, configured to recognize the head pose corresponding to the target face image in the image to be recognized according to the 2D key point coordinate set, the uncertainty, and the 3D key point coordinate set; Specifically, the output module is configured to sort each 2D key point coordinate in the 2D key point coordinate set according to the uncertainty, and remove 2D key point coordinates with uncertainties greater than a preset threshold from the 2D key point coordinate set to obtain an intermediate 2D key point coordinate set; Obtain an intermediate 3D key point coordinate set from the 3D key point coordinate set according to the intermediate 2D key point coordinate set; Recognize the head pose corresponding to the target face image in the image to be recognized according to the intermediate 2D key point coordinate set and the intermediate 3D key point coordinate set.

13. The device according to claim 12, characterized in that, Specifically, the output module is configured to use PnP solution to recognize the head pose corresponding to the target face image in the image to be recognized according to the intermediate 2D key point coordinate set and the intermediate 3D key point coordinate set.

14. The device according to any one of claims 12 to 13, characterized in that, The acquisition module is further configured to acquire a first training sample set and establish a first initial network model. The first training sample set is labeled with a set of true two-dimensional key point coordinates and a set of true three-dimensional key point coordinates of a face image. The first initial network model includes a feature extraction network, an initial first branch network, and an initial second branch network. The initial first branch network is used to identify and output two-dimensional key point coordinates and uncertainty, and the initial first branch network is used to identify and output three-dimensional key point coordinates; The head pose estimation device further includes a training module, configured to input the first training sample set into the feature extraction network to obtain a feature representation of each training sample in the first training sample set; Input the feature representation into the initial first branch network to obtain a set of predicted two-dimensional key point coordinates and the predicted uncertainty corresponding to the predicted two-dimensional key point coordinates, and input the feature representation into the initial second branch network to obtain a set of predicted three-dimensional key point coordinates; Calculate a first loss value according to the set of predicted two-dimensional key point coordinates, the predicted uncertainty, and the set of true two-dimensional key point coordinates, and calculate a second loss value according to the set of predicted three-dimensional key point coordinates and the set of true three-dimensional key point coordinates; Adjust the initial first branch network according to the first loss value to obtain the first branch network, and adjust the initial second branch network according to the second loss value to obtain the second branch network; Obtain the first network model according to the first branch network and the second branch network.

15. The device according to claim 14, wherein The acquisition module is specifically configured to collect a training image set through a depth camera. Each training image in the training image set includes three-dimensional point cloud data of a face image and a true head pose; Perform pose projection on the three-dimensional point cloud data to obtain two-dimensional key point data of the face image in the training image; Input the training image set into an image processing network to output the training sample set.

16. The device according to claim 15, characterized in that, The acquisition module is specifically configured to input the training image set into the image processing network to obtain sparse key points of the face image in the training image. The sparse key points include facial feature points and face contour points of the face image in the training image; Obtain the target face image from the training image according to the face contour points; Horizontally align and scale the target face image to a target size according to the facial feature points to obtain a training sample in the training sample set.

17. The device according to claim 16, characterized in that, The sparse key points include at least five facial feature points and four face contour points.

18. The device according to claim 14, characterized in that, The feature extraction network includes a Residual Neural Network (ResNet) and a pooling layer. The first branch network is a fully connected layer, and the second branch network is a fully connected layer.

19. The device according to claim 14, characterized in that, The training module is specifically configured to calculate a first loss value according to the set of predicted two-dimensional key point coordinates, the predicted uncertainty, and the set of true two-dimensional key point coordinates by using a Gaussian negative log-likelihood loss; calculate a second loss value according to the set of predicted three-dimensional key point coordinates and the set of true three-dimensional key point coordinates by using a regression loss function.

20. The device according to any one of claims 12 to 13, characterized in that The obtaining module is specifically configured to obtain a second training sample set and establish a second initial network model. The second training sample set is labeled with a set of true two-dimensional key point coordinates, a set of true three-dimensional key point coordinates, and a true head pose of a face image. The second initial network model includes a feature extraction network, an initial first branch network, an initial second branch network, and a calculation network. The initial first branch network is used to identify and output two-dimensional key point coordinates and uncertainty. The initial first branch network is used to identify and output three-dimensional key point coordinates. The calculation network is used to solve the head pose according to the two-dimensional key point coordinates and the three-dimensional key point coordinates; The head pose estimation device further includes a training module, configured to input the second training sample set into the feature extraction network to obtain a feature representation of each training sample in the second training sample set; Input the feature representation into the initial first branch network to obtain a set of predicted two-dimensional key point coordinates and the predicted uncertainty corresponding to the predicted two-dimensional key point coordinates, and input the feature representation into the initial second branch network to obtain a set of predicted three-dimensional key point coordinates; Calculate a predicted head pose based on the set of predicted two-dimensional key point coordinates and the set of predicted three-dimensional key point coordinates; Calculate a first loss value according to the set of predicted two-dimensional key point coordinates, the predicted uncertainty, and the set of true two-dimensional key point coordinates, calculate a second loss value according to the set of predicted three-dimensional key point coordinates and the set of true three-dimensional key point coordinates, and calculate a third loss value according to the predicted head pose and the true head pose; Adjust the second initial network model according to the first loss value, the second loss value, and the third loss value to obtain the first network model.

21. The device according to claim 20, wherein, The feature extraction network includes a residual neural network ResNet and a pooling layer. The first branch network is a fully connected layer. The second branch network is a fully connected layer. The calculation network is a differentiable PnP solving network.

22. The device according to any one of claims 12 to 13 or any one of claims 15 to 19 or claim 21, characterized in that, The obtaining module is specifically configured to obtain an image to be processed, where the image to be processed includes a target face image collected by a camera; Input the image to be processed into an image preprocessing network to obtain sparse key points of the target face image. The sparse key points include feature points of the five sense organs and face contour points of the target face image; Obtain the target face image from the image to be processed according to the face contour points; Horizontally align and scale the target face image to a target size according to the feature points of the five sense organs to obtain the image to be recognized.

23. A computer device, characterized in that, Including: A memory, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is used to execute the programs in the memory. The processor is used to execute the method according to any one of claims 1 to 11 according to the instructions in the program code; The bus system is used to connect the memory and the processor so that the memory and the processor can communicate.

24. A computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Face image quality detection method and system, and computer equipment

    CN112200176A

  • Method for acquiring attitude data and neural network construction method

    CN114299152A