Facial expression reconstruction method, device, equipment and storage medium

By using the deep neural network model to capture the single face image, output the expression coefficients and shooting parameters, and selecting the basic expression template for face expression reconstruction, solving the problem of expensive equipment and complex calculations in the existing technology, and achieving efficient and high-precision face expression reconstruction.

CN114202615BActive Publication Date: 2025-08-08GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111503555.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-08-08
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

Existing methods for facial expression reconstruction require expensive depth camera equipment or complex optimization calculations, resulting in complex and inefficient operations.

Method used

A pre-trained deep neural network model is used to capture expressions on a single face image, output expression coefficients and shooting parameters, select basic expression templates based on the expression coefficients and reconstruct face expressions based on the shooting parameters.

Benefits of technology

High-precision facial expression reconstruction is achieved, reducing the amount of calculation, avoiding complex optimization processes, and improving operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114202615B_ABST
    Figure CN114202615B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device and storage medium for reconstructing facial expressions, which obtains a facial image; inputs the facial image into a pre-trained deep neural network model to output the facial image's expression coefficient and shooting parameters; wherein the expression coefficient represents the weight of each basic expression template, the weight value is greater than or equal to zero, and at least one weight value is greater than zero, so at least one basic expression template can be selected from a preset basic expression template library based on the expression coefficient; facial expressions are reconstructed based on the expression coefficient, at least one basic expression template and shooting parameters. This method captures facial expressions from a single facial image based on deep learning. On the one hand, the expression coefficient and shooting parameters obtained are very accurate, so that the results obtained in the subsequent facial expression reconstruction are more accurate. On the other hand, it avoids the need for a large amount of optimization like traditional models, greatly reducing the amount of calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network live broadcast technology, and in particular to a method, apparatus, device and storage medium for reconstructing facial expressions. Background Art

[0002] With the development of computer vision technology, facial expression reconstruction has been widely used in fields including gaming, live streaming, and AR. Current facial expression reconstruction can be roughly divided into two categories, depending on the input data: facial expression capture methods based on depth images and facial expression capture methods based on single images. Depth image-based facial expression capture methods often require expensive depth camera equipment, while single-image-based facial expression reconstruction methods, while simple in equipment, require additional key point detection and complex optimization calculations, making the operation very complex and inefficient. Summary of the Invention

[0003] In view of this, embodiments of the present application provide a method, apparatus, device, and storage medium for reconstructing facial expressions.

[0004] In a first aspect, an embodiment of the present application provides a method for reconstructing facial expressions, the method comprising:

[0005] Get face image;

[0006] Inputting the facial image into a pre-trained deep neural network model to output expression coefficients and shooting parameters of the facial image;

[0007] Selecting at least one basic expression template from a preset basic expression template library according to the expression coefficient;

[0008] Facial expression reconstruction is performed according to the expression coefficient, at least one basic expression template and the shooting parameters.

[0009] In a second aspect, an embodiment of the present application provides a facial expression reconstruction device, the device comprising:

[0010] A face image acquisition module, used to acquire face images;

[0011] A coefficient and parameter output module, configured to input the facial image into a pre-trained deep neural network model to output expression coefficients and shooting parameters of the facial image;

[0012] A template selection module, configured to select at least one basic expression template from a preset basic expression template library according to the expression coefficient;

[0013] A facial expression reconstruction module is used to reconstruct facial expressions based on the expression coefficient, at least one basic expression template and the shooting parameters.

[0014] In a third aspect, an embodiment of the present application provides a terminal device comprising: a memory; one or more processors coupled to the memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the facial expression reconstruction method provided in the first aspect above.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored. The program code can be called by a processor to execute the facial expression reconstruction method provided in the first aspect above.

[0016] The facial expression reconstruction method, apparatus, device and storage medium provided in the embodiments of the present application first obtain a facial image; then input the facial image into a pre-trained deep neural network model to output the expression coefficient and shooting parameters of the facial image; wherein the expression coefficient represents the weight of each basic expression template, the value of the weight is greater than or equal to zero, and at least one weight has a value greater than zero, so at least one basic expression template can be selected from a preset basic expression template library based on the expression coefficient; finally, the facial expression is reconstructed based on the expression coefficient, at least one basic expression template and the shooting parameters.

[0017] This facial expression reconstruction method captures facial expressions from a single face image using deep learning. On the one hand, the expression coefficients and shooting parameters obtained are very accurate, making the results of subsequent facial expression reconstruction more precise. On the other hand, it avoids the need for a lot of optimization like traditional models, greatly reducing the amount of calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0019] Figure 1 A schematic diagram of an application scenario of the facial expression reconstruction method provided in an embodiment of the present application;

[0020] Figure 2 A flowchart of a method for reconstructing facial expressions provided in one embodiment of the present application;

[0021] Figure 3 A schematic diagram of the structure after facial expression reconstruction is provided for one embodiment of the present application;

[0022] Figure 4 A schematic diagram of the Blendshape structure provided for one embodiment of the present application;

[0023] Figure 5 A schematic diagram of a facial image sample structure provided for one embodiment of the present application;

[0024] Figure 6 A schematic diagram of a structure for selecting a curve on a facial contour provided in one embodiment of the present application;

[0025] Figure 7 A structural diagram of a facial expression reconstruction device provided by one embodiment of the present application;

[0026] Figure 8 This is a schematic diagram of the structure of a terminal device provided in one embodiment of the present application;

[0027] Figure 9 A schematic diagram of the structure of a computer-readable storage medium provided in one embodiment of the present application. DETAILED DESCRIPTION

[0028] The following is a clear and complete description of the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0029] In order to explain the present application in more detail, the following specifically describes a facial expression reconstruction method, apparatus, terminal device and computer storage medium provided by the present application in conjunction with the accompanying drawings.

[0030] Please refer to Figure 1 , Figure 1A schematic diagram illustrates an application scenario for the facial expression reconstruction method provided by an embodiment of the present application. The application scenario includes a server 102, a live broadcast terminal 104, and a client 106 provided by an embodiment of the present application. A network is provided between the server 102, the live broadcast terminal 104, and the client 106. The network serves as a medium for providing a communication link between the server 102, the live broadcast terminal 104, and the client 106. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The server 102 can communicate with the live broadcast terminal 104 and the client 106 to provide live broadcast services to the live broadcast terminal 104 and / or the client 106. For example, the live broadcast terminal 104 can send a live video stream from a live broadcast room to the server 102, and a user can access the server 102 through the client 106 to view the live video from the live broadcast room. For another example, the server 102 can also send a notification message to the user's client 106 when a live broadcast room to which a user has subscribed begins broadcasting. The live video stream can be a video stream currently being broadcast on the live broadcast platform or a complete video stream formed after the live broadcast is completed.

[0031] In some implementations, the live broadcast terminal 104 and the client 106 can be used interchangeably. For example, a host can use the live broadcast terminal 104 to provide live video services to viewers, or as a user, to view live videos provided by other hosts. For another example, a user can use the client 106 to watch live videos provided by a host they follow, or as a host, to provide live video services to other viewers.

[0032] In this embodiment, the live broadcast terminal 104 and the client 106 are both terminals, which can be various electronic devices with display screens, including but not limited to smartphones, personal digital assistants, tablet computers, personal computers, laptop computers, virtual reality terminal devices, augmented reality terminal devices, etc. The live broadcast terminal 104 and the client 106 can be installed with an Internet product for providing Internet live broadcast services. For example, the Internet product can be an application APP, web page, mini-program, etc. related to Internet live broadcast services used in computers or smartphones.

[0033] I understand. Figure 1 The application scenario shown is only a feasible example. In other feasible embodiments, the application scenario may also include only Figure 1 The components shown may also include other components. For example, Figure 1 The application scenario shown may also include a video acquisition terminal 108 for acquiring live video frames of the anchor. The video acquisition terminal 108 may be directly installed or integrated into the live broadcast terminal 104, or may be independent of the live broadcast terminal 104, etc. This embodiment does not impose any restrictions here.

[0034] It should be understood that the number of live broadcast terminals 104, clients 106, networks, and servers 102 is merely illustrative. Any number of live broadcast terminals 104, clients 106, networks, and servers 102 may be provided as needed. For example, a server may be a server cluster consisting of multiple servers. The live broadcast terminals 104 and clients 106 interact with the server via the network to receive or send messages, etc. The server 102 may be a server that provides various services. The live broadcast terminals 104 or clients 102 may be used to perform the steps of a facial expression reconstruction method provided in the embodiments of the present application.

[0035] Based on this, a method for reconstructing facial expressions is provided in the embodiment of the present application. Figure 2 , Figure 2 A schematic diagram showing a process flow of a facial expression reconstruction method provided by an embodiment of the present application is shown. Figure 1 The live broadcast end in the example is used to illustrate, including the following steps:

[0036] Step S110: Acquire a face image.

[0037] A facial image is an image containing facial data. Facial data is data rich in facial features (e.g., facial shape, facial expression, and shooting angle). Analyzing facial data can reveal facial shape, facial expression, and shooting angle characteristics.

[0038] Alternatively, the facial image can be a facial image of any user whose expression needs to be reconstructed. In live broadcasts, especially virtual live broadcasts, the facial image is usually the host's facial image. The host's facial image can be analyzed to capture the host's expression and reconstruct the host's expression on the virtual host to complete the host's expression drive on the virtual host, even if the virtual host makes the same expression as the host. For details, please refer to Figure 3 shown.

[0039] In an optional implementation, the facial image is usually a 2D picture, which can be a facial photo directly captured by a camera; or a facial photo extracted from a video captured by a video capture terminal.

[0040] In this embodiment, a single facial image is typically used, meaning that the user's expression can be reconstructed from a single facial image. However, during live broadcasts, when the host's facial expression needs to be reconstructed, the host's facial image can be captured in real time or at a scheduled time to achieve real-time reconstruction of the host's facial expression. The captured facial images may have different shooting angles, lighting, colors, and expressions at different times.

[0041] Step S120: input the facial image into a pre-trained deep neural network model to output the expression coefficient and shooting parameters of the facial image.

[0042] Deep Neural Networks (DNNs) are discriminative models with at least one hidden layer, trained using the backpropagation algorithm. Based on their location, the neural network within a DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and all layers in between are hidden layers. Layers are fully connected, meaning that any neuron in layer i is connected to any neuron in layer i+1.

[0043] In this embodiment, the deep neural network model can be a deep convolutional neural network model (DCNN). Convolutional neural network models have achieved good results in feature recognition related tasks and are commonly used in image recognition and speech recognition. Optional deep convolutional neural network models can be a residual network variant structure based on deep residual network optimization, a residual network variant structure using a new training method, a residual network variant structure based on increased width, and a residual network variant structure using a new dimension.

[0044] The essence of model training is to give an input vector and a target output value, then input the input vector into one or more network structures or functions to obtain the actual output value, and calculate the deviation based on the target output value and the actual output value, and determine whether the deviation is within the allowable range; if it is within the allowable range, the training ends and the relevant parameters are fixed; if it is not within the allowable range, some parameters in the network structure or function are continuously adjusted until the deviation is within the allowable range or a certain end condition is reached, the training ends and the relevant parameters are fixed, and finally the trained model can be obtained based on the fixed relevant parameters. In this embodiment, facial image samples with different expressions and postures (i.e., shooting angles) are mainly used to input into the deep neural network model, calculate its loss function, and then update the network parameters of the deep neural network model until the network converges, thereby obtaining a trained deep neural network model.

[0045] Deep neural network models have powerful nonlinear fitting capabilities, strong feature extraction capabilities, strong ability to process high-dimensional data, and strong error recognition capabilities. Even if some neurons are damaged, it will not have a significant impact on the overall training results. Therefore, in the embodiments of this application, using a trained deep neural network model to extract features from facial images can obtain more and more accurate features, thereby making the output expression coefficient and shooting coefficient more accurate.

[0046] The expression coefficient refers to data or information related to facial expressions and is often used to fit facial expressions. Facial expressions include anger, surprise, fear, disgust, happiness, sadness, and so on. Currently, there are up to 52 commonly used expressions. Alternatively, the expression coefficient can be a data set, where each data point in the set represents the weight of its corresponding expression. For example, suppose there are 52 expressions: anger, surprise, fear, disgust, happiness, sadness, etc., and the expression coefficients are 1, 0.2, 0, 0, 0, 0, ... 0. Then, the weight of anger is 1, the weight of surprise is 0.2, and the weights of all other expressions are 0.

[0047] Shooting parameters refer to some shooting equipment related parameters when shooting facial images, including but not limited to camera rotation parameters and translation parameters.

[0048] Step S130 : selecting at least one basic expression template from a preset basic expression template library according to the expression coefficient.

[0049] Specifically, the basic expression template is also called the basic expression base image, also known as Blendshape (deformation target or expression deformation). Blendshape can be used to control facial expression details. Generally, a face can usually set several or dozens of Blendshapes, that is, a person's facial expression can be composed by combining several or dozens of Blendshapes. Each Blendshape usually only controls one facial detail. For example, the eyes, mouth, corners of the mouth, eyebrows, nose, etc. can be controlled by different Blendshapes respectively. The value range of each Blendshape can be 0-1. For example, a Blendshape that controls the mouth, if the Blendshape is 0, it means the mouth is tightly closed, and when the Blendshape is 1, it means the mouth is fully open. When the Blendshape is other values, it indicates the degree of mouth opening. Based on this, several or dozens of Blendshapes can be combined to form very complex facial expressions.

[0050] In an optional implementation, each expression can generate a corresponding Blendshape, so the number of Blendshapes can be multiple, for example, 52. For details, please refer to Figure 4 As shown, Figure 3 Only some BlendShapes are shown in the figure. Multiple BlendShapes can be combined into a set, which is recorded as a preset basic expression template library.

[0051] Since the expression coefficient represents the weight of each expression (i.e., blendshape), selecting at least one basic expression template from the preset basic expression template library based on the expression coefficient essentially involves determining the corresponding value of each blendshape. When a blendshape is greater than 0, it is extracted from the preset basic expression template library. When a blendshape is equal to 0, it can be ignored. This method is used to find all relevant blendshapes.

[0052] Step S140 , reconstructing facial expressions based on the expression coefficients, at least one basic expression template, and shooting parameters.

[0053] After finding all the basic expression templates (i.e., Blendshape), the expression coefficients, shooting parameters, and all the basic expression templates are merged to form a 3D face model, and the facial expression reconstruction is completed.

[0054] In one embodiment, when executing step S140, the shooting parameters include rotation parameters and translation parameters; facial expression reconstruction is performed based on the expression coefficient, at least one basic expression template and the shooting parameters, including: performing face synthesis based on the expression coefficient and at least one basic expression model to obtain an initial facial expression model; adjusting the initial facial expression model based on the rotation parameters and translation parameters to form a final facial expression model to complete facial expression reconstruction.

[0055] Specifically, all basic expression templates (i.e., Blendshape) are found, and then all basic expression templates are merged according to the weight corresponding to each basic expression template to form a 3D face model. The 3D face model is then adjusted according to the rotation parameters and translation parameters to complete the facial expression reconstruction.

[0056] For ease of understanding, a detailed example is given. Assume that the expression coefficient is W exp , camera rotation parameter R, camera translation parameter T, the initial facial expression model is V = B × W exp T , the final facial expression model V′=R×V+T. Where B∈R n×3×52 Indicates 52 Blendshapes, n is the number of vertices of 3D facial expressions, 3 represents the 3D coordinates [x, y, z], 52 corresponds to 52 expressions, W exp It can be understood as the weight of each expression. After obtaining the initial facial expression model V, the face posture can be controlled by V′=R×V+T so that the estimated face posture is consistent with the face posture in the input face image.

[0057] In another optional embodiment, during facial expression reconstruction based on the expression coefficient, at least one basic expression template, and shooting parameters, the face shape of the input facial image may also be comprehensively considered. Specifically, the shape coefficient of the input facial image may be obtained, and then facial expression reconstruction may be performed based on the shape coefficient, expression coefficient, at least one basic expression template, and shooting parameters. Considering the facial shape during facial expression reconstruction can make the reconstructed face shape more similar to the face shape in the input facial image, further improving the accuracy of facial expression reconstruction.

[0058] The facial expression reconstruction method, apparatus, device and storage medium provided in the embodiments of the present application first obtain a facial image; then input the facial image into a pre-trained deep neural network model to output the expression coefficient and shooting parameters of the facial image; wherein the expression coefficient represents the weight of each basic expression template and at least one weight is a positive value, so at least one basic expression template can be selected from a preset basic expression template library based on the expression coefficient; finally, the facial expression is reconstructed based on the expression coefficient, at least one basic expression template and the shooting parameters.

[0059] This facial expression reconstruction method captures facial expressions from a single face image using deep learning. On the one hand, the expression coefficients and shooting parameters obtained are very accurate, making the results of subsequent facial expression reconstruction more precise. On the other hand, it avoids the need for a lot of optimization like traditional models, greatly reducing the amount of calculation.

[0060] Furthermore, a specific implementation method of model training is given, which is described as follows:

[0061] In one embodiment, the pre-trained deep neural network model is obtained by:

[0062] Step S1, obtaining a face image sample.

[0063] Specifically, first, a relatively large number (e.g., thousands or tens of thousands) of facial image samples must be prepared. These facial image samples can be collected using a camera. Generally, the more image samples, the more accurate the trained model; however, too many facial image samples can slow down model training. Therefore, in practical applications, selecting an appropriate number of facial image samples is sufficient. However, when preparing facial image samples, the samples should be as diverse as possible, including images with a variety of expressions and from a variety of shooting angles. For example, facial image samples should include both normal poses and expressions, as well as poses with the face in profile, looking down, looking up, and various expressions such as eyes open and closed, and a crooked mouth. Furthermore, when preparing facial image samples, a data training set can be created and the facial image samples stored in the data training set.

[0064] Step S2: annotate facial key points on the facial image sample to obtain a first facial key point set.

[0065] After obtaining the face image sample, the face image sample is annotated with the face key points. For each face image sample, many key points can be annotated, for example, 278 key points can be annotated. For details, please refer to Figure 5 As shown in , these facial key points form the first facial key point set.

[0066] Step S3: Input the facial image sample with key points annotated into a deep neural network model to output shape coefficients, expression coefficients and shooting parameters.

[0067] Specifically, we first need to build a deep neural network model. We can use a convolutional neural network as the backbone network in the model, and then add four fully connected layers after the backbone network. The four fully connected layers are used to output the shape coefficient W. id , expression coefficient W exp , shooting parameters (i.e., camera rotation parameter R and translation parameter T). Optionally, the convolutional neural network can be mobilenetv3, resnet, or densenet, etc.

[0068] Among them, expression coefficient refers to some data or information related to facial expressions, which is often used to fit facial expressions. Shooting parameters refer to some parameters related to the shooting equipment when shooting facial images, including but not limited to camera rotation parameters and translation parameters. Shape coefficient refers to some data or information related to the face shape of the face, which is often used to fit the face shape of the face. Since the facial image samples include facial images of many different people, each person's face shape will be different, and different face shapes usually differ when expressing the same expression, that is, the face shape may interfere with the expression of the expression. Therefore, when training the deep neural network model, such factors are fully considered, that is, the shape coefficient related to the face shape is selected to train the deep neural network model, so that the trained model is more accurate, so that the expression coefficient extracted in the subsequent reconstruction of the facial expression is more accurate, further improving the accuracy of facial expression reconstruction.

[0069] Step S4: Reconstruct the face according to the shape coefficient, expression coefficient, shooting parameters and the three-dimensional deformation model to obtain a three-dimensional face model.

[0070] The 3D deformable model is a basic model used to perform 3D reconstruction of a human face. Optionally, the 3D deformable model can be Facewarehouse (a multilinear face model), which is a 3D deformable face model commonly used in the fields of computer vision and computer graphics.

[0071] The specific process is: shape coefficient W id ∈R150×1 , expression coefficient W exp ∈R 52×1 , camera rotation parameter R∈R 3×3 , translation parameter T∈R 1×3 and the three-dimensional deformable model Cr∈R n×3×52×150 Perform face synthesis, that is, V1 = Cr × W id T ×W exp T ; V1′=R×V+T, where Cr refers to the three-dimensional deformation model, V1∈R n×3 is the coordinate representation of the vertex set of the face unit model before processing according to the shooting parameters, V1′∈R n×3 It is a three-dimensional face model after posture control such as rotation and translation is performed through camera parameters (i.e. shooting parameters).

[0072] Step S5: Select facial key points on the three-dimensional face model, and project the selected facial key points onto a two-dimensional pixel plane to form a second facial key point set.

[0073] After obtaining the 3D facial model, facial key points are selected on the 3D facial model to obtain a plurality of facial key points. The number of facial key points selected from the 3D facial model is equal to the number of facial key points in the input facial image sample. After selecting the facial key points, the selected key points can be projected onto a 2D pixel plane P using perspective projection to generate a second facial key point set.

[0074] Optionally, when selecting key points in a 3D facial model, the selection of key points is mainly divided into two steps: one step is to select from fixed facial features, and the other step is to select from the facial contour. The specific process is to select multiple facial key points from different locations in the facial features and the facial contour of the 3D facial model.

[0075] The five facial features are primarily the nose, eyes, mouth, and eyebrows. Key points on these features are typically selected from relatively fixed locations, such as the corners of the mouth and eyelids. However, key points within the facial contour must be selected from different locations on the face, referencing the positions of the facial key in the input facial image sample. This means that the position of the second facial key point essentially corresponds to the position of the first.

[0076] Furthermore, an implementation method for selecting facial key points from a facial contour is provided, which is described in detail below.

[0077] In one embodiment, a plurality of facial key points are selected from different positions in a facial contour of a three-dimensional face model, including: randomly selecting a plurality of curves on the facial contour of the three-dimensional face model, wherein the curves are distributed at different positions of the facial contour and the curves do not intersect with each other; selecting one or more points from each curve as candidate key points; and selecting a plurality of facial key points from the candidate key points.

[0078] Specifically, first select some curves on the 3D face model, such as Figure 6 As shown in the figure, we select horizontal or vertical curves on each cheek, and use the points on each curve as candidate cheek keypoints for different poses. Similarly, we select curves from the mouth to the chin, and use the points on each curve as candidate chin keypoints for different poses. When selecting curves, we try to avoid intersecting them. Intersecting curves can lead to duplicate selected keypoints, which can lead to inaccurate keypoint selection.

[0079] After obtaining multiple curves, when selecting candidate key points from the curves, there are different selection methods for curves on the cheeks and chin. The specific methods are as follows:

[0080] In one embodiment, one or more points are selected from each curve as candidate key points, including: when the curve is located on the cheek, the point with the largest absolute value of the horizontal coordinate is selected from the curve as the candidate key point; and / or; when the curve is located on the chin, the point with the largest sum of the square value of the horizontal coordinate and the square value of the vertical coordinate is selected from the curve as the candidate key point.

[0081] Specifically, the point with the largest absolute value of x on the left and right cheek curves can be selected as a candidate key point; the point with the largest absolute value of x on the chin curve can be selected as a candidate key point. 2 +y 2 The largest point is taken as a candidate key point.

[0082] The above-mentioned candidate key point selection method is adopted, that is, a dynamic selection algorithm is used for facial contour key points, which solves the problem that facial contour key points are not fixed on the 3D model under different postures during model training, which easily leads to incorrect selection.

[0083] In addition, since the selected curve may be sparse and uneven, the candidate key points selected based on this curve are also sparse and uneven, and the number does not match. Based on this, a fixed number of uniform facial contour key points are obtained through interpolation. The interpolation process is as follows:

[0084] In one embodiment, selecting one or more points from each curve as candidate key points includes: performing an interpolation operation on each curve to obtain a plurality of interpolation points; and using the plurality of interpolation points as candidate key points.

[0085] Specifically, for any curve, the distance between two adjacent points on the curve is first calculated, which is recorded as D = {d0, d1, ..., d n}, where d0 is the distance between the first and second points on the curve, and so on, dn is the distance between the last n points and the n-1th point before it. Then calculate the cumulative distance along one direction of the curve Where d0 is the cumulative distance between the first and second points on the curve, di is the distance between the first and second points on the curve, plus the distance from the first point to the third point, and finally the distance from the first point to the i-th point to form the cumulative distance. The number of interpolation points is determined based on the cumulative distance, and then the distance between the interpolation points is calculated based on the curve length and the number of interpolation points. Then according to the cumulative distance of the interpolation points Calculate which two points P the interpolation point P′ falls on i , P i+1 Between; calculate the interpolation point to P' to P i , P i+1 The distance between i , w i+1 ; Calculate the coordinates of the interpolation points

[0086] Step S6: determining a mean square loss function based on the first facial key point set and the second facial key point set.

[0087] Specifically, the mean square loss function, MSELoss, is calculated based on the facial key points in the second facial key point set and the facial key points in the first facial key point set:

[0088]

[0089] Among them, N represents the number of facial key points, Pj is the jth key point selected on the face 3D model V1′, and P′ j It refers to the jth face key point in the first face key set, that is, the jth face key point on the input face image sample.

[0090] Step S7: performing affine transformation on the facial image samples after key point annotation, and determining a third facial key point set based on the simulated transformed facial image samples.

[0091] Specifically, assuming that the face image sample after key point annotation is I, it is subjected to the affine transformation of the M matrix to obtain I′. If the jth key point on the face 3D model V1′ constructed using the face image sample I is recorded as P Ij , then P IjAfter the affine transformation of the M matrix, we get P I′j , P I′j Recorded as the third face key point, each third face key point is collected together to form a third face key point set.

[0092] Step S8: determining a consistency loss function based on the first facial key point set, the second facial key point set, and the third facial key point set.

[0093] Specifically, the expression of the consistency loss function is: , Among them, P Ij It refers to the jth key point on the face 3D model V1′ constructed using the face image sample I, which is recorded as P Ij , P′ Ij Refers to the jth facial key point on the input face image sample; P′ I′j means that P Ij The j-th facial key point is obtained through the affine transformation of the M matrix.

[0094] Step S9: determining a total loss function based on the mean square loss function and the consistency loss function.

[0095] Step S10: updating the network parameters of the deep neural network model according to the total loss function until convergence to obtain a pre-trained deep neural network model; wherein the number of facial key points in the first facial key point set, the second facial key point set, and the third facial key point set is the same.

[0096] Among them, the total loss function consists of the above two parts: that is, L = L 2d +L consistency . Calculate the total loss function, and update the network parameters of the deep neural network model based on the total loss function combined with the back propagation algorithm until convergence to obtain a pre-trained deep neural network model.

[0097] It should be understood that although Figure 2 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Figure 2 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0098] The above embodiments disclosed in this application describe in detail a method for reconstructing facial expressions. The above method disclosed in this application can be implemented using various forms of equipment. Therefore, this application also discloses a facial expression reconstruction device corresponding to the above method. Specific embodiments are given below for detailed description.

[0099] See also Figure 7 , is a facial expression reconstruction device disclosed in an embodiment of the present application, mainly comprising:

[0100] The face image acquisition module 710 is used to acquire a face image.

[0101] The coefficient and parameter output module 720 is used to input the facial image into a pre-trained deep neural network model to output the expression coefficient and shooting parameters of the facial image.

[0102] The template selection module 730 is configured to select at least one basic expression template from a preset basic expression template library according to the expression coefficient.

[0103] The facial expression reconstruction module 740 is used to reconstruct facial expressions based on the expression coefficient, at least one basic expression template and shooting parameters.

[0104] In one embodiment, the apparatus further comprises:

[0105] The sample acquisition module is used to obtain face image samples.

[0106] A first key point set obtaining module is used to annotate facial key points of a facial image sample to obtain a first facial key point set;

[0107] The parameter output module is used to input the facial image samples with key points annotated into the deep neural network model to output shape coefficients, expression coefficients and shooting parameters.

[0108] The 3D face model reconstruction module is used to reconstruct the face according to the shape coefficient, expression coefficient, shooting parameters and 3D deformation model to obtain a 3D face model.

[0109] The second key point set acquisition module is used to select facial key points on the three-dimensional face model and project the selected facial key points onto the two-dimensional pixel plane to form a second facial key point set.

[0110] The mean square loss function calculation module is used to determine the mean square loss function based on the first face key point set and the second face key point set.

[0111] The third key point set obtaining module is used to perform affine transformation on the face image samples after key point annotation, and determine the third face key point set based on the simulated transformed face image samples.

[0112] The consistency loss function calculation module is used to determine the consistency loss function according to the first face key point set, the second face key point set and the third face key point set.

[0113] The total loss function determination module is used to determine the total loss function based on the mean square loss function and the consistency loss function.

[0114] The model acquisition module is used to update the network parameters of the deep neural network model according to the total loss function until convergence to obtain a pre-trained deep neural network model; wherein the number of facial key points in the first facial key point set, the second facial key point set and the third facial key point set is the same.

[0115] In one embodiment, the second key point set obtaining module is used to select multiple facial key points from different positions in the facial features and facial contour of the three-dimensional face model.

[0116] In one embodiment, the second key point set acquisition module is used to randomly select multiple curves on the facial contour of the three-dimensional face model, where each curve is distributed at a different position of the facial contour and does not intersect with each other; select one or more points from each curve as candidate key points; and select multiple facial key points from the candidate key points.

[0117] In one embodiment, the second key point set acquisition module is used to select the point with the largest absolute value of the horizontal coordinate from the curve as a candidate key point when the curve is located on the cheek; and / or; when the curve is located on the chin, select the point with the largest sum of the square value of the horizontal coordinate and the square value of the vertical coordinate from the curve as a candidate key point.

[0118] In one embodiment, the second key point set obtaining module is configured to perform an interpolation operation on each curve to obtain a plurality of interpolation points; and use the plurality of interpolation points as candidate key points.

[0119] In one embodiment, the facial expression reconstruction module 740 is used to synthesize a face based on the expression coefficient and at least one basic expression model to obtain an initial facial expression model; adjust the initial facial expression model according to the rotation parameter and the translation parameter to form a final facial expression model to complete the facial expression reconstruction.

[0120] The specific definitions of the facial expression reconstruction device can be found in the definitions of the method above and will not be repeated here. Each module in the above-mentioned device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the terminal device in hardware form, or can be stored in the memory of the terminal device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0121] Please refer to Figure 8 , Figure 8 The present invention provides a block diagram of a terminal device according to an embodiment of the present invention. The terminal device 80 may be a computer device. The terminal device 80 in the present invention may include one or more of the following components: a processor 82, a memory 84, and one or more application programs. The one or more application programs may be stored in the memory 84 and configured to be executed by the one or more processors 82. The one or more application programs are configured to execute the method described in the embodiment of the method for reconstructing facial expressions.

[0122] The processor 82 may include one or more processing cores. The processor 82 utilizes various interfaces and lines to connect the various parts of the entire terminal device 80, and executes various functions and processes data of the terminal device 80 by running or executing instructions, programs, code sets or instruction sets stored in the memory 84, and calling data stored in the memory 84. Optionally, the processor 82 may be implemented in the form of at least one hardware of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 82 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU) for reporting and validating buried data, and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to handle wireless communications. It is understandable that the above-mentioned modem may not be integrated into the processor 82, but may be implemented separately through a communication chip.

[0123] The memory 84 may include a random access memory (RAM) or a read-only memory (ROM). The memory 84 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 84 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the various method embodiments described below, and the like. The data storage area may also store data created by the terminal device 80 during use.

[0124] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the terminal device to which the solution of the present application is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0125] In summary, the terminal device provided in the embodiment of the present application is used to implement the corresponding facial expression reconstruction method in the aforementioned method embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0126] See also Figure 9 , which shows a block diagram of a computer-readable storage medium provided in an embodiment of the present application. The computer-readable storage medium 90 stores program code, which can be called by a processor to execute the method described in the above embodiment of the method for reconstructing facial expressions.

[0127] The computer-readable storage medium 90 may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium 90 may include a non-transitory computer-readable storage medium. The computer-readable storage medium 90 has storage space for program code 92 for executing any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 92 may be compressed, for example, in a suitable form.

[0128] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0129] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for reconstructing facial expressions, characterized in that: The method comprises: Get face image; Inputting the facial image into a pre-trained deep neural network model to output expression coefficients and shooting parameters of the facial image; Selecting at least one basic expression template from a preset basic expression template library according to the expression coefficient; Reconstructing facial expressions according to the expression coefficient, at least one basic expression template, and the shooting parameters; The pre-trained deep neural network model is obtained by the following method: Obtain face image samples; Annotating facial key points of the facial image sample to obtain a first facial key point set; Input the facial image samples with key points annotated into the deep neural network model to output shape coefficients, expression coefficients and shooting parameters; Reconstructing the face according to the shape coefficient, the expression coefficient, the shooting parameters and the three-dimensional deformation model to obtain a three-dimensional face model; Selecting facial key points on the three-dimensional face model, and projecting the selected facial key points onto a two-dimensional pixel plane to form a second facial key point set; Determine a mean square loss function based on the first facial key point set and the second facial key point set; Performing affine transformation on the facial image samples after key point annotation, and determining a third facial key point set based on the simulated transformed facial image samples; Determining a consistency loss function based on the first face key point set, the second face key point set, and the third face key point set; Determine a total loss function based on the mean square loss function and the consistency loss function; Updating the network parameters of the deep neural network model according to the total loss function until convergence to obtain a pre-trained deep neural network model; The first facial key point set, the second facial key point set and the third facial key point set have the same number of facial key points.

2. The method according to claim 1, characterized in that The selecting of facial key points on the three-dimensional face model comprises: A plurality of facial key points are respectively selected from the facial features of the three-dimensional face model and different positions in the facial contour.

3. The method according to claim 2, characterized in that The step of selecting a plurality of facial key points from different positions in the facial contour of the three-dimensional face model includes: Randomly selecting a plurality of curves on the facial contour of the three-dimensional face model, wherein the curves are distributed at different positions of the facial contour and do not intersect with each other; Selecting one or more points from each of the curves as candidate key points; A plurality of facial key points are selected from the candidate key points.

4. The method according to claim 3, characterized in that The selecting one or more points from each of the curves as candidate key points includes: When the curve is located on the cheek, selecting a point with the largest absolute value of the horizontal coordinate from the curve as the candidate key point; and / or; When the curve is located on the chin, a point with the largest sum of the square value of the abscissa and the square value of the ordinate is selected from the curve as the candidate key point.

5. The method according to claim 4, characterized in that The selecting one or more points from each of the curves as candidate key points includes: Perform interpolation operations on each curve to obtain multiple interpolation points; The plurality of interpolation points are used as the candidate key points.

6. The method according to any one of claims 1 to 5, characterized in that The shooting parameters include rotation parameters and translation parameters; and the facial expression reconstruction based on the expression coefficient, at least one basic expression template and the shooting parameters includes: Performing face synthesis based on the expression coefficient and at least one basic expression template to obtain an initial facial expression model; The initial facial expression model is adjusted according to the rotation parameter and the translation parameter to form a final facial expression model, so as to complete facial expression reconstruction.

7. A facial expression reconstruction device, characterized in that: The device comprises: A face image acquisition module, used to acquire face images; A coefficient and parameter output module, configured to input the facial image into a pre-trained deep neural network model to output expression coefficients and shooting parameters of the facial image; A template selection module, configured to select at least one basic expression template from a preset basic expression template library according to the expression coefficient; A facial expression reconstruction module, configured to reconstruct facial expressions based on the expression coefficient, at least one basic expression template, and the shooting parameters; The device further comprises: A sample acquisition module is used to obtain face image samples; A first key point set obtaining module, configured to annotate the facial key points of the facial image sample to obtain a first facial key point set; The parameter output module is used to input the facial image samples with key points annotated into the deep neural network model to output the shape coefficient, expression coefficient and shooting parameters; A three-dimensional face model reconstruction module, configured to reconstruct the face according to the shape coefficient, the expression coefficient, the shooting parameters, and the three-dimensional deformation model to obtain a three-dimensional face model; a second key point set obtaining module, configured to select facial key points on the three-dimensional face model and project the selected facial key points onto a two-dimensional pixel plane to form a second facial key point set; a mean square loss function calculation module, configured to determine a mean square loss function based on the first facial key point set and the second facial key point set; A third key point set acquisition module is used to perform affine transformation on the facial image samples after key point annotation, and determine the third facial key point set in the simulated transformed facial image samples; a consistency loss function calculation module, configured to determine a consistency loss function based on the first face key point set, the second face key point set, and the third face key point set; A total loss function determination module, used to determine the total loss function based on the mean square loss function and the consistency loss function; A model acquisition module is used to update the network parameters of the deep neural network model according to the total loss function until convergence to obtain a pre-trained deep neural network model; The first facial key point set, the second facial key point set and the third facial key point set have the same number of facial key points.

8. A terminal device, characterized in that: include: Memory; one or more processors coupled to the memory; One or more applications, wherein the one or more applications are stored in a memory and configured to be executed by one or more processors, and the one or more applications are configured to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Real-time automatic high-quality three-dimensional face reconstruction method based on single face image

    CN107358648A