3D Facial Model Reconstruction Method, Device, Storage Medium, and Computer Equipment

Through the face reconstruction model trained based on deep learning, three-dimensional face reconstruction is carried out on two-dimensional images, which solves the problems of low accuracy and lack of details in the existing technology, and achieves stable and rich details of the three-dimensional face model reconstruction.

CN114037802BActive Publication Date: 2025-05-30GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111409845.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2025-05-30
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

The existing three-dimensional face reconstruction technology lacks image prior feature information, resulting in strong model sense, low accuracy and lack of details in the reconstruction results, especially in the case of large expressions or postures.

Method used

A deep learning-based method is used to perform three-dimensional face reconstruction on two-dimensional images through a pre-trained face reconstruction model. The model is trained based on multiple cost functions to generate a textured three-dimensional face model containing texture information.

Benefits of technology

A textured three-dimensional face model with stable effects, rich details and realistic texture information is realized based on two-dimensional images, improving the effect and detailed performance of three-dimensional face reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114037802B_ABST
    Figure CN114037802B_ABST
Patent Text Reader

Abstract

The present application discloses a three-dimensional face model reconstruction method, device, storage medium, and computer device. The method includes: inputting a target two-dimensional image into a trained face reconstruction model to output a corresponding target three-dimensional face model and a target texture map, and then generating a textured three-dimensional face model containing texture information based on the three-dimensional face model and the target texture map. The face reconstruction model is trained based on a cost function constructed from a predicted three-dimensional face model, a standard three-dimensional face model, a smoothed predicted three-dimensional face model, face key points in a predicted two-dimensional image, face key points in a sample two-dimensional image, a predicted texture map, and a standard texture map. By adopting the embodiments of the present application, a pre-trained face reconstruction model can be used to reconstruct a textured three-dimensional face model with rich details and stable effects that contains texture information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computer device applications, and particularly relates to a three-dimensional face model reconstruction method, apparatus, storage medium, and computer device. Background Art

[0002] Three-dimensional face model reconstruction refers to reconstructing a three-dimensional model of a human face from one or more two-dimensional face images. It has one more dimension than two-dimensional face images and is widely used in fields such as movies and games. Three-dimensional face reconstruction technology can restore the three-dimensional shape of a human face in two-dimensional face images, has high application value in fields such as animation production and online games, and has broad application prospects. Summary of the Invention

[0003] Embodiments of this application provide a three-dimensional face model reconstruction method, apparatus, storage medium, and computer device. The three-dimensional face model reconstructed using a pre-trained face reconstruction model has an outstanding smoothing effect. The technical solution is as follows:

[0004] In a first aspect, an embodiment of this application provides a three-dimensional face model reconstruction method, characterized in that the method includes:

[0005] Obtain a target two-dimensional image;

[0006] Input the target two-dimensional image into a trained face reconstruction model, and output a target three-dimensional face model and a target texture map corresponding to the target two-dimensional image;

[0007] Generate a texture three-dimensional face model containing texture information based on the target three-dimensional face model and the target texture map.

[0008] In a second aspect, an embodiment of this application provides a three-dimensional face model reconstruction apparatus, where the target three-dimensional face model reconstruction apparatus includes:

[0009] An image acquisition module, configured to obtain a target two-dimensional image;

[0010] A model prediction module, configured to input the target two-dimensional image into a trained face reconstruction model, and output a target three-dimensional face model and a target texture map corresponding to the target two-dimensional image;

[0011] A target model generation module, configured to generate a texture three-dimensional face model containing texture information based on the target three-dimensional face model and the target texture map.

[0012] In a third aspect, an embodiment of this application provides a storage medium, where the storage medium stores at least one instruction, and the at least one instruction is adapted to be loaded and executed by a processor to perform the above method steps.

[0013] In a fourth aspect, an embodiment of the present application provides a computer device, which may include: a processor and a memory; wherein, the memory stores at least one instruction, and the at least one instruction is adapted to be loaded and executed by the processor to perform the above method steps.

[0014] The beneficial effects brought by the technical solutions provided in some embodiments of the present application at least include:

[0015] Using the three-dimensional face model reconstruction method provided in the embodiment of the present application, input the target two-dimensional image into the trained face reconstruction model, output the corresponding target three-dimensional face model and target texture map, and then generate a texture three-dimensional face model containing texture information based on the target three-dimensional face model and target texture map. Among them, the face reconstruction model is trained and generated based on the first cost function constructed by the predicted three-dimensional face model corresponding to the sample two-dimensional image and the standard three-dimensional face model corresponding to the sample two-dimensional image, the second cost function constructed by the predicted three-dimensional face model and the smoothed predicted three-dimensional face model after smoothing the predicted three-dimensional face model, the third cost function constructed by the face key points after two-dimensional projection of the predicted three-dimensional face model and the face key points in the sample two-dimensional image, and the fourth cost function constructed by the predicted texture map corresponding to the sample two-dimensional image and the standard texture map corresponding to the sample two-dimensional image. The predicted three-dimensional face model is obtained by predicting the sample two-dimensional image using the created initial face reconstruction model, and the standard three-dimensional face model is obtained by reconstructing and constraining the average face three-dimensional model based on the face key points in the sample two-dimensional image. By using the embodiment of the present application, a face reconstruction model pre-trained based on multiple cost functions can reconstruct a stable texture three-dimensional face model containing texture information based on a two-dimensional image. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0017] Figure 1 It is a schematic flowchart of a three-dimensional face model reconstruction method provided by an embodiment of the present application;

[0018] Figure 2 It is an example schematic diagram of three-dimensional face reconstruction provided by an embodiment of the present application;

[0019] Figure 3 It is a schematic flowchart of a three-dimensional face model reconstruction method provided by an embodiment of the present application;

[0020] Figure 4 This is a schematic flowchart of a three-dimensional face model reconstruction method provided by an embodiment of the present application;

[0021] Figure 5 This is a flowchart of reconstructing a standard three-dimensional face model provided by an embodiment of the present application;

[0022] Figure 6 This is a flowchart of generating a standard texture map provided by an embodiment of the present application;

[0023] Figure 7 This is an exemplary schematic diagram of generating a standard texture map provided by an embodiment of the present application;

[0024] Figure 8 This is an exemplary schematic diagram of predicting a texture map provided by an embodiment of the present application;

[0025] Figure 9 This is a schematic flowchart of a three-dimensional face model reconstruction method provided by an embodiment of the present application;

[0026] Figure 10 This is a flowchart of training a face reconstruction model provided by an embodiment of the present application;

[0027] Figure 11 This is a schematic flowchart of a three-dimensional face model reconstruction method provided by an embodiment of the present application;

[0028] Figure 12 This is an exemplary schematic diagram of stylization processing provided by an embodiment of the present application;

[0029] Figure 13 This is an exemplary schematic diagram of reconstructing a standard three-dimensional face model using face key point constraints provided by an embodiment of the present application;

[0030] Figure 14 This is a schematic structural diagram of a three-dimensional face model reconstruction device provided by an embodiment of the present application;

[0031] Figure 15 This is a schematic structural diagram of a three-dimensional face model reconstruction device provided by an embodiment of the present application;

[0032] Figure 16 This is a schematic structural diagram of a model training module provided by an embodiment of the present application;

[0033] Figure 17 This is a schematic structural diagram of a standard acquisition unit provided by an embodiment of the present application;

[0034] Figure 18 This is a schematic structural diagram of a model training unit provided by an embodiment of the present application;

[0035] Figure 19 The embodiment of the present application provides a schematic structural diagram of a computer device. Specific embodiments

[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0037] In the description of the present application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In the description of the present application, it should be noted that unless otherwise clearly specified and limited, "including" and "having", as well as any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood in specific cases. In addition, in the description of the present application, unless otherwise stated, "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0038] To more clearly describe the technical solutions in the embodiments of the present application, before the description, some concepts in the present application are described in detail for better understanding of the solution.

[0039] Cost function: Usually, in order to obtain the parameters of the training logistic regression model, a cost function is required, and the parameters are obtained by training the cost function.

[0040] Stylization: Transforming a two-dimensional image from one style domain to another style domain, such as converting a realistic two-dimensional image into a cartoon two-dimensional image.

[0041] Facial key points: Some important feature point positions such as the eyes, the tip of the nose, the corners of the mouth, the eyebrows, and the contour points of each component of the face.

[0042] Texture map: One or several two-dimensional graphics representing the details of an object's surface. When the texture is mapped onto the object's surface in a specific way, it can make the object look more realistic. It includes both the texture of the object's surface in the general sense, which makes the object's surface appear uneven with grooves, and the colored patterns on the smooth surface of the object.

[0043] Most traditional 3D face reconstruction methods are based on image information, such as 3D face reconstruction using one or more information modeling techniques based on image brightness, edge information, linear perspective, color, relative height, disparity, etc. The model-based 3D face reconstruction method is a currently popular 3D face reconstruction method; currently popular models include the generic face model (CANDIDE-3) and the 3D morphable model (3D MM) and their variant models. The 3D face reconstruction algorithms based on them include both traditional algorithms and deep learning algorithms.

[0044] In the prior art, whether it is 3D face reconstruction based on traditional algorithms or 3D face reconstruction based on deep learning algorithms, due to the lack of prior image feature information, the reconstruction results often have a strong model sense, low accuracy, and lack of details. When facing faces with large expressions or poses, they often perform unsatisfactorily.

[0045] Based on this, the present application proposes a 3D face model reconstruction method based on deep learning algorithms. By pre-training a deep learning-based face reconstruction model, during the training process, iterative training is performed on smoothness, the accuracy of the 3D face model, and the stability of the target texture map. Then, based on the trained face reconstruction model, 3D face reconstruction is performed on the 2D image, and a texture 3D face model with stable effects, rich details, and high realism and containing texture information can be reconstructed based on the 2D image.

[0046] The following will be described in detail with specific embodiments. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily have to be executed in the order shown. For example, some steps are parallel and there is no strict sequence in logic, so the actual execution order can be variable.

[0047] Please refer to Figure 1, which is a schematic flowchart of a three-dimensional face model reconstruction method provided by an embodiment of the present application. In a specific embodiment, the three-dimensional face reconstruction method is applied to a three-dimensional face model reconstruction device and a computer device configured with a three-dimensional face model reconstruction device. Hereinafter, the computer device will be taken as an example to illustrate the specific process of this embodiment. Of course, it can be understood that the computer device applied in this embodiment can be a smart phone, a tablet computer, a desktop computer, a wearable device, etc., which is not limited herein. The following will be elaborated in detail for the Figure 1 process shown below. The three-dimensional face reconstruction method may specifically include the following steps:

[0048] S101, obtain a target two-dimensional image;

[0049] The target two-dimensional image refers to a two-dimensional face image for which three-dimensional face reconstruction is required. The target two-dimensional image can be obtained by means such as offline shooting, cloud downloading, or calling from a local gallery.

[0050] S102, input the target two-dimensional image into the trained face reconstruction model, and output the target three-dimensional face model and the target texture map corresponding to the target two-dimensional image;

[0051] Specifically, input the target two-dimensional image into the trained face reconstruction model, and the face reconstruction reconstructs the target three-dimensional face model and the target texture map corresponding to the target two-dimensional image based on the face image information in the target two-dimensional image.

[0052] Optionally, before inputting the target two-dimensional image into the trained face reconstruction model, preprocess the target two-dimensional image. The preprocessing may include: if the target two-dimensional image is not with the face facing upward, rotate the target two-dimensional image so that the face in the target two-dimensional image faces upward; crop the target two-dimensional image to crop the background area outside the face area to reduce the amount of data that the face reconstruction model needs to operate on; perform normalization processing on the cropped target two-dimensional image to normalize the gray value of each pixel point in the target two-dimensional image to facilitate the face reconstruction model to perform calculations.

[0053] Optionally, the face reconstruction model may include a backbone based on a Convolutional Neural Network (CNN) and a Multi-layer Perceptron (MLP) model. The backbone is a CNN-based encoder network for extracting image features from the target two-dimensional image, and the backbone may adopt, but is not limited to, backbone networks such as mobilenet, resnet, xception, etc.

[0054] Optionally, inputting the target two-dimensional image into the trained face reconstruction model to output the target three-dimensional face model and the target texture map corresponding to the target two-dimensional image includes: on the one hand, inputting the target two-dimensional image into the CNN-based backbone network, and the backbone extracts image features from the target two-dimensional image, then inputs the extracted image features into the MLP model. The MLP model outputs regression parameters, and the regression parameters are transformed to generate a deformation parameter matrix, and further, the target three-dimensional face model corresponding to the target two-dimensional image is obtained based on the deformation parameter matrix; on the other hand, inputting the target two-dimensional image into the CNN-based backbone network, after the backbone extracts image features from the target two-dimensional image, the backbone outputs illumination parameters and texture parameters based on the image features. An illumination map containing illumination information is obtained by performing illumination modeling based on the illumination parameters, a first texture map is generated based on the texture parameters and a pre-designed texture template, and the illumination map and the first texture map are subjected to a dot product combination operation to generate the target texture map.

[0055] Please refer to Figure 2 , which is an example schematic diagram of 3D face reconstruction provided by an embodiment of the present application.

[0056] As Figure 2 shown, inputting the shown target two-dimensional image into the shown face reconstruction model can obtain the target three-dimensional face model as shown in the figure. The shown face reconstruction model includes a backbone based on a Convolutional Neural Network (CNN) and a Multi-layer Perceptron (MLP) model.

[0057] It can be understood that the face reconstruction model is obtained based on complex training. In the embodiments of the present application, the face reconstruction model is trained and generated based on a cost function, and the cost function includes a first cost function, a second cost function, a third cost function, and a fourth cost function. The first cost function is obtained based on the predicted three-dimensional face model corresponding to the sample two-dimensional image and the standard three-dimensional face model corresponding to the sample two-dimensional image. The second cost function is obtained based on the predicted three-dimensional face model and the smoothed predicted three-dimensional face model after smoothing the predicted three-dimensional face model. The third cost function is obtained based on the face key points after two-dimensional projection of the predicted three-dimensional face model and the face key points in the sample two-dimensional image. The fourth cost function is obtained based on the predicted texture map corresponding to the sample two-dimensional image and the standard texture map corresponding to the sample two-dimensional image. The predicted three-dimensional face model is predicted from the sample two-dimensional image based on the created initial face reconstruction model. The standard three-dimensional face model is obtained by constraint reconstruction of the average face three-dimensional model based on the face key points in the sample two-dimensional image.

[0058] S103. Based on the target three-dimensional face model and the target texture map, generate a texture three-dimensional face model including texture information.

[0059] Specifically, render the target texture map into the target three-dimensional face model according to the mapping relationship between the defined texture UV coordinates and the three-dimensional vertex coordinates in the target three-dimensional face model to generate a texture three-dimensional face model including texture information.

[0060] Optionally, before generating a texture three-dimensional face model including texture information based on the target three-dimensional face model and the target texture map, perform smoothing processing on the target three-dimensional face model to further improve the smoothing effect of the three-dimensional face model.

[0061] Optionally, before generating a texture three-dimensional face model including texture information based on the target three-dimensional face model and the target texture map, perform projection transformation on the target three-dimensional face model based on the projection parameters output by feature extraction of the target two-dimensional image by the backbone network, so that the pose of the target three-dimensional face model is closer to the pose in the target two-dimensional image, and improve the effect of the finally generated texture three-dimensional face model.

[0062] By using the 3D face model reconstruction method provided in the embodiments of the present application, a face reconstruction model pre-trained based on multiple cost functions can reconstruct a stable target texture map and a target 3D face model based on a target 2D image. Then, by rendering the target texture map onto the target 3D face model, a texture 3D face model with rich details and stable effects and containing texture information can be generated, improving the reconstruction effect of 3D faces.

[0063] Please refer to Figure 3 , a flowchart of a 3D face model reconstruction method provided in another embodiment of the present application. As Figure 3 shown, the 3D face model reconstruction method may include the following steps:

[0064] S201, collect sample 2D images;

[0065] The sample 2D images contain face images, and the sample 2D images are used as training data to assist in training the 3D face model.

[0066] Optionally, the face image area in the sample 2D image should exceed a certain proportion of the total area of the sample 2D image, and the proportion can be 75%.

[0067] Optionally, if the face image area in the sample 2D image does not exceed the preset proportion of the total area of the sample 2D image, the sample 2D image is cropped to obtain a new sample 2D image so that the face image area in the new sample 2D image exceeds the preset proportion of the total area of the sample 2D image.

[0068] Optionally, if the face image in the sample 2D image is skewed, the sample 2D image is rotated to obtain a new sample 2D image so that the face image in the new sample 2D image is a frontal face image.

[0069] The sample 2D images can be collected by offline shooting.

[0070] S202, obtain a standard 3D face model and a standard texture map corresponding to the sample 2D image;

[0071] The standard 3D face model is a 3D face model with high precision and close to the real face effect, and the standard texture map is a texture map with high precision and close to the real face texture effect.

[0072] In one embodiment, the obtaining of the standard three-dimensional face model and the standard texture map corresponding to the sample two-dimensional image may be to collect the real face of the person in the sample two-dimensional image based on a depth camera offline to obtain the three-dimensional point cloud data of the face, and then generate the standard three-dimensional face model and the standard texture map according to the three-dimensional point cloud data.

[0073] In one embodiment, the obtaining of the standard three-dimensional face model and the standard texture map corresponding to the sample two-dimensional image may also be to extract the face key points in the sample two-dimensional image, obtain the average face three-dimensional model, and then perform constrained reconstruction on the average face three-dimensional model based on the face key points in the sample two-dimensional image to obtain the standard three-dimensional face model corresponding to the sample two-dimensional image, obtain the UV coordinates of each vertex in the standard three-dimensional face model, and bilinearly interpolate the sample two-dimensional image based on the UV coordinates of each vertex to obtain the standard texture map of a preset size.

[0074] S203. Create an initial face reconstruction model, and train the initial face reconstruction model based on the sample two-dimensional image, the standard three-dimensional face model, and the standard texture map to obtain a trained face reconstruction model.

[0075] The initial face reconstruction model can reconstruct and generate a texture map and a three-dimensional face model corresponding to the two-dimensional image input into the model based on the two-dimensional image.

[0076] In one embodiment, train the initial face reconstruction model based on the sample two-dimensional image, the standard three-dimensional face model corresponding to the sample two-dimensional image, and the standard texture map corresponding to the sample two-dimensional image to improve the reconstruction effect of the initial face reconstruction model on the texture map and the three-dimensional face model, and obtain a trained face reconstruction model. The trained face reconstruction model has a good reconstruction effect on the texture map and the three-dimensional face model.

[0077] S204. Obtain a target two-dimensional image.

[0078] S205. Input the target two-dimensional image into the trained face reconstruction model, and output the target three-dimensional face model and the target texture map corresponding to the target two-dimensional image.

[0079] S206. Generate a texture three-dimensional face model containing texture information based on the target three-dimensional face model and the target texture map.

[0080] For the details of steps S210 to S212, please refer to the detailed description in steps S101 to S103, and will not be elaborated here.

[0081] Using the 3D face model reconstruction method provided by the embodiments of the present application, a pre-trained face reconstruction model can be used to reconstruct a stable target texture map and a target 3D face model based on a target 2D image. Then, the target texture map is rendered into the target 3D face model, and a texture 3D face model with rich details and stable effects including texture information can be generated, improving the reconstruction effect of 3D faces.

[0082] Please refer to Figure 4 , a schematic flowchart of a 3D face model reconstruction method provided by another embodiment of the present application. In this embodiment, the training process of the face reconstruction model is introduced in detail. As Figure 4 shown, the 3D face model reconstruction method may include the following steps:

[0083] S301, collect sample 2D images;

[0084] S302, extract the face key points in the sample 2D images, and obtain an average face 3D model;

[0085] In one embodiment, face key point detection is performed on the collected sample 2D images to obtain the face key points in the sample 2D images. The face key points refer to some important feature points such as eyes, the tip of the nose, the corners of the mouth, eyebrows, and the contour points of each component of the face, etc.

[0086] The face key point detection refers to locating the key area positions of the face and extracting the key points at the key area positions. The detection methods of the face key points include, but are not limited to, the Active Shape Model (ASM) algorithm, the Active Appearance Model (AAM) algorithm, and the deep learning algorithm, etc. In this regard, the embodiments of the present application do not make any limitations.

[0087] In one embodiment, an average face 3D model is obtained.

[0088] The average face 3D model can be obtained by using a depth camera to collect the face of a real person to obtain the 3D point cloud data of the face, and then generating a 3D face model according to the 3D point cloud data. It is necessary to collect the faces of multiple real people and generate multiple 3D face models, and finally average the multiple 3D face models to obtain the average face 3D model.

[0089] Optionally, the average face 3D model can also be a face model obtained based on big data analysis or pre-designed by a designer. The embodiments of the present application do not make any limitations.

[0090] Optionally, in one embodiment, when extracting the facial key points in the sample two-dimensional image, first perform stylization processing on the sample two-dimensional image to obtain a stylized sample two-dimensional image, and then extract the facial key points in the stylized sample two-dimensional image.

[0091] Optionally, the stylization processing of the sample two-dimensional image can be based on the unsupervised CycleGAN algorithm.

[0092] S303. Perform constrained reconstruction on the average face three-dimensional model based on the facial key points in the sample two-dimensional image to obtain a standard three-dimensional face model corresponding to the sample two-dimensional image.

[0093] In one embodiment, obtain an initial deformation matrix, reconstruct the average face three-dimensional model based on the deformation matrix to obtain an initial three-dimensional face model, project the initial three-dimensional face model into the two-dimensional space to obtain a two-dimensional face image, extract the facial key points in the two-dimensional face image, compare and calculate the difference between each facial key point in the two-dimensional face image and each facial key point in the sample two-dimensional image, update the deformation matrix based on the difference, and iteratively execute the above process until the difference between each facial key point in the two-dimensional face image and each facial key point in the sample two-dimensional image is less than a preset threshold, then stop the iteration. At this time, the deformation matrix is the required deformation matrix. Perform constrained reconstruction on the average face three-dimensional model based on the deformation matrix to obtain a standard three-dimensional face model corresponding to the sample two-dimensional image.

[0094] Please refer to Figure 5 , which is a flowchart for reconstructing a standard three-dimensional face model provided by an embodiment of the present application.

[0095] As Figure 5 shown, extract the facial key points from the sample two-dimensional image, and use the facial key points extracted from the sample two-dimensional image to perform constrained reconstruction on the average face three-dimensional model to obtain a standard three-dimensional face model.

[0096] Optionally, in one embodiment, the method of performing constrained reconstruction on the average face three-dimensional model based on the facial key points in the sample two-dimensional image to obtain a standard three-dimensional face model corresponding to the sample two-dimensional image may include: performing constrained reconstruction on the average face three-dimensional model based on the facial key points in the sample two-dimensional image to obtain a constrained three-dimensional face model and first projection parameters; performing projection transformation on the constrained three-dimensional face model based on the first projection parameters to obtain a standard three-dimensional face model. It should be understood that the process of performing constrained reconstruction on the average face three-dimensional model based on facial key points is an iterative optimization process. In this iterative optimization process, the reconstructed standard three-dimensional face model can be updated by continuously updating the deformation matrix. In this iterative optimization process, projection parameter information characterizing the facial pose can also be obtained. The first projection parameters include rotation parameters, translation parameters, and scale parameters. After the iterative optimization ends, projection transformation is performed on the constrained three-dimensional face model obtained by performing constrained reconstruction based on facial key points based on the first projection parameters to obtain a standard three-dimensional face model with accurate pose identical to the facial pose in the sample two-dimensional image. For example, if the face in the sample two-dimensional image is in a head-tilted state, then based on the first projection parameters after iterative optimization, the standard three-dimensional face model can be transformed into a head-tilted state as well.

[0097] Optionally, the iterative optimization process can be performed by finding the optimal value of the energy function to continuously update the deformation matrix and the projection parameters. Please refer to the following steps:

[0098] First, create an initial deformation matrix and projection parameters. Then perform iterations according to the following process.

[0099] Formula 1: min R,t,P′,w,Π E def (P′, w)+λE lan (Π, R, t, P′)

[0100] where w is the deformation matrix, Π is the scale parameter in the projection parameters, R is the rotation parameter in the projection parameters, t is the translation parameter in the projection parameters, and P′ is the coordinate of the point in the standard three-dimensional face model. Using this formula, the deformation matrix and the projection parameters can be fixed, and the coordinate P′ of the point in the standard three-dimensional face model can be obtained.

[0101] Determine whether the difference between E in this iteration and the previous iteration is less than the set threshold, that is, |E def -E def j -E def j-1 |, where j represents the iteration number. If it is satisfied, stop the iteration. If not, continue to execute the following process and perform the iteration.

[0102] Formula 2: E def= ||P′ - w.P|| 2

[0103] Wherein, P is the coordinate of a point in the average face three-dimensional model, P′ is obtained by using formula 1. Fix P′ in the formula, and use the least squares method to solve the deformation matrix w.

[0104] Formula 3:

[0105] Wherein, q i is the coordinate of the i-th key point in the sample two-dimensional image, p′ i is the coordinate of the point corresponding to q in the sample two-dimensional image i in the standard three-dimensional face model. Substitute P′ obtained by using formula 1 into formula 3 to update the projection parameters Π, R, t.

[0106] It can be understood that formula 2 is essentially the update of the deformation matrix in each iteration process, and formula 3 is the update of the projection parameters in each iteration process.

[0107] Optionally, in an embodiment, after performing constrained reconstruction on the average face three-dimensional model based on the face key points in the sample two-dimensional image to obtain a standard three-dimensional face model, perform smoothing processing on the standard three-dimensional face model to obtain a new standard three-dimensional face model, so that the newly obtained standard three-dimensional face model has a better smoothing effect. In this way, in step S308, the first cost function constructed based on the vertices in the predicted three-dimensional face model and the vertices in the standard three-dimensional face model not only plays a role in vertex correction of the face reconstruction model during the training process, but also can enhance the smoothing effect of the face reconstruction model.

[0108] Specifically, the performing smoothing processing on the standard three-dimensional face model to obtain a new standard three-dimensional face model includes:

[0109]

[0110] Wherein, V i ′ is the i-th three-dimensional vertex in the standard three-dimensional face model before smoothing processing, and the is the first-order neighborhood point directly adjacent to V i ′. i represents the i-th three-dimensional vertex, j represents the label of the first-order neighborhood point, and the is the i-th three-dimensional vertex in the standard three-dimensional face model after smoothing processing, and α represents the smoothing coefficient, α ∈ (0, 1).

[0111] Optionally, the smoothing of the standard 3D face model to obtain the standard 3D face model can also be based on curvature smoothing, Taubin smoothing algorithm, etc. For the specific implementation of the smoothing process, the embodiments of the present application do not make limitations.

[0112] S304. Obtain the UV coordinates of each vertex in the standard 3D face model, and bilinearly interpolate the sample two-dimensional image based on the UV coordinates of each vertex to obtain a standard texture map of a preset size;

[0113] In one embodiment, the obtained standard 3D face model is used to obtain the UV coordinates corresponding to each vertex in the standard 3D face model according to the predefined mapping relationship between the three-dimensional vertex coordinates and the UV coordinates. Then, according to the UV coordinates corresponding to each vertex in the standard 3D face model and the pixel information corresponding to each pixel point in the sample two-dimensional image, a bilinear interpolation algorithm is used to interpolate to obtain a standard texture map of a preset size. Among them, the pixel information includes but is not limited to color information, transparency information, roughness information, metallicity information, specular information, etc. of each pixel point.

[0114] It should be understood that the standard texture map is used to be rendered on the surface of the standard 3D face model. Therefore, in one embodiment, the preset size is set based on the size of the standard 3D face model.

[0115] Please refer to Figure 6 , which provides a flowchart for generating a standard texture map according to an embodiment of the present application.

[0116] As Figure 6 shown, the standard 3D face model is subjected to coordinate transformation according to the predefined coordinate mapping relationship to obtain the UV coordinates of each vertex in the standard 3D face model, and bilinear interpolation is performed based on the UV coordinates of each vertex and the pixel information of each pixel point in the sample two-dimensional image to obtain a standard texture map of a preset size. Exemplarily, please also refer to Figure 7 the example schematic diagram of generating a standard texture map shown.

[0117] S305. Create an initial face reconstruction model, and input the sample two-dimensional image into the initial face reconstruction model for prediction to obtain a predicted 3D face model and a predicted texture map;

[0118] The initial face reconstruction model refers to a rough deep learning model that has not been trained.

[0119] In one embodiment, an initial face reconstruction model is created, and a sample two-dimensional image is input into the initial face model. The initial face reconstruction model predicts a predicted three-dimensional face model and a predicted texture map based on the initial parameter information in the initial face reconstruction model and the image feature information in the sample two-dimensional image.

[0120] It can be understood that since the initial face reconstruction model is an untrained and rough deep learning model, the predicted three-dimensional face model and the predicted texture map obtained do not have good model effects and texture effects. Therefore, it is necessary to train the initial face reconstruction model to continuously update the parameter information in the initial face reconstruction model, so as to continuously improve the model effect of the predicted three-dimensional face model and the texture effect of the predicted texture map.

[0121] Further, in one embodiment, creating the initial face reconstruction model and inputting the sample two-dimensional image into the initial face reconstruction model for prediction to obtain a predicted three-dimensional face model and a predicted texture map may include: creating the initial face reconstruction model, and inputting the sample two-dimensional image into the initial face reconstruction model for prediction to obtain a deformation parameter matrix, lighting parameters, and texture parameters; predicting a predicted three-dimensional face model based on the deformation parameter matrix; performing lighting modeling based on the lighting parameters to obtain a lighting map, and obtaining a first texture map based on the texture parameters and a preset texture template; generating a predicted texture map based on the lighting map and the first texture map.

[0122] Among them, in this embodiment, performing lighting modeling based on the lighting parameters to obtain a lighting map may be based on spherical harmonic lighting technology. Specifically, a normal map is extracted from the sample two-dimensional image based on a backbone network, and the lighting map is calculated based on the normal map and the lighting parameters. The following formula can be referred to:

[0123]

[0124] Among them, the is the lighting parameter, the Norm is the normal map, and the H i (Norm) is used to expand the normal map from the (h, w, c) dimension to the (9, h, w) dimension, where h represents the height of the normal map, w represents the width of the normal map, and c represents the number of channels. The SH shading is the lighting map.

[0125] Optionally, performing lighting modeling based on the lighting parameters to obtain a lighting map may also be performed based on other achievable lighting models. In this regard, the embodiments of the present application do not make limitations.

[0126] Optionally, the texture template is a basic texture template pre-designed based on factors such as different skin colors, different genders, and different ages. Further, based on the texture parameters corresponding to the sample two-dimensional image output by the backbone network and the basic texture template, a first texture map corresponding to the sample two-dimensional image can be generated. Specifically, the first texture map can be calculated based on the following method:

[0127]

[0128] wherein, the Tex mean is the average of M texture templates, M represents the number of texture templates, and the refers to the texture parameters output during the training process of the backbone network, a total of M, and the Tex basis refers to each texture template minus the average value of the texture templates.

[0129] Optionally, the first texture map contains albedo information, roughness information, metallicity information, specular information, height information, transparency information, etc.

[0130] It should be understood that generating a first texture map corresponding to the sample two-dimensional image based on the pre-designed texture template and the output texture parameters can ensure that the predicted texture map generated based on the first texture map and the light map is always within the variation range defined by the pre-designed texture template, without large errors, and the texture effect is relatively stable.

[0131] It can be understood that the more the number of pre-designed texture templates, the richer the details of the generated first texture map.

[0132] Please refer to Figure 8 for an example schematic diagram of a predicted texture map provided by an embodiment of the present application. As Figure 8 shown, the sample two-dimensional image is input into the backbone network in the face reconstruction model. The backbone network extracts features from the sample two-dimensional image and outputs light parameters and texture parameters. Based on the light parameters, light modeling can obtain a light map. Based on the texture parameters and the pre-designed texture template, a first texture map can be obtained. Based on the first texture map and the light map, a combined operation can obtain the final predicted texture map.

[0133] S306. Smooth the predicted three-dimensional face model to obtain a smoothed predicted three-dimensional face model;

[0134] Specifically, refer to the specific implementation manner of smoothing the standard three-dimensional face model to obtain a new standard three-dimensional face model in step S303.

[0135] S307. Project the predicted three-dimensional face model into a two-dimensional image to obtain a predicted two-dimensional image, and extract the facial key points in the predicted two-dimensional image.

[0136] In one embodiment, after inputting the sample two-dimensional image into the CNN-based backbone network in the face reconstruction model, the backbone network outputs a second projection parameter. Based on the second projection parameter, project the predicted three-dimensional face model into a two-dimensional image to obtain a predicted two-dimensional image, and extract the facial key points in the predicted two-dimensional image.

[0137] It can be understood that the embodiment of the present application is essentially an iterative training process. In each iteration process, the parameters of the initial face reconstruction model are adjusted and updated, and then the updated initial face reconstruction model is used for the next round of iteration process. After adjusting and updating the parameters of the initial face reconstruction model, the updated initial face reconstruction model will also update the output second projection parameter.

[0138] S308. Construct a cost function based on the predicted three-dimensional face model, the standard three-dimensional face model, the smoothed predicted three-dimensional face model, the facial key points in the predicted two-dimensional image, the facial key points in the sample two-dimensional image, the predicted texture map, and the standard texture map.

[0139] Schematically, on the basis of Figure 4 as shown in Figure 9 Step S307 may include S3081, S3082, S3083, S3084, and S3085.

[0140] S3081. Construct a first cost function based on the vertices in the predicted three-dimensional face model and the vertices in the standard three-dimensional face model.

[0141]

[0142] where N is the number of vertices, V i is the i-th vertex in the predicted three-dimensional face model, and V GT,i is the i-th vertex in the standard three-dimensional face model.

[0143] It is not difficult to understand that the predicted three-dimensional face model is directly predicted by the initial face reconstruction model from the sample two-dimensional image, and the standard three-dimensional face model can be obtained by constraining and reconstructing the average face three-dimensional model based on the face key points in the sample two-dimensional image and then performing smoothing processing. Therefore, compared with the predicted three-dimensional face model, the standard three-dimensional face model has higher accuracy in model shape and better smoothing effect. Furthermore, based on the vertices in the predicted three-dimensional face model and the vertices in the standard three-dimensional face model, a first cost function is constructed, and by adjusting and training the internal parameters of the initial face reconstruction model based on the first cost function, the accuracy and smoothing effect of the initial face reconstruction model for the reconstruction model can be improved.

[0144] S3082, construct a second cost function based on the vertices in the predicted three-dimensional face model and the vertices in the smoothed predicted three-dimensional face model;

[0145]

[0146] where N is the number of vertices, V i is the i-th vertex in the predicted three-dimensional face model, and is the i-th vertex in the smoothed predicted three-dimensional face model.

[0147] It is not difficult to understand that the predicted three-dimensional face model is directly predicted by the initial face reconstruction model from the sample two-dimensional image, and the smoothed predicted three-dimensional face model is obtained by smoothing the predicted three-dimensional face model. Therefore, the smoothed predicted three-dimensional face model has a significant smoothing effect compared with the predicted three-dimensional face model. Furthermore, based on the vertices in the predicted three-dimensional face model and the vertices in the smoothed predicted three-dimensional face model, a second cost function is constructed, and by adjusting and training the internal parameters of the initial face reconstruction model based on the second cost function, the smoothing effect of the initial face reconstruction model for the reconstruction model can be optimized.

[0148] S3083, construct a third cost function based on the face key points in the predicted two-dimensional image and the face key points in the sample two-dimensional image;

[0149]

[0150] where M is the number of face key points, lmk i is the i-th face key point in the predicted two-dimensional image, and lmk GT,i is the i-th face key point in the sample two-dimensional image.

[0151] It is not difficult to understand that the predicted two-dimensional image is obtained by performing a two-dimensional projection on the predicted three-dimensional face model. Therefore, compared with the sample two-dimensional image, there are significant differences in the facial key points in the predicted two-dimensional image. A third cost function is constructed based on the facial key points in the predicted two-dimensional image and the facial key points in the sample two-dimensional image. By adjusting and training the internal parameters of the initial face reconstruction model based on the third cost function, the accuracy of the initial face reconstruction model for the reconstruction model can be optimized, making the reconstructed target three-dimensional face model closer to the standard three-dimensional face model corresponding to the sample two-dimensional image.

[0152] S3084, construct a fourth cost function based on each pixel in the predicted texture map and each pixel in the standard texture map;

[0153]

[0154] wherein, the Pixel is the number of pixels, the Tex i is the i-th pixel in the predicted texture map, and the Tex GT,i is the i-th pixel in the standard texture map.

[0155] It can be understood that the standard texture map is obtained by performing bilinear interpolation based on the standard three-dimensional face model and the sample two-dimensional image. Compared with the predicted texture map directly predicted by the initial face reconstruction model, it has a more prominent texture effect, with richer and more accurate details. By constructing a fourth cost function based on the predicted texture map and the standard texture map, and adjusting and training the internal parameters of the initial face reconstruction model based on the fourth cost function, the prediction effect of the initial face reconstruction model on the texture map can be optimized, making the predicted texture map closer to the standard texture map with rich details, accuracy, and a prominent texture effect, and improving the prediction effect of the initial face reconstruction model on the texture map.

[0156] S3085, perform a weighted sum of the first cost function, the second cost function, the third cost function, and the fourth cost function to obtain a cost function.

[0157] In one embodiment, different weight values are set for the first cost function, the second cost function, the third cost function, and the fourth cost function respectively, and a weighted sum is performed to obtain a total cost function including these four cost functions, so as to adjust and train the internal parameters of the initial face reconstruction model based on the total cost function.

[0158] Further, in a feasible implementation manner, a cost function based on MSE is calculated and backpropagation is performed. Optionally, other forms of cost functions different from MSE can also be used, and the embodiments of the present application do not limit this. For example, if a cost function based on SAD is used, the first cost function, the second cost function, and the third cost function can be respectively expressed as:

[0159] The first cost function:

[0160] The second cost function:

[0161] The third cost function:

[0162] S309. Train the initial face reconstruction model based on the cost function to obtain a trained face reconstruction model;

[0163] It should be noted that in the embodiments of the present application, steps S305 to S309 are an iterative loop training process, and steps S301 to S304 are used to provide training data for the iterative loop training process of steps S305 to S309.

[0164] In one embodiment, the model parameters of the face reconstruction model are updated by backpropagation based on the cost function, and iterative loop training is performed until the function value of the cost function is less than a preset threshold or a preset number of loops is reached.

[0165] Optionally, in a realizable manner, the internal parameters of the initial face reconstruction model can be adjusted and trained first based on the first cost function, the second cost function, and the third cost function. After the cost function is less than a certain threshold or the number of iterations reaches a preset number, the first-stage training ends. At this time, the face reconstruction model has converged to a good effect on the predicted three-dimensional face model. Then, the second-stage training is performed in combination with the first cost function, the second cost function, the third cost function, and the fourth cost function to further optimize the prediction effect of the face reconstruction model on the three-dimensional face model and the texture map.

[0166] Please refer to Figure 10 , which provides a flowchart for training a face reconstruction model in the embodiments of the present application.

[0167] As Figure 10 shown, the sample two-dimensional image is input into the initial face reconstruction model. The initial face reconstruction model predicts a predicted three-dimensional face model and a predicted texture map. Based on the predicted three-dimensional face model and as Figure 5Construct the first cost function for the standard three-dimensional face model obtained from the flowchart shown; smooth the predicted three-dimensional face model to obtain a smoothed predicted three-dimensional face model, and construct the second cost function based on the smoothed predicted three-dimensional face model and the predicted three-dimensional face model; project the predicted three-dimensional face model into two dimensions by combining the projection parameters given by the initial face reconstruction model to obtain a predicted two-dimensional image, extract the facial key points in the predicted two-dimensional image, such as the facial key point 2 shown in the figure, and the facial key point 1 shown in the figure is the facial key point in the sample two-dimensional image, and construct the third cost function based on the facial key point 2 in the predicted two-dimensional image and the facial key point 1 in the sample two-dimensional image; construct the fourth cost function based on the predicted texture map and the standard texture map predicted by the initial face reconstruction model; perform weighted summation on the first cost function, the second cost function, the third cost function, and the fourth cost function and update the initial face reconstruction model. The above process is a complete training process, and this process is iterative training. The end flag of the training can be that the number of iterations reaches a preset number or the weighted summation value of the first cost function, the second cost function, the third cost function, and the fourth cost function is less than a preset threshold. After the training ends, the face reconstruction model can be obtained.

[0168] S310. Obtain the target two-dimensional image;

[0169] S311. Input the target two-dimensional image into the trained face reconstruction model, and output the target three-dimensional face model and the target texture map corresponding to the target two-dimensional image;

[0170] S312. Generate a texture three-dimensional face model containing texture information based on the target three-dimensional face model and the target texture map.

[0171] For the details of steps S310 to S312, please refer to the detailed description in steps S101 to S103 and will not be elaborated here.

[0172] In the embodiment of the present application, first, an initial face reconstruction model is created, and sample two-dimensional images are collected. Then, face key points in the sample two-dimensional images are extracted, and an average face three-dimensional model is obtained. The average face three-dimensional model is constrained and reconstructed based on the face key points in the sample two-dimensional images to obtain a standard three-dimensional face model corresponding to the sample two-dimensional images. The UV coordinates of each vertex in the standard three-dimensional face model are obtained, and the sample two-dimensional images are bilinearly interpolated based on the UV coordinates of each vertex to obtain a standard texture map of a preset size. Then, the sample two-dimensional images are input into the initial face reconstruction model for prediction to obtain a predicted three-dimensional face model. The predicted three-dimensional face model is smoothed to obtain a smoothed predicted three-dimensional face model. Then, the predicted three-dimensional face model is two-dimensionally projected to obtain a predicted two-dimensional image. The face key points in the predicted two-dimensional image are further extracted. Finally, four cost functions are constructed based on the predicted three-dimensional face model, the standard three-dimensional face model, the smoothed predicted three-dimensional face model, the face key points in the predicted two-dimensional image, the face key points in the sample two-dimensional images, the predicted texture map, and the standard texture map. The initial face reconstruction model is iteratively trained using the four constructed cost functions. Finally, a trained face reconstruction model is obtained. During the iterative training process, training on the smoothing effect, the accuracy of the model shape, the stability effect of the texture map, and the detail effect is included, so that the trained face reconstruction model has good accuracy, prominent smoothing effect, and good stability for three-dimensional face model reconstruction and target texture map prediction. Then, the trained face reconstruction model can be used to reconstruct a three-dimensional face model and predict a target texture map with high accuracy, significant smoothing effect, and high detail restoration degree for a target two-dimensional image. Among them, the texture map is generated based on parameters and a pre-designed texture template, so that the texture effect of the texture map is always within the effect range defined by the texture template, and there will be no large texture effect error, improving the texture effect of the texture map and ensuring the stability of the texture map. Moreover, the face reconstruction model only needs to be trained once, and then it can be reused for any target two-dimensional image. During use, only by inputting the target two-dimensional image into the face reconstruction model, a three-dimensional face model and a target texture map can be obtained, and the target two-dimensional image, the three-dimensional face model, and the target texture map are generated end-to-end without increasing additional computational complexity.

[0173] By using the embodiment of the present application to train and complete the face reconstruction model based on multiple cost functions, a target texture map and a target three-dimensional face model with stable effects can be reconstructed based on a target two-dimensional image. Then, the target texture map is rendered onto the target three-dimensional face model, and a texture three-dimensional face model with rich details and stable effects and containing texture information can be generated, improving the reconstruction effect of the three-dimensional face model.

[0174] Optionally, stylization is added during the training process of the face reconstruction model, so that the face reconstruction model can reconstruct a stylized three-dimensional face model and the target texture map.

[0175] Please refer to Figure 11 , which provides a schematic flowchart of a three-dimensional face model reconstruction method according to an embodiment of the present application. As Figure 11 shown, the three-dimensional face model reconstruction method includes the following steps.

[0176] S401, collect sample two-dimensional images;

[0177] S402, perform stylization processing on the sample two-dimensional image to obtain a stylized sample two-dimensional image, and extract the face key points in the stylized sample two-dimensional image;

[0178] The stylization processing refers to converting the sample two-dimensional image into a specific stylized sample two-dimensional image, such as a sketch portrait style, a cartoon image (animation) style, an oil painting style, etc.

[0179] Please refer to Figure 12 , which provides an example schematic diagram of stylization processing according to an embodiment of the present application.

[0180] As Figure 12 shown, the sample two-dimensional image shown can be processed into a stylized sample two-dimensional image of a cartoon image as shown after being processed in the cartoon image style.

[0181] S403, obtain an average face three-dimensional model;

[0182] S404, perform constrained reconstruction on the average face three-dimensional model based on the face key points in the stylized sample two-dimensional image to obtain a standard three-dimensional face model corresponding to the sample two-dimensional image;

[0183] Please refer to Figure 13 , which provides an example schematic diagram of using face key points to constrain the reconstruction of a standard three-dimensional face model according to an embodiment of the present application.

[0184] As Figure 13 shown, the average face three-dimensional model shown is constrained by the face key points in the stylized sample two-dimensional image during the reconstruction of the standard three-dimensional face model to obtain a standard three-dimensional face model corresponding to the stylized sample two-dimensional image.

[0185] S405, obtain the UV coordinates of each vertex in the standard three-dimensional face model, and bilinearly interpolate the stylized sample two-dimensional image based on the UV coordinates of each vertex to obtain a standard texture map of a preset size;

[0186] S406. Create an initial face reconstruction model, and input the sample two-dimensional image into the initial face reconstruction model for prediction to obtain a predicted three-dimensional face model and a predicted texture map;

[0187] S407. Smooth the predicted three-dimensional face model to obtain a smoothed predicted three-dimensional face model;

[0188] S408. Project the predicted three-dimensional face model onto a two-dimensional plane to obtain a predicted two-dimensional image, and extract the facial key points in the predicted two-dimensional image;

[0189] S409. Construct a cost function based on the predicted three-dimensional face model, the standard three-dimensional face model, the smoothed predicted three-dimensional face model, the facial key points in the predicted two-dimensional image, the facial key points in the stylized sample two-dimensional image, the predicted texture map, and the standard texture map;

[0190] S410. Train the initial face reconstruction model based on the cost function to obtain a trained face reconstruction model;

[0191] S411. Obtain a target two-dimensional image;

[0192] S412. Input the target two-dimensional image into the trained face reconstruction model, and output a stylized three-dimensional face model and a stylized texture map corresponding to the target two-dimensional image;

[0193] Wherein, the face reconstruction model is generated by training based on a cost function, the cost function includes a first cost function, a second cost function, and a third cost function. The first cost function is obtained based on the predicted stylized three-dimensional face model corresponding to the sample two-dimensional image and the standard stylized three-dimensional face model corresponding to the sample two-dimensional image. The second cost function is obtained based on the predicted stylized three-dimensional face model and the smoothed stylized three-dimensional face model after smoothing the predicted stylized three-dimensional face model. The third cost function is obtained based on the facial key points after two-dimensional projection of the predicted stylized three-dimensional face model and the facial key points in the stylized sample two-dimensional image corresponding to the sample two-dimensional image;

[0194] The predicted stylized three-dimensional face model is obtained by predicting the sample two-dimensional image using the created initial face reconstruction model. The standard stylized three-dimensional face model is obtained by performing constrained reconstruction on the average face three-dimensional model based on the facial key points in the stylized sample two-dimensional image corresponding to the sample two-dimensional image.

[0195] S413. Generate a textured three-dimensional face model containing texture information based on the stylized three-dimensional face model and the stylized texture map.

[0196] By adopting the 3D face model reconstruction method provided in the embodiments of the present application, stylization processing is added during the process of training the face reconstruction model. The face reconstruction model trained based on the smoothness effect, accuracy, and texture effect stability can output a stylized 3D face model and a stylized texture map. Based on the stylized 3D face model and the stylized texture map, a texture 3D face model containing texture information can be generated, increasing the interest and functionality of 3D face model reconstruction.

[0197] Please refer to Figure 14 , which is a schematic structural diagram of a 3D face model reconstruction device provided in the embodiments of the present application. As Figure 14 shown, the 3D face model reconstruction device 1 can be implemented as all or part of a computer device through software, hardware, or a combination of both. According to some embodiments, the 3D face model reconstruction device 1 includes an image acquisition module 11 and a model prediction module 12, specifically including:

[0198] The image acquisition module 11 is used to acquire a target two-dimensional image;

[0199] The model prediction module 12 is used to input the target two-dimensional image into the trained face reconstruction model and output the target 3D face model corresponding to the target two-dimensional image;

[0200] The target model generation module 13 is used to generate a texture 3D face model containing texture information based on the target 3D face model and the target texture map.

[0201] Optionally, the model prediction module 12 is specifically used for:

[0202] Input the target two-dimensional image into the trained face reconstruction model and output the stylized 3D face model and the stylized texture map corresponding to the target two-dimensional image;

[0203] The target model generation module 13 is specifically used for:

[0204] Generate a texture 3D face model containing texture information based on the stylized 3D face model and the stylized texture map.

[0205] Optionally, as Figure 15 shown, the device further includes a model training module 14.

[0206] Optionally, please refer to Figure 16 , which is a schematic structural diagram of a model training module provided in the embodiments of the present application. As Figure 16 shown, the model training module 14 includes:

[0207] The image acquisition unit 141 is configured to acquire a two-dimensional sample image;

[0208] The standard acquisition unit 142 is configured to acquire a standard three-dimensional face model and a standard texture map corresponding to the two-dimensional sample image;

[0209] The model training unit 143 is configured to create an initial face reconstruction model, and train the initial face reconstruction model based on the two-dimensional sample image, the standard three-dimensional face model, and the standard texture map to obtain a trained face reconstruction model.

[0210] Optionally, please refer to Figure 17 , which provides a schematic structural diagram of a standard acquisition unit for an embodiment of the present application. As Figure 17 shown, the standard acquisition unit 142 includes:

[0211] The first key point extraction subunit 1421 is configured to extract the face key points in the two-dimensional sample image and obtain an average face three-dimensional model;

[0212] The model reconstruction subunit 1422 is configured to perform constrained reconstruction on the average face three-dimensional model based on the face key points in the two-dimensional sample image to obtain a standard three-dimensional face model corresponding to the two-dimensional sample image;

[0213] The texture generation subunit 1423 is configured to obtain the UV coordinates of each vertex in the standard three-dimensional face model, and bilinearly interpolate the two-dimensional sample image based on the UV coordinates of each vertex to obtain a standard texture map of a preset size.

[0214] Optionally, the model reconstruction subunit 1422 is specifically configured to:

[0215] Perform constrained reconstruction on the average face three-dimensional model based on the face key points in the two-dimensional sample image to obtain a constrained three-dimensional face model and a first projection parameter;

[0216] Perform projection transformation on the constrained three-dimensional face model based on the first projection parameter to obtain a standard three-dimensional face model.

[0217] Optionally, please refer to Figure 18 , which provides a schematic structural diagram of a model training unit for an embodiment of the present application. As Figure 18 shown, the model training unit 143 includes:

[0218] The model prediction subunit 1431 is configured to create an initial face reconstruction model, and input the two-dimensional sample image into the initial face reconstruction model for prediction to obtain a predicted three-dimensional face model and a predicted texture map;

[0219] The smoothing subunit 1432 is configured to smooth the predicted three-dimensional face model to obtain a smoothed predicted three-dimensional face model;

[0220] The second key point extraction subunit 1433 is configured to project the predicted three-dimensional face model into a predicted two-dimensional image and extract the face key points in the predicted two-dimensional image;

[0221] The cost function construction subunit 1434 is configured to construct a cost function based on the predicted three-dimensional face model, the standard three-dimensional face model, the smoothed predicted three-dimensional face model, the face key points in the predicted two-dimensional image, the face key points in the sample two-dimensional image, the predicted texture map, and the standard texture map;

[0222] The model training subunit 1435 is configured to train the initial face reconstruction model based on the cost function to obtain a trained face reconstruction model.

[0223] Optionally, the model prediction subunit 1431 is specifically configured to:

[0224] Create an initial face reconstruction model, and input the sample two-dimensional image into the initial face reconstruction model to predict a deformation parameter matrix, lighting parameters, and texture parameters;

[0225] Predict a predicted three-dimensional face model based on the deformation parameter matrix;

[0226] Perform lighting modeling based on the lighting parameters to obtain a lighting map;

[0227] Obtain a first texture map based on the texture parameters and a preset texture template; and

[0228] Generate a predicted texture map based on the lighting map and the first texture map.

[0229] Optionally, the second key point extraction subunit 1433 is specifically configured to:

[0230] Obtain the second projection parameters in the initial face reconstruction model, project the predicted three-dimensional face model into a predicted two-dimensional image based on the second projection parameters, and extract the face key points in the predicted two-dimensional image.

[0231] Optionally, the cost function construction subunit 1434 is specifically configured to:

[0232] Construct a first cost function based on the vertices in the predicted three-dimensional face model and the vertices in the standard three-dimensional face model: where N is the number of vertices, and V iis the i-th vertex in the predicted 3D face model, and the V GT,i is the i-th vertex in the standard 3D face model;

[0233] Construct a second cost function based on the vertices in the predicted 3D face model and the vertices in the smoothed predicted 3D face model: where N is the number of vertices, and the V i is the i-th vertex in the predicted 3D face model, and the is the i-th vertex in the smoothed predicted 3D face model;

[0234] Construct a third cost function based on the facial key points in the predicted 2D image and the facial key points in the sample 2D image: where M is the number of facial key points, and the lmk i is the i-th facial key point in the predicted 2D image, and the lmk GT,i is the i-th facial key point in the sample 2D image;

[0235] Construct a fourth cost function based on the pixels in the predicted texture map and the pixels in the standard texture map: where Pixel is the number of pixels, and the Tex i is the i-th pixel in the predicted texture map, and the Tex GT,i is the i-th pixel in the standard texture map;

[0236] Perform weighted summation on the first cost function, the second cost function, the third cost function, and the fourth cost function to obtain a cost function.

[0237] Optionally, the cost function construction subunit 1434 is specifically used for:

[0238] L = β·L 1 + μ· 2 + λ· 3 + ω· 4

[0239] where β is the weight corresponding to the first cost function L 1 μ is the weight corresponding to the second cost function L 2 λ is the weight corresponding to the third cost function L 3 ω is the weight corresponding to the fourth cost function L 4 corresponding weight.

[0240] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0241] Using the 3D face model reconstruction method provided in the embodiments of the present application, a face reconstruction model pre-trained based on multiple cost functions can reconstruct a target texture map with stable effects and a target 3D face model based on a target 2D image. The target texture map is generated based on parameters and a pre-designed texture template, so that the texture effect of the target texture map is always within the effect range defined by the texture template, and large texture effect errors will not occur, improving the texture effect of the target texture map and ensuring the stability of the target texture map. Then, rendering the target texture map into the target 3D face model can generate a textured 3D face model with rich details and stable effects, enhancing the reconstruction effect of the 3D face. Optionally, during the process of training the face reconstruction model, stylization processing can be added. The face reconstruction model trained based on smoothness and accuracy can output a stylized 3D face model and a stylized texture map. Based on the stylized 3D face model and the stylized texture map, a textured 3D face model containing texture information can be generated, increasing the interest and functionality of 3D face model reconstruction.

[0242] The embodiments of the present application also provide a computer storage medium, which can store multiple instructions. The instructions are suitable for being loaded and executed by a processor to perform the 3D face model reconstruction method as described in the above Figures 1 to 13 shown embodiments. The specific execution process can refer to the Figures 1 to 13 specific description of the shown embodiments and will not be elaborated here.

[0243] The present application also provides a computer program product. The computer program product stores at least one instruction, and the at least one instruction is loaded and executed by the processor to perform the 3D face model reconstruction method as described in the above Figures 1 to 13 shown embodiments. The specific execution process can refer to the Figures 1 to 13 specific description of the shown embodiments and will not be elaborated here.

[0244] Please refer to Figure 19 , which shows a structural block diagram of a computer device provided by an exemplary embodiment of the present application. The computer device in the present application may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected through the bus 150.

[0245] The processor 110 may include one or more processing cores. The processor 110 connects various parts within the entire computer device using various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by invoking the data stored in the memory 120, it performs various functions of the computer device 100 and processes data. Optionally, the processor 110 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 110 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the display content; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 110 and can be implemented separately through a communication chip.

[0246] The memory 120 may include random access memory (RAM) and may also include read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets.

[0247] Among them, the input device 130 is used to receive input instructions or data. The input device 130 includes, but is not limited to, a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 140 is used to output instructions or data. The output device 140 includes, but is not limited to, a display device and a speaker, etc. In the embodiments of the present application, the input device 130 may be a temperature sensor for obtaining the operating temperature of the computer device. The output device 140 may be a speaker for outputting an audio signal.

[0248] In addition, those skilled in the art can understand that the structure of the computer device shown in the above drawings does not limit the computer device. The computer device may include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements. For example, the computer device may also include components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (WiFi) module, a power supply, and a Bluetooth module, which will not be elaborated here.

[0249] In the embodiments of the present application, the execution subject of each step may be the computer device introduced above. Optionally, the execution subject of each step is the operating system of the computer device. The operating system may be the Android system, the IOS system, or other operating systems, which are not limited in the embodiments of the present application.

[0250] In Figure 19 In the computer device shown, the processor 110 may be used to call the three-dimensional face model reconstruction program stored in the memory 120 and execute it to implement the three-dimensional face model reconstruction method as described in various method embodiments of the present application.

[0251] By using the three-dimensional face model reconstruction method provided in the embodiments of the present application, the face reconstruction model pre-trained based on multiple cost functions can reconstruct a target texture map with stable effects and a target three-dimensional face model based on the target two-dimensional image. The target texture map is generated based on parameters and a pre-designed texture template, so that the texture effect of the target texture map is always within the effect range defined by the texture template, and there will be no large texture effect errors, improving the texture effect of the target texture map and ensuring the stability of the target texture map. Then, rendering the target texture map into the target three-dimensional face model can generate a texture three-dimensional face model with rich details and stable effects and containing texture information, improving the reconstruction effect of the three-dimensional face. Optionally, during the process of training the face reconstruction model, stylization processing can be added. The face reconstruction model trained based on the smoothing effect and accuracy can output a stylized three-dimensional face model and a stylized texture map. Based on the stylized three-dimensional face model and the stylized texture map, a texture three-dimensional face model containing texture information can be generated, increasing the interest and functionality of the three-dimensional face model reconstruction.

[0252] Those skilled in the art can clearly understand that the technical solutions of the present application can be implemented by means of software and / or hardware. The "units" and "modules" in this specification refer to software and / or hardware that can independently complete or cooperate with other components to complete specific functions, where the hardware may be, for example, a Field-Programmable Gate Array (FPGA), an Integrated Circuit (IC), etc.

[0253] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0254] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0255] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0256] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0257] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0258] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable memory. The memory can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.

[0259] The foregoing are only exemplary embodiments of the present application, and thus cannot limit the scope of the present application. That is, all equivalent changes and modifications made in accordance with the teachings of the present application still fall within the scope covered by the present application. Other embodiments of the present application will be readily contemplated by those skilled in the art after considering the specification and practicing the disclosure herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not recorded in the present application. The specification and embodiments are only regarded as exemplary, and the scope and spirit of the present application are defined by the claims.

Claims

1. A method for three-dimensional human face model reconstruction, characterized in that, the method includes: Collecting sample two-dimensional images; Obtaining a standard three-dimensional human face model and a standard texture map corresponding to the sample two-dimensional image; Creating an initial human face reconstruction model, and inputting the sample two-dimensional image into the initial human face reconstruction model for prediction to obtain a predicted three-dimensional human face model and a predicted texture map; Performing smoothing processing on the predicted three-dimensional human face model to obtain a smoothed predicted three-dimensional human face model; Performing two-dimensional projection on the predicted three-dimensional human face model to obtain a predicted two-dimensional image, and extracting human face key points in the predicted two-dimensional image; Construct a first cost function based on each vertex in the predicted three-dimensional face model and each vertex in the standard three-dimensional face model: where N is the number of vertices, V i is the i-th vertex in the predicted three-dimensional face model, and V GT,i is the i-th vertex in the standard three-dimensional face model; Construct a second cost function based on the vertices in the predicted three-dimensional face model and the vertices in the smoothed predicted three-dimensional face model: where N is the number of vertices, and V i is the i-th vertex in the predicted three-dimensional face model, and the is the i-th vertex in the smoothed predicted three-dimensional face model; Construct a third cost function based on the facial key points in the predicted two-dimensional image and the facial key points in the sample two-dimensional image: where M is the number of facial key points, and lmk i is the i-th facial key point in the predicted two-dimensional image, and lmk GT,i is the i-th facial key point in the sample two-dimensional image; Construct a fourth cost function based on each pixel in the predicted texture map and each pixel in the standard texture map: where the Pixel is the number of pixels, and the Tex i is the i-th pixel in the predicted texture map, and the Tex GT,i is the i-th pixel in the standard texture map; Performing weighted summation on the first cost function, the second cost function, the third cost function, and the fourth cost function to obtain a cost function; Training the initial human face reconstruction model based on the cost function to obtain a trained human face reconstruction model; Obtaining a target two-dimensional image; Inputting the target two-dimensional image into the human face reconstruction model, and outputting a target three-dimensional human face model and a target texture map corresponding to the target two-dimensional image; Generating a textured three-dimensional human face model containing texture information based on the target three-dimensional human face model and the target texture map.

2. The method according to claim 1, characterized in that, the inputting the target two-dimensional image into the trained human face reconstruction model, and outputting a target three-dimensional human face model and a target texture map corresponding to the target two-dimensional image includes: Inputting the target two-dimensional image into the trained human face reconstruction model, and outputting a stylized three-dimensional human face model and a stylized texture map corresponding to the target two-dimensional image; the generating a textured three-dimensional human face model containing texture information based on the target three-dimensional human face model and the target texture map includes: Generating a textured three-dimensional human face model containing texture information based on the stylized three-dimensional human face model and the stylized texture map.

3. The method according to claim 1, characterized in that, the obtaining a standard three-dimensional human face model and a standard texture map corresponding to the sample two-dimensional image includes: Extracting human face key points in the sample two-dimensional image, and obtaining an average face three-dimensional model; Performing constrained reconstruction on the average face three-dimensional model based on the human face key points in the sample two-dimensional image to obtain a standard three-dimensional human face model corresponding to the sample two-dimensional image; Obtaining the UV coordinates of each vertex in the standard three-dimensional human face model, and bilinearly interpolating the sample two-dimensional image based on the UV coordinates of each vertex to obtain a standard texture map of a preset size.

4. The method according to claim 3, characterized in that, the performing constrained reconstruction on the average face three-dimensional model based on the human face key points in the sample two-dimensional image to obtain a standard three-dimensional human face model corresponding to the sample two-dimensional image includes: Performing constrained reconstruction on the average face three-dimensional model based on the human face key points in the sample two-dimensional image to obtain a constrained three-dimensional human face model and a first projection parameter; Performing projection transformation on the constrained three-dimensional human face model based on the first projection parameter to obtain a standard three-dimensional human face model.

5. The method according to claim 1, wherein, the creating of the initial face reconstruction model and inputting the sample two-dimensional image into the initial face reconstruction model for prediction to obtain a predicted three-dimensional face model and a predicted texture map includes: creating the initial face reconstruction model and inputting the sample two-dimensional image into the initial face reconstruction model for prediction to obtain a deformation parameter matrix, lighting parameters, and texture parameters; predicting a predicted three-dimensional face model based on the deformation parameter matrix; performing lighting modeling based on the lighting parameters to obtain a lighting map; obtaining a first texture map based on the texture parameters and a preset texture template; and generating a predicted texture map based on the lighting map and the first texture map.

6. The method according to claim 1, wherein, the two-dimensional projection of the predicted three-dimensional face model to obtain a predicted two-dimensional image and extracting the facial key points from the predicted two-dimensional image includes: obtaining second projection parameters in the initial face reconstruction model, performing two-dimensional projection on the predicted three-dimensional face model based on the second projection parameters to obtain a predicted two-dimensional image, and extracting the facial key points from the predicted two-dimensional image.

7. The method according to claim 1, wherein, the weighted summation of the first cost function, the second cost function, the third cost function, and the fourth cost function to obtain a cost function includes: L = β·L 1 + μ·L 2 + λ·L 3 + ω·L 4 wherein, β is the weight corresponding to the first cost function L 1 μ is the weight corresponding to the second cost function L 2 λ is the weight corresponding to the third cost function L 3 ω is the weight corresponding to the fourth cost function L 4 corresponding weight.

8. A three-dimensional face model reconstruction device, wherein, the device includes: A model training module, which is used to collect sample two-dimensional images; obtain a standard three-dimensional face model and a standard texture map corresponding to the sample two-dimensional images; create an initial face reconstruction model, and input the sample two-dimensional images into the initial face reconstruction model for prediction to obtain a predicted three-dimensional face model and a predicted texture map; perform smoothing processing on the predicted three-dimensional face model to obtain a smoothed predicted three-dimensional face model; project the predicted three-dimensional face model into two dimensions to obtain a predicted two-dimensional image, and extract face key points in the predicted two-dimensional image; construct a first cost function based on each vertex in the predicted three-dimensional face model and each vertex in the standard three-dimensional face model: wherein, N is the number of vertices, and V i is the i-th vertex in the predicted three-dimensional face model, and V GT,i is the i-th vertex in the standard three-dimensional face model; construct a second cost function based on each vertex in the predicted three-dimensional face model and each vertex in the smoothed predicted three-dimensional face model: wherein N is the number of vertices, and V i is the i-th vertex in the predicted three-dimensional face model, and the is the i-th vertex in the smoothed predicted three-dimensional face model; construct a third cost function based on the face key points in the predicted two-dimensional image and the face key points in the sample two-dimensional image: wherein M is the number of face key points, and lmk i is the i-th face key point in the predicted two-dimensional image, and lmk GT,i is the i-th face key point in the sample two-dimensional image; construct a fourth cost function based on each pixel in the predicted texture map and each pixel in the standard texture map: wherein, Pixel is the number of pixels, and Tex i is the i-th pixel in the predicted texture map, and Tex GT,i is the i-th pixel in the standard texture map; perform weighted summation on the first cost function, the second cost function, the third cost function, and the fourth cost function to obtain a cost function; train the initial face reconstruction model based on the cost function to obtain a trained face reconstruction model; an image acquisition module for acquiring a target two-dimensional image; a model prediction module for inputting the target two-dimensional image into the face reconstruction model and outputting a target three-dimensional face model and a target texture map corresponding to the target two-dimensional image; a target model generation module for generating a texture three-dimensional face model containing texture information based on the target three-dimensional face model and the target texture map.

9. A storage medium having a computer program stored thereon, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer device, wherein, comprising: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Three-dimensional face reconstruction method and device, electronic equipment and storage medium

    CN112819947A

  • Three-dimensional face processing method, training method, generating method, device and equipment

    CN113538221A