Multi-branch deep learning 3D face reconstruction model training method, system and medium
Through the 3D face reconstruction model training method of multi-branch deep learning, the use of face recognition, alignment, expression recognition and generation adversarial networks jointly solves the problem that 3D face generation details are difficult to accurately represent in the prior art, and improves the authenticity and accuracy of the generated images.
Patent Information
- Application Number
- CN202210574406.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-25
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-05-25
AI Technical Summary
The prior art has the problem that geometric structural details are difficult to accurately represent when generating high-fidelity 3D faces with real textures, and the input image quality is high, so it cannot adapt to non-high-definition images. The model relies too much on training data, resulting in the generated 3D face distortion.
The 3D face reconstruction model training method with multi-branch deep learning is adopted. Through joint training of face recognition network, face alignment network, expression recognition network and generation adversarial network, the rendered image is generated and network parameters are updated to improve the authenticity and accuracy of 3D face images.
The authenticity and accuracy of the generated 3D face images can be improved, and the non-high-definition images can be better adapted to non-high-definition images, and the dependence on training data is reduced, so the generated 3D face images are more realistic.
Smart Images

Figure CN114926591B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a multi-branch deep learning 3D face reconstruction model training method, system and medium. Background Art
[0002] As one of the core research topics in the cross-fields of computer vision and machine learning, face 3D reconstruction technology has been widely used in human-computer interaction, games, animation and other fields in recent years. Face 3D reconstruction refers to the restoration of face 3D information from 2D face images, including face texture information, light reflection information, expression information, geometric shape information, etc. Traditional 3D face generation is completed by expensive capture systems or professionals. With the improvement of computer computing power, algorithm-generated 3D faces are becoming more and more realistic, and the cost is relatively low compared to traditional methods, so it has attracted the attention of many researchers.
[0003] The process of reconstructing a 3D face from a 2D face image is an uncertain problem, because the same 3D face model can generate multiple 2D images, and it is difficult to determine which one corresponds to the real 3D face. The key to success is to add prior knowledge to eliminate ambiguous solutions. Generally speaking, 3D face reconstruction methods can be divided into three types: statistical-based methods, photometric-based methods, and deep learning-based methods. The statistical-based method encodes prior knowledge in a 3D face model. The most classic one is the 3D Morphable Models (3DMM). The process of generating a 3D face by 3DMM is to solve a set of linear parameters. It consists of a shape model and optional texture and color models. The average 3D face shape and other information are obtained by principal component analysis (PCA), and then the parameters are optimized to generate a 3D face corresponding to the input 2D face image. The photometric-based method combines the 3D face model with the photometric stereo vision method to estimate the normal of the face surface. This strategy is based on modeling the reflectivity of the face surface, which will affect the quality of the reconstructed face. The original data uses information from multiple images, which will further cause ambiguity in the solution, so it is not as widely used as the other two methods. The deep learning-based method learns prior knowledge from a large amount of raw data, that is, directly learns the mapping between 2D images and 3D faces, and then outputs high-quality 3D face information. The emergence of this method has led to great development in face 3D reconstruction technology.
[0004] At present, deep learning-based methods can be divided into four categories according to different neural network architectures: 3D face reconstruction algorithms based on convolutional neural networks, 3D face reconstruction algorithms based on autoencoders, 3D face reconstruction algorithms based on graph convolutional networks, and 3D face reconstruction algorithms based on generative adversarial networks (GAN). Among them, GAN has been proven to generate images with real features when trained on 2D face images, and obtain photo-realistic high-resolution faces. There are also more and more GAN algorithms trying to generate texture maps for 3D faces.
[0005] However, it is still technically difficult to generate a high-fidelity 3D face with realistic textures. Wrinkles and other geometric details are important indicators of age and facial expressions, and are essential for generating realistic virtual humans. Although the PCA processing model used in the 3DMM algorithm has its advantages, it is limited by the capacity of linear space and cannot fully represent high-frequency information, which usually causes the texture model to be too smooth, resulting in facial texture distortion; some existing algorithms directly perform super-resolution processing on the input image in order to obtain high-resolution texture maps, but this method has high requirements on the quality of the input image and is not in line with reality, that is, it is not suitable for non-HD images obtained by ordinary equipment; some algorithms use a large amount of high-quality UV data as a training set to train GAN. Although the effect is very good, the algorithm is overly dependent on training data.
[0006] In addition, textures, geometric shapes and expressions should have potential correlation information. If the model parameters are trained independently, the rendered image may lose its realism. Therefore, some algorithms directly train all parameters through the network and complete data alignment directly in the UV space without the need for additional conversion into 3DMM parameter form. However, this method cannot train all parameters perfectly, and the model is relatively complex. When the input is a continuous sequence of video frames, existing algorithms rarely have measures to prevent occlusion, so they are not robust to occlusion. Summary of the invention
[0007] The purpose of the present invention is to solve one of the technical problems existing in the prior art to at least a certain extent.
[0008] To this end, an object of an embodiment of the present invention is to provide a multi-branch deep learning 3D face reconstruction model training method, which improves the authenticity and accuracy of the generated 3D face image.
[0009] Another object of an embodiment of the present invention is to provide a multi-branch deep learning 3D face reconstruction model training system.
[0010] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present invention include:
[0011] In a first aspect, an embodiment of the present invention provides a multi-branch deep learning 3D face reconstruction model training method, comprising the following steps:
[0012] Acquire a first face image, input the first face image into a pre-built face recognition network to obtain first identity information, and input the first face image into a pre-built face alignment network to obtain first key point position information;
[0013] Determine first facial geometric shape information according to the first identity information, and input the first identity information and the first key point position information into a pre-constructed expression recognition network to obtain first facial expression information;
[0014] Inputting the first key point position information, the first face geometry information, and the first face expression information into a pre-constructed generative adversarial network to obtain a first rendered image;
[0015] The network parameters of the face recognition network, the face alignment network, the expression recognition network and the generative adversarial network are updated according to the first rendered image to obtain an optimal parameter combination, and then a 3D face reconstruction model is obtained according to the face recognition network, the face alignment network, the expression recognition network, the generative adversarial network and the optimal parameter combination.
[0016] Furthermore, in one embodiment of the present invention, the face recognition network is a FaceNet network, the face alignment network is an MTCNN network, and the expression recognition network is a lightweight RingNet network.
[0017] Further, in one embodiment of the present invention, the step of determining the first face geometric shape information according to the first identity information is specifically:
[0018] The first face image is subjected to feature extraction and dimensionality reduction processing by a principal component analysis algorithm to obtain a reduced dimension matrix, and the first face geometric shape information is determined according to the first identity information and the reduced dimension matrix.
[0019] Furthermore, in one embodiment of the present invention, the generative adversarial network includes a generator and a discriminator, the generator includes a texture generation module and a rendering module, the generator is used to generate a rendered image based on the first key point position information, the first face geometry information, the first face expression information and preset parameters of the generative adversarial network, and the discriminator is used to update the face recognition network, the face alignment network, the expression recognition network and the network parameters of the generative adversarial network through a back propagation algorithm according to the rendered image output by the generator.
[0020] Furthermore, in one embodiment of the present invention, the step of inputting the first key point position information, the first face geometry information, and the first face expression information into a pre-constructed generative adversarial network to obtain a first rendered image specifically includes:
[0021] Inputting the first key point position information into the texture generation module to obtain a first texture map;
[0022] Performing super-resolution processing on the first texture map to obtain a second texture map;
[0023] Determine a texture normal vector according to the second texture map, determine a face geometry normal vector according to the first face geometry information, and determine a face expression normal vector according to the first face expression information;
[0024] Inputting the texture normal vector, the face geometry normal vector and the face expression normal vector into the rendering module to obtain a first normal map;
[0025] Performing differentiable rendering on the first normal map to obtain a first rendered image.
[0026] Further, in one embodiment of the present invention, the step of updating the network parameters of the face recognition network, the face alignment network, the expression recognition network, and the generative adversarial network according to the first rendered image to obtain the optimal parameter combination specifically includes:
[0027] Inputting the first rendered image into the discriminator, and calculating a loss value according to a preset loss function;
[0028] Update the network parameters of the face recognition network, the face alignment network, the expression recognition network, and the generative adversarial network through a gradient descent algorithm and a back propagation algorithm according to the loss value;
[0029] When the loss value reaches a preset first threshold, or the number of iterations reaches a preset second threshold, or the test accuracy reaches a preset third threshold, training is stopped to obtain the optimal parameter combination.
[0030] Furthermore, in one embodiment of the present invention, the loss function is:
[0031] L=m K L K +m P L P +m id L id +m f L f +m R L R
[0032] Among them, L represents the loss value, L K represents the alignment loss, m K represents the weight of the alignment loss, L P represents the perceptual loss, m P Represents the weight of perceptual loss, L id Indicates identity information loss, m id Represents the weight of identity information loss, L f Indicates the video continuity loss, m f Represents the weight of video continuity loss, L R represents the regularization loss, m R Represents the weight of the regularization loss.
[0033] In a second aspect, an embodiment of the present invention provides a multi-branch deep learning 3D face reconstruction model training system, comprising:
[0034] An identity information and key point position information determination module, used to obtain a first face image, input the first face image into a pre-built face recognition network to obtain first identity information, and input the first face image into a pre-built face alignment network to obtain first key point position information;
[0035] A face geometry information and face expression information determination module, used to determine first face geometry information according to the first identity information, and input the first identity information and the first key point position information into a pre-built expression recognition network to obtain first face expression information;
[0036] A rendering image determination module, used for inputting the first key point position information, the first face geometry information and the first face expression information into a pre-constructed generative adversarial network to obtain a first rendering image;
[0037] A network parameter optimization module is used to update the network parameters of the face recognition network, the face alignment network, the expression recognition network and the generative adversarial network according to the first rendered image to obtain an optimal parameter combination, and then obtain a 3D face reconstruction model according to the face recognition network, the face alignment network, the expression recognition network, the generative adversarial network and the optimal parameter combination.
[0038] In a third aspect, an embodiment of the present invention provides a 3D face reconstruction model training device for multi-branch deep learning, comprising:
[0039] at least one processor;
[0040] at least one memory for storing at least one program;
[0041] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned multi-branch deep learning 3D face reconstruction model training method.
[0042] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores a program executable by a processor, and when the program executable by the processor is executed by the processor, it is used to execute the above-mentioned multi-branch deep learning 3D face reconstruction model training method.
[0043] The advantages and beneficial effects of the present invention will be partly given in the following description, partly become apparent from the following description, or be understood through the practice of the present invention:
[0044] The embodiment of the present invention forms a multi-branch deep learning 3D face reconstruction model through a face recognition network, a face alignment network, an expression recognition network and a generative adversarial network. The face image is first input into the face recognition network and the face alignment network to obtain identity information and key point position information, then the face geometry information is determined according to the identity information, and the identity information and key point position information are input into the expression recognition network to obtain face expression information, and then the key point position information, face geometry information and face expression information are input into the generative adversarial network to obtain a rendered image, and the network parameters of the face recognition network, the face alignment network, the expression recognition network and the generative adversarial network are updated based on the rendered image and a preset loss function until the optimal parameter combination is obtained, thereby obtaining a trained 3D face reconstruction model. The embodiment of the present invention updates the network parameters of each branch network through joint training of multiple branch networks, and can learn the potential associations between the face features of each branch network, maintain the correlation of the face features of multiple modalities, and improve the authenticity and accuracy of the generated 3D face image. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solution in the embodiments of the present invention, the following introduction is made to the drawings required for use in the embodiments of the present invention. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solution of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0046] Figure 1 A flowchart of the steps of a multi-branch deep learning 3D face reconstruction model training method provided by an embodiment of the present invention;
[0047] Figure 2 A schematic diagram of the training process of a 3D face reconstruction model using multi-branch deep learning provided by an embodiment of the present invention;
[0048] Figure 3 A structural block diagram of a multi-branch deep learning 3D face reconstruction model training system provided by an embodiment of the present invention;
[0049] Figure 4 A structural block diagram of a multi-branch deep learning 3D face reconstruction model training device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0051] In the description of the present invention, the meaning of "a plurality" is two or more than two. If there is a description of "a first" or "a second", it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used in this document have the same meaning as those commonly understood by those skilled in the art.
[0052] The traditional 3DMM model uses UV mapping for 3D face reconstruction. The texture information contained in each vertex is stored in the UV coordinates. The UV space defines the information of each pixel in the image, and these points are interconnected with the 3D model. Texture can reflect the surface properties of an object, and the parameter space value is converted to the texture UV space through a mapping function. This process is called mapping, that is, texture mapping. By performing PCA analysis on the vectorized UV mapping, the average basis of the mapping can be obtained. Specifically, the face appearance model S of the 3DMM model model and texture model T model They are:
[0053]
[0054]
[0055] Among them, the face appearance model S model Including geometric shape models and expression models, represents the average shape vector calculated based on the faces in the dataset. Similarly, is the average texture vector; U s , U e , U t Represent the basis subset of PCA analysis and the predicted shape parameter α s , expression parameter α e , texture parameter α t A linear combination of .
[0056] The parameter fitting process of 3DMM can be regarded as the following optimization problem:
[0057]
[0058] Among them, I 0 Represents the input 2D face image, I R represents the rendered image, α=[α s ,α e ,α t ,α l ,α c ], α l and α c Represent the illumination parameters and camera parameters respectively, ∑α s,e,l 2 Represents the regularization process, used to constrain the parameter α s , α e and α t , preventing the normal range of the average human face and the appearance of lighting that deviates from reality.
[0059] Since the principle of the 3DMM model is relatively simple, it is not complicated to fit various attributes (texture, shape, etc.) to the 3DMM, and the 3D face reconstructed by combining 3DMM with a deep learning algorithm can have more texture details than the face reconstructed using only 3DMM. Therefore, the embodiment of the present invention adopts the framework of 3DMM, obtains facial features by improving the extraction methods of various attributes, and optimizes the 3D face reconstruction model through joint training of multiple branch networks, so as to generate a 3D face reconstructed image with high-precision detail information. Each branch network of the embodiment of the present invention is used to estimate a single attribute (identity, expression, texture and other features). In this way, each branch can focus on one task to improve accuracy. Considering that features such as geometric shape, texture and expression have potential correlation, they can be trained separately until each branch network converges to a good weight, and then each branch network is connected and jointly trained to obtain a better combination of network parameters.
[0060] Reference Figure 1 The embodiment of the present invention provides a multi-branch deep learning 3D face reconstruction model training method, which specifically includes the following steps:
[0061] S101, obtaining a first face image, inputting the first face image into a pre-built face recognition network to obtain first identity information, and inputting the first face image into a pre-built face alignment network to obtain first key point position information.
[0062] Specifically, both the face recognition network and the face alignment network can use existing neural network models, and use the face image training set to pre-train until the model converges, and then start the joint training process of the embodiment of the present invention. When a 2D face image is input, the face identity information and face key point position information need to be extracted through the above two networks, and then respectively input into other branch networks for subsequent model training.
[0063] S102, determining first facial geometry information according to the first identity information, and inputting the first identity information and first key point position information into a pre-built expression recognition network to obtain first facial expression information.
[0064] As a further optional implementation, the face recognition network is a FaceNet network, the face alignment network is an MTCNN network, and the expression recognition network is a lightweight RingNet network.
[0065] Specifically, the embodiment of the present invention extracts facial identity information of an input 2D image through a face recognition network FaceNet. The principle is to use a convolutional neural network to learn Euclidean space features. The smaller the Euclidean distance between the feature vectors of two images, the greater the possibility that the two images are of the same person.
[0066] The embodiment of the present invention performs facial key point detection through a face alignment network MTCNN, which is a CNN-based cascade detection method that realizes face detection and key point calibration in real time.
[0067] The lightweight RingNet network adopted in the embodiment of the present invention is based on the RingNet network and adds SE modules (Squeeze and Excite Modules). The RingNet network is a multi-encoder-decoder architecture, which can capture 3D facial expressions and can be used for animation driving. The role of the SE module is to reduce the model size and complexity without affecting the accuracy of the results, so that the algorithm can meet the needs of real-time applications. Since the output of RingNet is not in the form of 3DMM parameters, the embodiment of the present invention converts it into 3DMM parameter form at the network output layer to meet the requirements of data alignment. Similar to the pre-training of the face recognition network and the face alignment network, the embodiment of the present invention pre-trains the lightweight RingNet network by learning the characteristics of different expressions (such as happiness, sadness, surprise, disgust and fear, etc.), and then performs the multi-branch network joint training of the embodiment of the present invention.
[0068] As an optional implementation, the step of determining the first face geometric shape information according to the first identity information is specifically as follows:
[0069] The first face image is subjected to feature extraction and dimensionality reduction processing by a principal component analysis algorithm to obtain a reduced dimension matrix, and the first face geometric shape information is determined according to the first identity information and the reduced dimension matrix.
[0070] Specifically, principal component analysis is divided into the following steps:
[0071] 1. Data preprocessing;
[0072] 2. Obtain the covariance matrix;
[0073] 3. Obtain the eigenvalue and eigenvector of the covariance matrix;
[0074] 4. According to the size of the eigenvalue, select an appropriate number of eigenvectors as a basis to form a subspace;
[0075] 5. Project the original matrix into the subspace to obtain the reduced dimension matrix.
[0076] In the embodiment of the present invention, the first face image is subjected to feature extraction and dimensionality reduction processing by principal component analysis to obtain a dimensionality reduction matrix, and then the face geometry information can be determined by combining the identity information obtained in the above steps.
[0077] S103: Inputting the first key point position information, the first face geometry information and the first face expression information into a pre-built generative adversarial network to obtain a first rendered image.
[0078] Further as an optional implementation, the generative adversarial network includes a generator and a discriminator, the generator includes a texture generation module and a rendering module, the generator is used to generate a rendered image based on the first key point position information, the first face geometry information, the first face expression information and preset parameters of the generative adversarial network, and the discriminator is used to update the network parameters of the face recognition network, the face alignment network, the expression recognition network and the generative adversarial network through a back propagation algorithm according to the rendered image output by the generator.
[0079] As an optional implementation, the step of inputting the first key point position information, the first face geometry information, and the first face expression information into a pre-built generative adversarial network to obtain a first rendered image specifically includes:
[0080] S1031, inputting the first key point position information into a texture generation module to obtain a first texture map;
[0081] S1032, performing super-resolution processing on the first texture map to obtain a second texture map;
[0082] S1033, determining a texture normal vector according to the second texture map, determining a face geometry normal vector according to the first face geometry information, and determining a face expression normal vector according to the first face expression information;
[0083] S1034, inputting the texture normal vector, the face geometry normal vector and the face expression normal vector into a rendering module to obtain a first normal map;
[0084] S1035 . Perform differentiable rendering on the first normal map to obtain a first rendered image.
[0085] Specifically, an embodiment of the present invention proposes a generative adversarial network MB-GAN, which is composed of a generator and a discriminator. The generator includes two modules, namely a texture generation module and a rendering module.
[0086] The texture generation module is used to generate high-resolution texture maps and perform super-resolution processing. The generated texture maps are used to replace the texture model in 3DMM. By training the MB-GAN network with a high-resolution UV texture map dataset, a higher-quality texture map can be generated. In view of the problem that there is still room for quality improvement of the generated texture maps, the embodiment of the present invention further performs super-resolution processing on the texture maps to ensure that the texture details are richer. Super-resolution processing refers to the amplification of image resolution. The embodiment of the present invention adopts the latest RealSR algorithm. By amplifying the resolution of the generated texture map by 8 times, a clearer UV texture map is obtained.
[0087] The rendering module is used to generate normal maps and perform differentiable rendering. Normal maps are composed of normal vectors of face geometry, facial expression, and texture. As an extension of bump texture, normal maps include a lot of detailed surface information. The normal value of each pixel can be used for lighting calculation. When rendering, the bump effect of the texture is represented by using lighting parameters and the normal value of the point. After completing differentiable rendering, it can be input into the discriminator to update the network parameters.
[0088] In addition, similar to the pre-training of the aforementioned face recognition network, face alignment network, and expression recognition network, the MB-GAN network of the embodiment of the present invention also needs to be pre-trained before the joint training of the multi-branch network. Based on different tasks, the training data sets of each branch network are different, among which the training set of the MB-GAN network is composed of large-scale 3D textures, which are composed of high-resolution texture maps generated from 3 perspectives (left, front, and right) of 1,000 different people after processing. In addition, the MB-GAN network samples camera and lighting parameters from the Gaussian distribution of the AFLW2000-3D data set, which contains 2,000 3D face images, each of which contains the corresponding 3DMM coefficients and 68 face key points; the RealFaceDB data set can be used to train the expression recognition network, which includes more than 200 faces of people of different ages and characteristics under 7 different expressions.
[0089] S104, updating the network parameters of the face recognition network, the face alignment network, the expression recognition network and the generative adversarial network according to the first rendered image to obtain an optimal parameter combination, and then obtaining a 3D face reconstruction model according to the face recognition network, the face alignment network, the expression recognition network and the generative adversarial network and the optimal parameter combination.
[0090] Specifically, the function of the MB-GAN network of the embodiment of the present invention is to render and generate high-fidelity 3D face images that conform to the 3DMM parameter distribution. During training, the geometric shapes and expressions output by each branch network, as well as the texture, camera, lighting and other parameters generated by the MB-GAN network itself, are integrated and input into the rendering module. Differentiable rendering is used to decouple the above-obtained parameters and back-propagate them to each branch network to update the network parameters, thereby reducing the difference between the rendered image and the real image. After the above training, the MB-GAN network can generate the optimal parameter combination, so that the input 2D face image can be input into the MB-GAN network after being processed by each branch network and finally render a high-fidelity 3D face image. The 3D face reconstruction model of the embodiment of the present invention is composed of various branch networks in the optimal parameter combination state.
[0091] Further as an optional implementation, the step of updating the network parameters of the face recognition network, the face alignment network, the expression recognition network, and the generative adversarial network according to the first rendered image to obtain the optimal parameter combination specifically includes:
[0092] A1. Input the first rendered image into the discriminator, and calculate the loss value according to the preset loss function;
[0093] A2. Update the network parameters of the face recognition network, face alignment network, expression recognition network, and generative adversarial network through the gradient descent algorithm and back propagation algorithm according to the loss value;
[0094] A3. When the loss value reaches the preset first threshold, or the number of iterations reaches the preset second threshold, or the test accuracy reaches the preset third threshold, stop training and obtain the optimal parameter combination.
[0095] As an optional implementation, the loss function is:
[0096] L=m K L K +m P L P +m id L id +m f L f +m R L R
[0097] Among them, L represents the loss value, L K represents the alignment loss, m K represents the weight of the alignment loss, L P represents the perceptual loss, m P Represents the weight of perceptual loss, L id Indicates identity information loss, m idRepresents the weight of identity information loss, L f Indicates the video continuity loss, m f Represents the weight of video continuity loss, L R represents the regularization loss, m R Represents the weight of the regularization loss.
[0098] Specifically, in the loss function of the embodiment of the present invention, the alignment loss L K It is used to ensure the consistency of the facial key point alignment between the rendered image and the input image. In addition to the loss at the image level, the perceptual loss at the feature level should also be considered. The cosine distance between the feature representations of the input image extracted by the FaceNet network and the rendered image can be calculated to obtain the perceptual loss L. P , used to improve parameter quality; identity information loss L id It is used to ensure that the identity information of the input image obtained by face recognition is still completely preserved in the rendered image; when the input image is a continuous video frame, the identity feature information of the image rendered by the previous frame is retained, and the identity information of the image rendered by the current frame is calculated with it to obtain the video continuity loss L f , which is used to prevent the current frame from being occluded, resulting in a large deviation between the rendered image and the previous frame; in order to ensure the authenticity of the geometric shape and texture of the reconstructed 3D face, a regularization loss L is proposed R , which is used to force the parameters of each branch network output to obey the normal distribution of 3DMM:
[0099] L R =w s ||α s ||+w e ||α e ||+w t ||α t ||
[0100] Among them, w s 、w e and w t Respectively represent the weights of geometric shape features, facial expression features, and texture features, α s , α e and α t They represent geometric shape features, facial expression features and texture features respectively.
[0101] The embodiment of the present invention simultaneously optimizes the network parameters of all branch networks through the gradient descent algorithm and the back propagation algorithm to minimize the weighted combination of the above loss terms (ie, the loss value of the embodiment of the present invention) to achieve the purpose of model fitting.
[0102] The above is a description of the method steps of the embodiment of the present invention. Figure 2The figure shows a training flow diagram of a multi-branch deep learning 3D face reconstruction model provided by an embodiment of the present invention. First, a 2D face image is input, identity information is collected through the face recognition network FaceNet, and the position of key points of the face is detected through the face alignment network MTCNN; then the face geometry information of the 3DMM is obtained by the PCA method, and the identity information and key point position information are input into the expression recognition network to obtain the facial expression information, and then the key point position information, face geometry information and facial expression information are input into the MB-GAN network proposed in the embodiment of the present invention to generate high-precision texture maps and normal maps, and a rendered image is generated by differentiable rendering, and then the loss value is calculated, and the network parameters of all branch networks are updated by the back propagation algorithm. After reaching the preset convergence condition, the optimal parameter combination is obtained, and a trained 3D face reconstruction model can be obtained at this time. A high-fidelity 3D face image can be rendered by directly inputting a 2D face image into the 3D face reconstruction model. It can be appreciated that the embodiment of the present invention updates the network parameters of each branch network through joint training of multiple branch networks, and can learn the potential correlation between the facial features of each branch network, thereby maintaining the correlation of facial features of multiple modalities and improving the authenticity and accuracy of the generated 3D face image.
[0103] It can be understood that, for the problem of texture map quality, the embodiment of the present invention trains the proposed MB-GAN with a texture UV data set, and then performs super-resolution processing after generating a high-precision texture map to further improve the texture accuracy. Such processing can balance the quality requirements for the training data set and the input image; for the problem of potential correlation information between texture, geometric shape and expression features, the embodiment of the present invention designs multiple loss terms, integrates the features of different branch outputs in MB-GAN for back propagation, and iteratively optimizes the final parameter combination. Such a design can maximize the advantages of each branch network and learn the potential connection between each model parameter; for the real-time problem, the embodiment of the present invention adds a feature selection module, which helps to reduce the scale of calculation while ensuring the accuracy of the model; for the occlusion problem, in addition to the fact that the training data set of the expression model generation network is a human face with occlusion, the embodiment of the present invention also designs a video continuity loss, which uses the time information between video frames to effectively deal with occlusion. In addition, the embodiments of the present invention fully utilize the advantages of the generative renderer, using differentiable rendering technology to generate images similar to the identity information of the input image, and inferring more accurate and continuous facial geometry and texture, further improving the authenticity and accuracy of the generated 3D face image.
[0104] Reference Figure 3 The embodiment of the present invention provides a multi-branch deep learning 3D face reconstruction model training system, comprising:
[0105] An identity information and key point position information determination module, used to obtain a first face image, input the first face image into a pre-built face recognition network to obtain first identity information, and input the first face image into a pre-built face alignment network to obtain first key point position information;
[0106] A face geometry information and face expression information determination module, used to determine the first face geometry information according to the first identity information, and input the first identity information and the first key point position information into a pre-built expression recognition network to obtain the first face expression information;
[0107] A rendering image determination module, used to input the first key point position information, the first face geometry information and the first face expression information into a pre-built generative adversarial network to obtain a first rendering image;
[0108] The network parameter optimization module is used to update the network parameters of the face recognition network, the face alignment network, the expression recognition network and the generative adversarial network according to the first rendered image, obtain the optimal parameter combination, and then obtain the 3D face reconstruction model according to the face recognition network, the face alignment network, the expression recognition network and the generative adversarial network and the optimal parameter combination.
[0109] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0110] Reference Figure 4 The embodiment of the present invention provides a 3D face reconstruction model training device for multi-branch deep learning, comprising:
[0111] at least one processor;
[0112] at least one memory for storing at least one program;
[0113] When the at least one program is executed by the at least one processor, the at least one processor implements the multi-branch deep learning 3D face reconstruction model training method.
[0114] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0115] An embodiment of the present invention also provides a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to execute the above-mentioned multi-branch deep learning 3D face reconstruction model training method.
[0116] A computer-readable storage medium according to an embodiment of the present invention can execute a multi-branch deep learning 3D face reconstruction model training method provided by an embodiment of the method of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.
[0117] The embodiment of the present invention also discloses a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 The method shown.
[0118] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the above-mentioned boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.
[0119] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified to the contrary, one or more of the above-mentioned functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0120] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the above methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0121] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0122] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the above-mentioned program is printed, since the above-mentioned program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or processing in other suitable ways as necessary, and then stored in a computer memory.
[0123] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0124] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0125] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
[0126] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A multi-branch deep learning 3D face reconstruction model training method, characterized in that: The following steps are involved: Acquire a first face image, input the first face image into a pre-built face recognition network to obtain first identity information, and input the first face image into a pre-built face alignment network to obtain first key point position information; Determine first facial geometric shape information according to the first identity information, and input the first identity information and the first key point position information into a pre-constructed expression recognition network to obtain first facial expression information; Inputting the first key point position information, the first face geometry information, and the first face expression information into a pre-constructed generative adversarial network to obtain a first rendered image; Update the network parameters of the face recognition network, the face alignment network, the expression recognition network, and the generative adversarial network according to the first rendered image to obtain an optimal parameter combination, and then obtain a 3D face reconstruction model according to the face recognition network, the face alignment network, the expression recognition network, the generative adversarial network, and the optimal parameter combination; The generative adversarial network includes a generator and a discriminator, the generator includes a texture generation module and a rendering module, the generator is used to generate a rendered image according to the first key point position information, the first face geometry information, the first face expression information and the preset parameters of the generative adversarial network, and the discriminator is used to update the face recognition network, the face alignment network, the expression recognition network and the network parameters of the generative adversarial network through a back propagation algorithm according to the rendered image output by the generator; The step of inputting the first key point position information, the first face geometry information, and the first face expression information into a pre-built generative adversarial network to obtain a first rendered image specifically includes: Inputting the first key point position information into the texture generation module to obtain a first texture map; Performing super-resolution processing on the first texture map to obtain a second texture map; Determine a texture normal vector according to the second texture map, determine a face geometry normal vector according to the first face geometry information, and determine a face expression normal vector according to the first face expression information; Inputting the texture normal vector, the face geometry normal vector and the face expression normal vector into the rendering module to obtain a first normal map; Performing differentiable rendering on the first normal map to obtain a first rendered image.
2. The multi-branch deep learning 3D face reconstruction model training method according to claim 1, characterized in that: The face recognition network is a FaceNet network, the face alignment network is an MTCNN network, and the expression recognition network is a lightweight RingNet network.
3. The multi-branch deep learning 3D face reconstruction model training method according to claim 1, characterized in that: The step of determining the first face geometric shape information according to the first identity information is specifically as follows: The first face image is subjected to feature extraction and dimensionality reduction processing by a principal component analysis algorithm to obtain a reduced dimension matrix, and the first face geometric shape information is determined according to the first identity information and the reduced dimension matrix.
4. The multi-branch deep learning 3D face reconstruction model training method according to claim 1, characterized in that: The step of updating the network parameters of the face recognition network, the face alignment network, the expression recognition network, and the generative adversarial network according to the first rendered image to obtain an optimal parameter combination specifically includes: Inputting the first rendered image into the discriminator, and calculating a loss value according to a preset loss function; Update the network parameters of the face recognition network, the face alignment network, the expression recognition network, and the generative adversarial network through a gradient descent algorithm and a back propagation algorithm according to the loss value; When the loss value reaches a preset first threshold, or the number of iterations reaches a preset second threshold, or the test accuracy reaches a preset third threshold, training is stopped to obtain the optimal parameter combination.
5. The multi-branch deep learning 3D face reconstruction model training method according to claim 4, characterized in that: The loss function is: L=m K L K +m P L P +m id L id +m f L f +m R L R Among them, L represents the loss value, L K represents the alignment loss, m K represents the weight of the alignment loss, L P represents the perceptual loss, m P Represents the weight of perceptual loss, L id Indicates identity information loss, m id Represents the weight of identity information loss, L f Indicates the video continuity loss, m f Represents the weight of video continuity loss, L R represents the regularization loss, m R Represents the weight of the regularization loss.
6. A multi-branch deep learning 3D face reconstruction model training system, characterized in that: include: An identity information and key point position information determination module, used to obtain a first face image, input the first face image into a pre-built face recognition network to obtain first identity information, and input the first face image into a pre-built face alignment network to obtain first key point position information; A face geometry information and face expression information determination module, used to determine first face geometry information according to the first identity information, and input the first identity information and the first key point position information into a pre-built expression recognition network to obtain first face expression information; A rendering image determination module, used for inputting the first key point position information, the first face geometry information and the first face expression information into a pre-constructed generative adversarial network to obtain a first rendering image; A network parameter optimization module, used to update the network parameters of the face recognition network, the face alignment network, the expression recognition network and the generative adversarial network according to the first rendered image, obtain an optimal parameter combination, and then obtain a 3D face reconstruction model according to the face recognition network, the face alignment network, the expression recognition network, the generative adversarial network and the optimal parameter combination; The generative adversarial network includes a generator and a discriminator, the generator includes a texture generation module and a rendering module, the generator is used to generate a rendered image according to the first key point position information, the first face geometry information, the first face expression information and the preset parameters of the generative adversarial network, and the discriminator is used to update the face recognition network, the face alignment network, the expression recognition network and the network parameters of the generative adversarial network through a back propagation algorithm according to the rendered image output by the generator; The rendered image determination module is specifically used for: Inputting the first key point position information into the texture generation module to obtain a first texture map; Performing super-resolution processing on the first texture map to obtain a second texture map; Determine a texture normal vector according to the second texture map, determine a face geometry normal vector according to the first face geometry information, and determine a face expression normal vector according to the first face expression information; Inputting the texture normal vector, the face geometry normal vector and the face expression normal vector into the rendering module to obtain a first normal map; Performing differentiable rendering on the first normal map to obtain a first rendered image.
7. A multi-branch deep learning 3D face reconstruction model training device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a multi-branch deep learning 3D face reconstruction model training method as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to execute a multi-branch deep learning 3D face reconstruction model training method as described in any one of claims 1 to 5 when executed by the processor.
Citation Information
Patent Citations
Nonlinear 3DMM face reconstruction and posture normalization method and device, medium and equipment
CN112215050A
3D face reconstruction system and method
WO2020165557A1