A method for extracting 3D face representations from images and videos
By constructing a deep learning neural network model, the three-dimensional representation of the face was extracted, and the problem that two-dimensional features in the existing technology was difficult to decouple face factors, and efficient and accurate three-dimensional face representation extraction was achieved, which improved the expressiveness of face tasks.
Patent Information
- Application Number
- CN202210427450.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-04-21
AI Technical Summary
The prior art uses only two-dimensional features in facial representation learning, making it difficult to effectively decouple potential factors of change, resulting in poor facial performance and affecting the performance of downstream tasks.
A three-dimensional face representation extraction method is proposed. By constructing a deep learning neural network model, the shape, material, expression, lighting and posture characteristics of the face are extracted using an encoder, and the face image is reconstructed through the expression transformation module and the renderer, and the gap between the reconstructed image and the input image is evaluated using a new loss function.
It realizes efficient and accurate extraction of three-dimensional face representation from images and videos, and can decouple up to 5 face representation factors, including expressions, shapes, materials, lighting and postures, improving the expressiveness of face tasks and the accuracy of downstream tasks.
Smart Images

Figure CN114708586B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image and video understanding, and in particular relates to a three-dimensional face representation extraction method. Background Art
[0002] Faces play a crucial role in human visual perception, indispensable for conveying identity, information, expression, and intent. Neural networks are widely used to understand faces in computer vision tasks, including face recognition, facial expression recognition, pose estimation, and face reconstruction. However, these efforts focus on the performance of each task while neglecting the overall understanding of the face, and require large amounts of labeled data. Facial representation learning addresses this shortcoming and can be used as a pre-training method for face tasks. It uses unsupervised learning from unlabeled samples.
[0003] Self-supervised models are supervised only by information from the samples themselves and learn to extract their internal structure from the data. Self-supervised learning is widely used in computer vision tasks, including classification, detection, generation, and 3D reconstruction. Various types of network architectures have been proposed for these tasks to obtain better representations: generative models, such as autoencoders (AE) and variational autoencoders (VAE), and adversarial models, such as generative adversarial networks (GANs). Representation learning is one of the most important topics in self-supervised learning and is also an independent field aimed at improving data features and improving downstream predictors. Representation learning algorithms have been applied to many machine learning tasks, such as language models, graph neural networks, and visual tasks. Transferable explanatory factors are the standard for representation learning, and disentangled representations are also an important topic and a lot of work has been done.
[0004] A good face representation can disentangle the underlying factors of variation. Current methods use only two-dimensional features and are limited in addressing facial factors. This leads to poor performance on faces, which in turn leads to poor performance on downstream tasks. In reality, facial images are composed of many three-dimensional structural factors, including internal factors such as facial expression, shape, and material, as well as external factors such as lighting and pose.
[0005] 3D face modeling with textures has been studied for a long time. One of the most widely used methods is 3DMorphable Model (3DMM) [1], which has been subsequently improved by many methods [2,3]. Face models are obtained by PCA of 3D scans, which requires a lot of manpower. And the representation space is limited by the model. It is difficult to generalize these methods to natural face images. To improve it, Unsup3d [4] and Lifting Autoencoders [5] proposed unsupervised face reconstruction algorithms. Then, [6] used labeled identities to achieve better reconstruction. However, the above methods did not explore the potential of 3D face models in representation learning.
[0006] Face representation learning aims to obtain better representations for face tasks. Many supervised learning methods have been proposed to solve this problem, but they require a large amount of training data [7,8]. Some recent face representation works have used 3DMM [9,10]. They require less supervision information but require a 3D face prior. GAN is an unsupervised representation learning method, and some papers have followed this approach
[11] . However, existing work is limited to a certain dataset and it is difficult to extract universal face representations for building classifiers. Summary of the Invention
[0007] The purpose of the present invention is to provide a method for extracting three-dimensional facial representations from images and videos so as to efficiently and accurately perform facial expression recognition, posture estimation, face verification and face frontalization.
[0008] In the present invention, the three-dimensional face representation includes internal factors and external factors; wherein the internal factors refer to: shape, expression, and material; and the external factors refer to: posture and lighting.
[0009] The method provided by the present invention for extracting three-dimensional facial representations from images and videos constructs a three-dimensional unsupervised facial representation learning network model, which is a deep learning neural network model. The specific steps of the present invention are as follows:
[0010] (1) Use the encoder to extract the shape, material, expression, lighting, and posture features of the face from the input image I, including:
[0011] Using the shape encoder E s Extract face shape code C shape , using the material encoder E t Extracting facial texture code C texture , using the expression encoder E e Extract facial expression code C expr , using the illumination encoder E l Extract illumination code C light, using the pose encoder E p Extract pose code C pose ;
[0012] (2) Use the expression transformation module W to transform the estimated face material and shape; the expression transformation module W can make the extracted facial expression code C expr Influence of face shape coding C shape and face texture coding C texture The composition of the image makes the extracted material code and shape code different depending on the expression;
[0013] (3) Reconstruct the face image based on the extracted code; first use the material generator G t , from the extracted material code C texture Generate face texture map M t ; Use shape generator G s , from the extracted face shape code C shape Generate face depth map M s ; Then, use the renderer R to make the face texture map M t , face depth map M s , Light Coding C light , Posture Code C pose Synthesize new face images The renderer R mainly includes two processes: lighting and projection;
[0014] (4) Use a new loss function to evaluate the reconstructed image The gap between the face image and the input image I; first, a confidence map generator is used to predict the confidence of the face area in the image, and the confidence is used to guide the loss function to focus on the face area; the VGG
[12] network is also used to extract low-level and high-level semantic features of the face image to calculate the loss;
[0015] (5) Pre-train the model using a single image;
[0016] By constructing a neural network learning framework and optimizing it with certain constraints, we can extract the shape C from the encoder. shape 、Material C texture , light C light and posture C pose Four factors; finally, use this network to input the face image, predict the face representation of the face image, and then judge the face posture and frontal appearance;
[0017] (6) Continue training the model using the video;
[0018] Use certain constraints to optimize and extract expression C from the encoder expr , Shape Cshape 、Material C texture , light C light and posture C pose Five factors; use this network to input a video frame sequence, predict the facial representation in the video frame, and then judge factors such as facial expression, posture, shape, etc.
[0019] Further:
[0020] In step (1), the material encoder E t and shape encoder E s It is a feature encoder with the structure as Figure 4 As shown in Figure 2, the two encoders have the same architecture. They are both convolutional neural networks using batch normalization layers. The input image I with three channels of R, G, and B can generate a 256-dimensional encoding vector C. shape and C texture ,Right now:
[0021] C shape =E s (I),
[0022] C texture =E t (I), (1).
[0023] In step (1), the finger illumination encoder E l and pose encoder E p It is a digital encoder, whose structure consists of Figure 5 As shown in Figure 2. They are also convolutional neural networks that generate corresponding lighting and pose parameters to control the subsequent rendering work. The two encoders have the same structure except that the output encoding dimensions are slightly different. Pose encoder E p Generated code C pose There are 6 dimensions, namely the translation vector and rotation angle of the three-dimensional coordinates; the illumination encoder E l Generated code C light There are 4 dimensions: ambient light parameters, diffuse reflection parameters and two lighting directions x, y. The output of the encoder uses the activation function Tanh to expand the final output to between -1 and 1, and then map it to the corresponding space, that is:
[0024] C light =E l (I), C pose =E p (I) (2).
[0025] In step (1), the expression encoder E eUnlike other encoders, the above encoders are prone to failure when training expression extraction due to the small and unstable gradient during back propagation. Therefore, the present invention uses a ResNet18
[12] with a residual structure to extract features. Like the feature encoder, a 256-dimensional encoding vector C can be generated for an R, G, B three-channel input image I. expr ,Right now:
[0026] C expr =E e (I) (3).
[0027] In step (2), the expression transformation module W is a key module for modeling facial expressions in the video. The specific process is as follows:
[0028] First, the face shape code C is sampled from a series of video frames. shape and face texture coding C texture , averaging them in the feature space to obtain and It is assumed that the neutral expression face can be estimated by the average value of the sample sequence.
[0029] Then, use the obtained facial expression code C expr As shape C shape and material C texture The linear deviation of the parameter is added to the obtained average code to obtain the transformed code C′ shape and C′ texture Therefore, the process of expression transformation module W can be expressed as the following formula:
[0030]
[0031]
[0032] Among them, the symbol Indicates that x is averaged over the batch dimension.
[0033] W input is material code C texture or shape coding C shape , and expression code C expr ; W outputs the transformed shape code C′ shape and material code C′ texture In this way, these feature solutions are divided into the sequence variation part and the sequence invariant part in the sequence, and the gradients are calculated separately:
[0034]
[0035]
[0036] Among them, the gradient ΔC of the i-th material encoding texture,i Gradients from all material encodings in a sequence The average value of the shape code is ΔC. texture,i Gradients from all shape encodings in a sequence The average value of the gradient of the i-th expression code ΔC expr,i Gradient from the corresponding shape encoding and the corresponding material-encoded gradient The sum of t is the scaling factor of the material expression effect, λ s is the scaling factor of the shape expression effect, usually, λ s =λ t = 1; ||V|| is the length of the input video sequence V; when the generator is fixed, the expression transformation module W is encoded from the face shape C shape and face texture coding C texture Learning the effects of facial expressions on texture encoding during variations within the same video.
[0037] In step (3), the material generator G t and shape generator G s , whose network structure includes stacked convolutional layers, transposed convolutional layers and group normalization layers. Figure 6 The detailed structure is shown in . The network uses a 256-dimensional vector as input, and the material generator G t Finally, a 3-channel material map M is generated t Output, shape generator G s Generate 1-channel face depth map M s Output. Final material map M t and depth map M s Use the Tanh function to scale to the range of -1 to 1, that is:
[0038] M t =G t (C texture ), M s =G s (C shape ), (8)
[0039] In step (3), the renderer R receives the material map M t and depth map M s , and lighting code C light , Posture Code C pose As a parameter, it can be expressed as the following formula.
[0040]
[0041] is the reconstructed image, and R represents the rendering process, which mainly includes two processes: illumination and projection. In the rendering process, first, the depth map M s Converted into a 3D mesh in the 3D rendering pipeline; then, the material map M t Fuse with the mesh to get a realistic representation of the 3D model.
[0042] Furthermore, the lighting process of the renderer R described in step (3) uses a simplified Phong lighting model, which is an empirical model of local lighting. Through this lighting model, the color I of each point p can be obtained from the following equation p :
[0043] I p =k a,p +∑ m∈lights k d,p (L m ·N p ), (10)
[0044] Among them, lights represents the collection of all light sources, L m Represents the direction vector from a point m on the surface to each light source, N p Represents the depth map M s The normal to the surface is obtained directly. k a,p is the ambient light coefficient at point p, k d,p The model of the present invention ignores the specular reflection of the face, because in most cases, the specular reflection coefficient of the face is so small that it can be ignored compared with the diffuse reflection. The direction and intensity of the light source are encoded by the illumination code C light Provided, the diffuse reflection coefficient of point p is determined by the material map M t supply.
[0045] Furthermore, the projection process of the renderer described in step (3) uses a weak perspective camera model, that is, the light should be orthogonal to the camera plane. Under perspective projection, the conversion between the imaged 2D point p and the actual 3D point position P is as follows:
[0046] p=s c K[R c t c ]P, (11)
[0047] Where K is the internal parameter of the camera. c and t c are the external parameters of rotation and translation, s c is the camera's zoom factor. R c and t c Can be encoded from the pose C pose When the material map Mt and depth map M s , and lighting code C light , Posture Code C pose After illumination and projection by the renderer R, a two-dimensional reconstructed image of the face can be obtained.
[0048] In step (4), the reconstruction loss includes the constraints from low pixel level to high feature level. The loss function consists of three parts: photometric loss L p , feature-level loss L f and identity loss L i :
[0049] (1) Luminosity loss L p , feature-level loss L f It can be expressed as follows:
[0050]
[0051]
[0052] Where I represents the input image, Represents the reconstructed image. Conv represents the low-level feature extraction network, which is processed by extracting the relu3_3 features from a pre-trained VGG-19 network
[12] . Figure 7 As shown, the present invention uses the encoder-generator structure to generate the confidence map, denoted as σ, σ p is the confidence map of photometric loss, σ f Is the confidence map of the feature-level loss. In this model, the photometric loss and the feature-level loss are constrained by the estimated confidence map σ, and the confidence-based evaluation function L conf The model can be made to calibrate itself:
[0053]
[0054] Among them, L conf There are three parameters to reconstruct the image Input image I, and confidence σ. Ω is the effective area, that is, the reconstructed image The non-background part of |Ω| represents the number of points in the effective area. uv represents the point coordinates in the effective area, σ uv is the confidence level of the point, Indicates that the point is in the reconstructed image and the pixel RGB difference on the input image I, ∑ uv∈Ω Indicates the sum of all points in the valid area. exp represents the natural exponential operation, and ln represents the natural logarithm operation.
[0055] (2) L of the above loss p and L f The integrity of the face is not emphasized, which can easily lead to the face identity features of the reconstructed image being far different from the original image. Therefore, the present invention also uses an identity loss L i To constrain the perceptual consistency of the entire face. First, use the function g to reconstruct the image I and the original image Combined, the function uses the corresponding part of the original image to fill the background that is not in the reconstructed image. Then, the perceptual similarity (LPIPS)
[14] between the supplemented image and the original image is calculated. LPIPS attempts to extract multi-layer features in the VGG network
[12] to calculate the distance between images. Identity L i The loss can be expressed as follows.
[0056]
[0057] Among them, f is the VGG network of perceptual similarity, g is the filling function, and g has two parameters to reconstruct the image Input image I. Since the background area of the reconstructed image is missing, it is filled with the corresponding area of input image I. <·> represents the cosine distance, and ||·|| represents the modulus.
[0058] (3) Loss function It can be expressed as a linear combination of photometric loss, lower feature loss and identity loss:
[0059]
[0060] Among them, λ f and λ i is the weight of feature-level loss and identity loss, usually taken as λ f =λ i =1.
[0061] (4) The present invention also considers calculating the loss of the reconstructed image of the left and right faces horizontally flipped, and linearly summing it with the above losses to obtain the final loss function:
[0062]
[0063] in, is the reconstructed image of the left and right faces flipped horizontally, L tot Is the final loss, flip is the material map M s and depth map M t Perform left and right flooding transformation, λ flip is the weight of flipping and reconstructing the image, usually taken as λ flip = 0.5. Then, the backpropagation algorithm is used to calculate the gradient of the network from the loss and update the network parameters.
[0064] In step (5), the model is pre-trained using the image set. The specific process is: the image in the large face image set is passed through the encoder in step (1) to obtain the face shape code C shape , material code C texture , illumination code C light , posture coding C pose . Then use the generator in step (3) to encode C from the shape shape , material code C texture Get the material map M t and depth map M s , the formula is as follows:
[0065] M s =G s (C shape )=G s (E s (I)), (18)
[0066] M t =G t (C texture )=G t (E t (I)), (19)
[0067] Where I is the input image. s 、E t They are shape and material encoders, G s , G t They are shape and texture generators. Then the material map M t and depth map M s With illumination coding C light , posture coding C pose Generate reconstructed images through the renderer R
[0068] Then, calculate the loss function L according to step (4) tot To perform backpropagation, train all encoders and generators (except the expression encoder).
[0069] During testing, you only need to use the encoder to extract the face shape code C shape , material code C texture , illumination code C light , posture coding C pose Then we can perform subsequent tasks such as pose estimation, face verification and face frontalization. For example, pose estimation only requires the input image I to pass through the pose encoder E p , get the posture code C pose , we can get the predicted pose from the parameters.
[0070] In step (6), the model continues to train after adding the expression transformation module using the video, wherein the input frames are collected from the same video sequence, and they have different expressions and postures. The difference from the step (5) is that the expression transformation module W here can extract the expression code C expr Used to process material code C texture and shape encoding C s ape , get the transformed shape code C′ shape and material code C′ texture . Generate depth M s and texture map M t The process can be expressed by the following formula:
[0071] M s =G s (C′ shape )=G s (W(C shape ,C expr ))=G s (W(E s (I),E e (I))), (20)
[0072] M t =G t (C′ texture )=G t (W(C texture ,C expr ))=G t (W(E t (I),E e (I))), (21)
[0073] Where I is the input image. s 、E t They are shape and material encoders, G s , G t They are shape and texture generators respectively.
[0074] Combined with the previously extracted posture C pose and light C light Information, depth and material maps can be used to generate reconstructed images through the renderer R When the image set model training is completed, the expression transformation module W and expression encoder E can be easily added to the image set model. e, continue training on the video to obtain a model suitable for the video. The important reason why the present invention uses video for modeling expression is that the faces in the video naturally have the same identity and makeup. No annotation is required. At the same time, video frames contain a large number of expression changes and are easily decoupled. During testing, it is only necessary to pass the image through the corresponding encoder to decouple the posture, lighting, shape, expression and material from the face video, and assist in the prediction of various downstream tasks.
[0075] The advantages of the present invention are as follows:
[0076] (1) This paper proposes a novel 3D-based unsupervised face representation learning model framework. This model can learn disentangled 3D face representations from unlabeled image sets and natural videos. Existing face representation learning methods are limited to 2D features.
[0077] (2) This paper proposes a novel unsupervised strategy that uses an expression transformation module to learn 3D facial expressions from unannotated video sequences. 3D facial expressions typically require a 3D face prior to be acquired, but this paper can separate 3D facial expressions from identity features without any labels or face priors.
[0078] (3) The model of the present invention adds new geometric information and explores potential environmental factors. The model framework of the present invention can discover and decouple up to five facial representation factors, including expression, shape, material, lighting and posture. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 Schematic diagram of face representation learning based on three-dimensional decoupling technology.
[0080] Figure 2 A diagram of the neural network structure.
[0081] Figure 3 This is the architecture diagram of the feature encoder.
[0082] Figure 4 This is the architectural diagram of the digital encoder.
[0083] Figure 5 This is the structure diagram of the expression transformation module.
[0084] Figure 6 Diagram of the architecture of the material and depth map generator.
[0085] Figure 7 Architecture diagram of the confidence map generator.
[0086] Figure 8 A visualization of the intermediate results. DETAILED DESCRIPTION
[0087] (1) We use the CelebA dataset
[15] to learn from image datasets and the VoxCeleb dataset
[16] to learn from videos. We crop the CelebA and VoxCeleb datasets using FaceNet
[17] and resize them to 128×128. The proposed model is implemented in the PyTorch framework and trained using the Adam optimizer. Both the encoder and decoder are fully convolutional networks. The batch size is set to 16, and the learning rate for both training stages is 0.0001. We train the model on 30 epochs of image datasets and video sequences, respectively.
[0088] (2) In the rendering process of the present invention, the face shape is a two-dimensional single-channel matrix, which represents the depth map of the face. The present invention defines a grid with the same size as the image, which is 128×128. Their x-axis and y-axis coordinates are scaled to between -1 and 1, and then the z-axis coordinate comes from the depth map. In this way, the present invention obtains a three-dimensional model of the face and the normal of each point for subsequent calculations. The face material is represented in the renderer as a three-channel two-dimensional matrix. Represents the diffuse reflectivity of each grid point on each surface of the RGB ray. The lighting includes ambient light intensity, diffuse reflection intensity and light direction x, y. The present invention is modeling directional light, so only two variables are needed to describe the direction of the light. In general, the present invention first constructs a three-dimensional skeleton of the face shape, and then maps the face material to the three-dimensional skeleton. Then the present invention uses the light information to determine the color of the face. Finally, the present invention uses the camera formula to obtain the image taken by the present invention at a specific angle, which is equivalent to changing the posture of the face.
[0089] (3) The model framework of the present invention is composed of Figure 1 As shown, it aims to separate the material, shape, expression, pose, and lighting of the face from unlabeled face images and videos. 3D decomposition is used to decompose internal factors (bottom) and external factors (top). The dotted line indicates that the facial expression is learned from the changes in the video sequence. This framework can benefit many downstream tasks, color representation and the connection between tasks.
[0090] (4) The neural network structure of the present invention is composed of Figure 2 As shown, the input image I is input to the encoder E t 、E s 、E p and E l , which extract material, shape, pose and lighting encoding respectively. t and depth map M s By using the generator G t and G s Material code C textureand shape encoding C shape Generate a shadow of the depth map for better visualization. Finally, these two maps together with the pose C pose and light C light The parameters are passed through the renderer to obtain the final reconstructed image When learning from videos, there is an additional expression encoder E e . Extracted expression code C expr The expression transformation module W affects the material and shape coding to generate the true shape coding C′ shape and material code C′ texture The extracted codes will be used for downstream tasks such as facial expression recognition and face verification. The model of the present invention does not require any supervisory information or 3DMM face model.
[0091] (5) The architecture of the encoder of the present invention is composed of Figure 3 、 4 As shown, convolution (a, b, c) indicates that the convolution layer has a kernel size of a, a stride of b, and a padding of c. The number below the convolution layer indicates the number of convolution kernels. The number below the group normalization layer indicates the number of groups. The yellow arrow represents the LeakyRelu activation function with a slope of 0.2. The blue arrow represents the Relu.
[0092] (6) The expression transformation module structure of the present invention is composed of Figure 5 As shown. Taking material features as an example, first encode the material C texture The average is then added to the expression parameters to form the final output encoding. During the backpropagation process, the feature gradients of the sequence changes will flow to the expression encoder, while the feature parts that do not change the sequence will flow to the material and shape encoder.
[0093] (7) The architecture of the generator of the present invention is composed of Figure 6 、 7 As shown, Convolution(a,b,c) and ConvolutionT(a,b,c) indicate convolutional and transposed convolutional layers with kernel size a, stride b, and padding c, respectively. The number below the module indicates the number of convolution kernels. The number below the group normalization layer indicates the number of groups. The yellow arrow indicates the Leaky Relu activation function with a slope of 0.2. The blue arrow is Relu. The red arrow is a SoftPlus operator. The shorter path in the confidence map generator is used for feature-level loss, and the longer path is used for photometric loss.
[0094] (8) The generation result of the present invention is Figure 8 As shown in the figure, from left to right: input image, neutral depth map, depth map, neutral face shape, face shape, neutral texture map, texture map, and reconstructed image. The shape image is obtained by shading the 3D face model (i.e., depth map).
[0095] References
[0096] [1] Volker Blanz and Thomas Vetter. 1999. A morphable model for the synthesis of 3D faces. In SIGGRAPH’99.
[0097] [2] Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong. 2019. Accurate 3D Face Reconstruction with Weakly-Supervised Learning: From Single Image to Image Set. In IEEE Computer Vision and Pattern Recognition Workshops.
[0098] [3] Yao Feng, Haiwen Feng, Michael J. Black, and Timo Bolkart. 2021. Learning an animatable detailed 3D face model from in-the-wild images. ACM Transactions on Graphics (TOG) 40(2021), 1–13.
[0099] [4] Shangzhe Wu, C. Rupprecht, and Andrea Vedaldi. 2020. Unsupervised Learning of Probably Symmetric Deformable 3D Objects From Images in the Wild. 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020), 1–10.
[0100] [5]Mihir Sahasrabudhe,Zhixin Shu,Edward Bartrum,Riza Alp Güler,DimitrisSamaras,and Iasonas Kokkinos.2019.Lifting AutoEncoders:UnsupervisedLearning of a Fully-Disentangled3D Morphable Model Using DeepNon-RigidStructureFrom Motion.2019 IEEE / CVF International Conference onComputerVision Workshop(ICCVW)(2019),4054–4064.
[0101] [6]Yujun Shen,Jinjin Gu,Xiaoou Tang,and Bolei Zhou.2020.InterpretingtheLatent Space of GANs for Semantic Face Editing.2020IEEE / CVF ConferenceonComputer Vision and Pattern Recognition(CVPR)(2020),9240–9249.
[0102] [7]Luan Tran,Xi Yin,and Xiaoming Liu.2017.Disentangled RepresentationLearning GAN for Pose-Invariant Face Recognition.2017 IEEE Conference onComputerVision and Pattern Recognition(CVPR)(2017),1283–1292.
[0103] [8]Huiyuan Yang,UmurAybarsCiftci,and Lijun Yin.2018.FacialExpressionRecognition by De-expression Residue Learning.2018 IEEE / CVFConference onComputer Vision and Pattern Recognition(2018),2168–2177.
[0104] [9]Zhongpai Gao,Juyong Zhang,Yudong Guo,Chao Ma,GuangtaoZhai,andXiaokang Yang.2020.Semi-supervised 3D Face Representation LearningfromUnconstrained Photo Collections.2020IEEE / CVF Conference on ComputerVisionand Pattern Recognition Workshops(CVPRW)(2020),1426–1435
[0105]
[10] Feng Liu,Qijun Zhao,Xiaoming Liu,and Dan Zeng.2020.Joint FaceAlignment and 3D Face Reconstruction with Application to Face Recognition.IEEETransactionson Pattern Analysis and Machine Intelligence 42(2020),664–678.
[0106]
[11] Thu Nguyen-Phuoc,Chuan Li,Lucas Theis,Christian Richardt,andYongliangYang.2019.HoloGAN:Unsupervised Learning of 3D Representations FromNatural Images.2019IEEE / CVF International Conference on Computer Vision(ICCV)(2019),7587–7596.
[0107]
[12] Kaiming He,Xiangyu Zhang,Shaoqing Ren,and Jian Sun.2016.DeepResidualLearning for Image Recognition.(2016),770–778. https: / / doi.org / 10.1109 / CVPR.2016.90
[0108]
[13] Karen Simonyan and Andrew Zisserman.2015.Very Deep ConvolutionalNetworks for Large-Scale Image Recognition.(2015).http: / / arxiv.org / abs / 1409.1556
[0109]
[14] Richard Zhang,Phillip Isola,Alexei A.Efros,Eli Shechtman,andOliver Wang.2018.The Unreasonable Effectiveness of Deep Features as aPerceptual Metric.In2018 IEEE / CVF Conference on Computer Vision and PatternRecognition.586–595. https: / / doi.org / 10.1109 / CVPR.2018.00068
[0110]
[15] Ziwei Liu,Ping Luo,Xiaogang Wang,and Xiaoou Tang.2015.DeepLearningFace Attributes in the Wild.In 2015 IEEE International Conference onComputerVision(ICCV).3730–37
[0111]
[16] Arsha Nagrani,Joon Son Chung,and Andrew Zisserman.2017.VoxCeleb:ALarge-Scale Speaker Identification Dataset.In Proc.Interspeech 2017.2616–262
[0112]
[17] Florian Schroff,Dmitry Kalenichenko,and James Philbin.2015.FaceNet:Aunified embedding for face recognition and clustering。
Claims
1. A method for extracting three-dimensional face representations from images and videos, wherein: The three-dimensional face representation includes internal factors and external factors; the internal factors refer to shape, expression, and material; the external factors refer to posture and lighting; it is characterized in that a three-dimensional unsupervised face representation learning network model is constructed to extract the three-dimensional face representation, and the specific steps are: (1) Using the encoder in the network model to extract the shape, material, expression, lighting and posture features of the face from the input image I, specifically including: Using the shape encoder E s Extract face shape code C shape , using the texture encoder E t Extracting face texture code C texture , using the expression encoder E e Extract facial expression code C expr , using the illumination encoder E l Extracting illumination code C light , using the pose encoder E p Extract pose code C pose ; (2) using the expression transformation module W in the network model to transform the estimated face material and shape; including using the expression transformation module W to encode the extracted face expression C expr Influence of face shape coding C shape and face texture coding C texture The composition of the image makes the extracted material code and shape code different according to different expressions; (3) Reconstructing the face image based on the extracted code; First, use the material generator G in the network model t , from the extracted material code C texture Generate face texture map M t ; Use shape generator G s , from the extracted face shape code C shape Generate face depth map M s ; Then, use the renderer R to make the face texture map M t , face depth map M s , Light coding C light , Posture Coding C pose Synthesizing new face images The renderer R process includes two processes: lighting and projection; (4) Use a loss function to evaluate the reconstructed image The difference between the face image and the input image I; first, the confidence map generator in the network model is used to predict the confidence of the face area in the image, and the confidence is used to guide the loss function to focus on the face area; the VGG network is also used to extract low-level and high-level semantic features of the face image to calculate the loss; (5) pre-training the network model using a single image; Based on the constructed network model, the constraints are optimized to extract the shape C from the encoder shape 、Material C texture 、Light C light and posture C pose Four factors; finally, the network model is used to input the face image, predict the face representation of the face image, and then judge the face posture and frontal appearance; (6) continue to train the network model using the video; Using constraints to optimize, extract expression C from the encoder expr , Shape C shape 、Material C texture 、Light C light and posture C pose Five factors; use the network model to input a video frame sequence, predict the face representation in the video frame, and then judge the facial expression, posture, and shape factors; In step (2): the expression transformation module W has the following operation flow: First, the face shape code C is sampled from a series of video frames obtained in step (1). shape and face texture coding C texture , averaging them in the feature space to obtain and Assume that the neutral expression face is estimated by the average of the sampled sequence; Then, use the obtained facial expression code C expr As shape C shape and material C texture The linear deviation of the parameter is added to the obtained average code to obtain the transformed code C′ shape and C′ texture ; Therefore, the process of expression transformation module W is expressed as the following formula: Among them, the symbol Indicates that x is averaged over the batch dimension; W input is the material code C texture or shape encoding C shape , and expression code C expr ; W outputs the transformed shape code C′ shape and material code C′ texture ; In this way, these feature solutions are divided into the sequence variation part and the sequence invariant part in the sequence, and the gradients are calculated separately: Among them, the gradient ΔC of the i-th material encoding texture,i Gradients from all material encodings in a sequence The average value of the gradient of the i-th shape code is ΔC shape,i Gradients from all shape encodings in a sequence The average value of the i-th expression code gradient ΔC expr,i Gradient from the corresponding shape encoding and the corresponding material-encoded gradient The sum of t is the scaling factor of the material expression effect, λ s is the proportional factor of the shape expression effect, usually, λ s =λ t = 1; ||V|| is the length of the input video sequence V; when the generator is fixed, the expression transformation module W encodes the face shape C shape and face texture coding C texture Learning the effects of facial expressions on texture encoding during variations within the same video.
2. The method for extracting three-dimensional face representation from images and videos according to claim 1, characterized in that: In step (1): The material encoder E t and shape encoder E s It is a feature encoder. These two encoders have the same architecture and use a convolutional neural network with a batch normalization layer. The input image I with three channels of R, G, and B can generate a 256-dimensional encoding vector C shape and C texture ,Right now: C shape =E s (I), C texture =E t (I), (1) The light encoder E l and the pose encoder E p are digital encoders, both are convolutional neural networks, which generate corresponding lighting and posture parameters to control the subsequent rendering work; these two encoders have the same structure; the posture encoder E p Generated code C pose There are 6 dimensions, namely the translation vector and rotation angle of the three-dimensional coordinates; the illumination encoder E l Generated code C light There are 4 dimensions: ambient light parameters, diffuse reflection parameters and two lighting directions x, y; the output of the two encoders uses the activation function Tanh to expand the final output to between -1 and 1, and then maps it to the corresponding space, namely: C light =E l (I),C pose =E p (I) (2) The expression encoder E e A ResNet18 with a residual structure is used to extract features; Like the feature encoder, a 256-dimensional encoding vector C can be generated for an R, G, B three-channel input image I expr ,Right now: C expr =E e (I) (3)。 3. The method for extracting three-dimensional face representation from images and videos according to claim 1, characterized in that: In step (3): The material generator G t and shape generator G s , whose network structure includes stacked convolutional layers, transposed convolutional layers, and group normalization layers; the network uses a 256-dimensional vector as input, and the material generator G t Finally, a 3-channel material map M is generated. t Output, shape generator G s Generate 1-channel face depth map M s Output; Final texture map M t and depth map M s Scaling to the range of -1 to 1 using the Tanh function is expressed as: M t =G t (C texture ),M s =G s (C shape ), (8) The renderer R receives the material map M t and depth map M s , and the illumination code C light , Posture Coding C pose As a parameter, it is expressed as the following formula; is the reconstructed image, R is the rendering process, which mainly includes two processes: illumination and projection. In the rendering process, first, the depth map M s Converted into a 3D mesh in the 3D rendering pipeline; then, the material map M t Fuse with the mesh to get a realistic representation of the 3D model.
4. The method for extracting three-dimensional face representation from images and videos according to claim 3, characterized in that: In step (3): The lighting process of the renderer R uses a simplified Phong lighting model, through which the color I of each point p is obtained from the following equation: p : I p =k a,p +∑ m∈lights k d,p (L m ·N p ), (10) Among them, lights represents the collection of all light sources, L m represents the direction vector from a point m on the surface to each light source, N p Represents the depth map M s The normal directly obtained to the surface; k a,p is the ambient light coefficient at point p, k d,p is the diffuse reflectance of point p; the direction and intensity of the light source are encoded by the light code C light Provided, the diffuse reflection coefficient of point p is determined by the material map M t supply; The projection process of the renderer uses a weak perspective camera model, that is, the light is orthogonal to the camera plane. Under perspective projection, the conversion relationship between the imaged two-dimensional point p and the actual three-dimensional point position P is as follows: p=s c K[R c t c ]P, (11) Among them, K is the internal parameter of the camera; R c and t c is an external parameter, s c is the camera zoom factor; R c and t c From the pose code C pose Get from; Material map M t and depth map M s , and the illumination code C light , Posture Coding C pose After illumination and projection by the renderer R, a two-dimensional reconstructed image of the face is obtained.
5. The method for extracting three-dimensional face representation from images and videos according to claim 4, characterized in that: In step (4), the reconstruction loss includes constraints from low pixel level to high feature level; The loss function consists of three parts: photometric loss L p , feature-level loss L f and identity loss L i : (a) Luminosity loss L p , feature-level loss L f It is expressed as follows: Where I represents the input image, represents the reconstructed image, conv represents the low-level feature extraction network; σ represents the confidence map, which is generated using the encoder-generator structure, σ p is the confidence map of photometric loss, σ f is the confidence map of the feature-level loss; in this network model, the photometric loss and the feature-level loss are constrained by the estimated confidence map σ, and the confidence-based evaluation function L conf Make the model calibrate itself: Among them, L conf There are three parameters to reconstruct the image Input image I, and confidence σ; Ω is the effective area, that is, the reconstructed image The non-background part, |Ω| represents the number of points in the effective area; uv represents the point coordinates in the effective area, σ uv is the confidence level of the point, Indicates that the point is in the reconstructed image and the pixel RGB difference on the input image I, ∑ uv∈Ω It means summing all points in the valid area; (b) Identity loss L i It is used to constrain the perceptual consistency of the entire face; first, the reconstructed image I is compared with the original image using the function g Combined, this function uses the corresponding part of the original image to fill the background that the reconstructed image does not have; then, the perceptual similarity (LPIPS) between the supplemented image and the original image is calculated. The perceptual similarity attempts to extract multi-layer features in the VGG network to calculate the distance between images; identity L i The loss is expressed as follows: Among them, f is the VGG network that perceives similarity, g is the filling function, and g has two parameters: reconstructed image and input image I. Since the background area of the reconstructed image is missing, the corresponding area of the input image I is used for filling; <·> represents the cosine distance, and ||·|| represents the modulus; (c) Loss function It is expressed as a linear combination of photometric loss, lower feature loss and identity loss: Among them, λ f and λ i is the weight of feature-level loss and identity loss; (d) Finally, we also consider calculating the loss of the reconstructed image with the left and right faces horizontally flipped, and linearly add it to the above loss to get the final loss function: in, is the reconstructed image of the left and right faces flipped horizontally, L tot is the final loss, flip is to M the material map t and depth map M s Perform left and right panning transformation, λ flip is the weight of flipping the reconstructed image; Then, the back-propagation algorithm is used to calculate the gradient of the network from the loss and update the parameters of the network.
6. The method for extracting three-dimensional face representation from images and videos according to claim 5, characterized in that: In step (5), a single image is used to pre-train the model. The specific process is as follows: the image data in the large face image is passed through the encoder in step (1) to obtain the face shape code C shape , material code C texture , illumination code C light , posture coding C pose ; Then use the generator in step (3) to encode C from shape shape , material code C texture Get the material map M t and depth map M s , the formula is as follows: M s =G s (C shape )=G s (E s (I)), (18) M t =G t (C texture )=G t (E t (I)), (19) Where I is the input image, E s 、E t They are shape and material encoders, G s , G t They are shape and texture generators; material map M t and depth map M s With illumination coding C light , posture coding C pose Generate the reconstructed image through the renderer R Then, calculate the loss function L according to step (4) tot To perform backpropagation, train all encoders and generators, except the expression encoder; During testing, only the encoder is used to extract the face shape code C shape , material code C texture , illumination code C light , posture coding C pose Then, subsequent tasks such as pose estimation, face verification and face frontalization can be performed.
7. The method for extracting three-dimensional face representation from images and videos according to claim 6, characterized in that: In step (6), the model continues to be trained after adding the expression transformation module using the video pair, wherein the input frames are collected from the same video sequence, and they have different expressions and postures; different from the step in step (5), the expression transformation module W here converts the extracted expression code C expr Used to process material encoding C texture and shape encoding C shape , get the transformed shape code C′ shape and material code C′ texture ; Generate depth map M s and texture map M t The process is expressed by the following formula: M s =G s (C′ shape )=G s (W(C shape ,C expr ))=G s (W(E s (I),E e (I))), (20) M t =G t (C′ texture )=G t (W(C texture ,C expr ))=G t (W(E t (I),E e (I))), (21) Where I is the input image; E s 、E t They are shape and material encoders, G s , G t They are shape and texture generators; Combined with the previously extracted posture C pose and light C light information, depth and material maps can be used to generate a reconstructed image through the renderer R When the image set model training is completed, it is easy to add the expression transformation module W and the expression encoder E on the image set model. e , continue training on the video to obtain a model suitable for the video; when testing, you only need to pass the picture through the corresponding encoder to decouple the posture, lighting, shape, expression and material from the face video, and assist in the prediction of various downstream tasks.
Citation Information
Patent Citations
Self-adaptive three-dimensional face reconstruction method based on single image
CN111489435A
Micro-renderer-based method for acquiring reflection material of human face from single image
WO2021223134A1