Semantic information guided face depth data estimation model, system and method
By generating pseudo-label data and introducing face semantic information, the problems of few face depth data sets and low depth estimation accuracy are solved, and the accuracy of the face depth estimation model and the training sample coverage are improved.
Patent Information
- Application Number
- CN202510037255.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, the face depth data set is small, and the depth estimation accuracy of the training data is insufficient, resulting in a low accuracy of the face depth estimation algorithm.
A wider range of training samples are provided by generating pseudo-label data using label-free data and data augmenting of label-free data. At the same time, face semantic information, that is, face analytical features, is added to the face depth estimation model to improve the accuracy of depth estimation.
By generating a large amount of pseudo-label data and introducing face semantic information, the accuracy of training sample coverage and depth estimation of the face depth estimation model is improved, and the problems of few face depth data sets and low depth estimation accuracy are solved.
Smart Images

Figure CN120107326A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face depth estimation, and in particular to a face depth data estimation model, system and method guided by semantic information. Background Art
[0002] Monocular face depth estimation is a very important technology in the field of computer vision, which aims to restore the three-dimensional attributes of a face from a single face image. The typical application of face depth estimation is face recognition.
[0003] Monocular facial depth estimation based on deep learning (DL) has been intensively studied and developed. Improvement of facial depth estimation is an important research topic for fast and low-cost 3D facial applications. Compared with ordinary scenes, human faces contain fine structures. Facial recognition and other facial depth applications require features that can distinguish one face from another, which puts higher requirements on the refined facial depth.
[0004] The quality of training data has a significant impact on the generalization and reliability of deep learning models. In order to improve the accuracy of facial depth estimation, more high-quality and more diverse scene types of data are needed. However, the facial depth estimation datasets currently available are quite small, and creating a new dataset is time-consuming and expensive. Some researchers use various software to generate a large number of images for facial depth estimation. However, due to the small amount of training data and the insufficient accuracy of the depth information collected in the training data, the algorithm has a relatively low accuracy in estimating the depth of the face.
[0005] The patent application with publication number CN117152806A discloses a cross-domain facial expression recognition method, system, device and storage medium, and its technical solution includes: the constructed model includes a feature extractor, an independent adversarial learning module, an independent pseudo-label generation module, a global-local consistent prediction module and an independent classification learning module, and the feature extractor extracts multiple feature vectors including global features and local features from the image; the independent adversarial learning module distinguishes whether the image is from the source domain or the target domain according to the extracted multiple feature vectors; the independent pseudo-label generation module generates a pseudo-label according to the feature vector extracted from the target domain image, and then obtains the pseudo-label of the target domain image through the global-local consistent prediction module; the independent classification learning module trains the model according to the source domain data set, the target domain data set and the corresponding pseudo-label; the image to be predicted is input into the trained model and the pseudo-label is output. However, the technical solution of this patent lacks constraints when extracting features, so the extracted features are not accurate enough, which leads to inaccurate estimated depth. Summary of the invention
[0006] The purpose of the present invention is to overcome the problems in the prior art of small facial depth data sets and insufficient depth estimation accuracy of training data, which leads to relatively low accuracy of the algorithm in estimating the depth of the face, and to provide a facial depth data estimation model, system and method guided by semantic information.
[0007] In a first aspect, the present invention provides a method for estimating face depth data guided by semantic information, comprising the following steps: S1, generating pseudo-label data using unlabeled data; S2, performing data enhancement on the unlabeled data; S3, training a face depth estimation model using the enhanced data, and using a face parsing model for supervision during the training process; S4, performing face depth estimation using the trained face depth estimation model.
[0008] According to a preferred implementation, step S1 at least includes: S11, training a large depth estimation model using a data set synthesized from a three-dimensional face model, so that the large depth estimation model has the ability to estimate the absolute depth of the face model.
[0009] According to a preferred implementation, step S1 further includes: using a trained large depth estimation model to perform depth estimation on unlabeled data, generating a pseudo depth map, and forming pseudo label data.
[0010] According to a preferred implementation, step S2 at least includes: performing data enhancement on the unlabeled data to expand more face RGB images. The unlabeled data at least includes an open source real face data set. The strong data enhancement is divided into high-intensity color perturbation and spatial perturbation. The high-intensity color perturbation includes color jittering, Gaussian blur, and large-scale downsampling. The spatial perturbation is to randomly splice two pictures.
[0011] According to a preferred implementation, step S3 includes: first training a face parsing model including a face parsing encoder and a decoder, and then using the trained face parsing encoder to train a face depth estimation model. The face depth estimation model includes a face depth estimation encoder and a decoder. When training the face depth estimation model, the cosine similarity of the output features of the face depth estimation encoder and the features output by the face parsing encoder is calculated. A threshold is set, and when the cosine similarity is greater than the threshold, the face depth estimation model will not update the network parameters.
[0012] According to a preferred implementation, the data processing architecture of the face parsing encoder is: using a convolution-based block embedding layer to perform convolution batch embedding on the face image, and then using 2 Transformer Blocks to perform the first feature extraction; using the first downsampling layer to process the data of the first feature extraction, and then using 2 TransformerBlocks to perform the second feature extraction; using the second downsampling layer to process the data of the second feature extraction, and then using 2 Transformer Blocks to perform the third feature extraction to obtain face parsing features.
[0013] According to a preferred implementation, the Transformer Block includes two branches, one is a multi-head self-attention branch, and the other is a convolution branch; wherein the structure of the convolution branch is an hourglass-shaped convolution layer.
[0014] According to a preferred implementation, the architecture of the face depth estimation encoder-decoder is as follows: a depth estimation encoder, which obtains a feature map through a convolution operation, and then performs downsampling twice. Two ViT Blocks are used for feature extraction before and after sampling. Each stage is connected by a maximum pooling and downsampled by 2 times; a depth estimation decoder, which obtains the output of the depth estimation encoder and performs upsampling twice. Two ViT Blocks are used for feature extraction before and after sampling, and then the depth map is output through the output convolution layer.
[0015] The present invention also provides a semantic information-guided human face depth data estimation model. The human face depth data estimation model can implement the semantic information-guided human face depth data estimation method provided by the present invention.
[0016] The present invention also provides a semantic information-guided human face depth data estimation system. The human face depth data estimation system can implement the semantic information-guided human face depth data estimation method provided by the present invention.
[0017] Compared with the prior art, the present invention has the following beneficial effects: The present invention generates pseudo-label data from unlabeled data and performs data enhancement on the unlabeled data, thereby providing a wider range of training samples for the face depth estimation model and solving the problem of the small number of face depth data sets. In addition, the present invention adds face semantic information, namely face parsing features, to the face depth estimation model. This feature can effectively provide the depth estimation model with accurate pixel classification information of any face image, which is conducive to the depth estimation model to estimate the depth more accurately. And this additional information can also be used as a supplement to the pseudo-label data, so that the model has auxiliary information to help it predict the depth of the face when facing the pseudo-label data that has been data enhanced. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A training schematic diagram of a semantic information-guided face depth data estimation model of the present invention; Figure 2 A schematic diagram of the architecture of a face parsing encoder of the present invention; Figure 3 It is a schematic diagram of the architecture of the Transformer Block of the face parsing encoder of the present invention; Figure 4 A schematic diagram of the architecture of a face depth estimation encoder-decoder of the present invention; Figure 5 Schematic diagram of the architecture of the self-attention mechanism of the present invention. DETAILED DESCRIPTION
[0019] The present invention is further described in detail below in conjunction with test examples and specific implementation methods. However, this should not be understood as the scope of the above subject matter of the present invention being limited to the following embodiments, and all technologies realized based on the content of the present invention belong to the scope of the present invention.
[0020] Unless otherwise specified, in the description of the specific embodiments of the present invention, the terms indicating the orientation or position relationship such as "up", "down", "left", "right", "center", "inside", "outside", etc. are all expressions based on the orientation or position relationship shown in the drawings, or the orientation or position relationship when the invented product / equipment / device is usually used. These terms of orientation or position relationship are only for the convenience of describing the scheme of the present invention or simplifying the description in the specific embodiments, so as to facilitate the technicians to quickly understand the scheme, rather than indicating or implying that a specific device / component / element must have a specific orientation, or be constructed and operated in a specific position relationship, and therefore cannot be understood as a limitation on the present invention.
[0021] In addition, if the terms "horizontal", "vertical", "overhanging", "parallel" and the like appear, it does not mean that the corresponding devices / components / elements are required to be absolutely horizontal or vertical or overhanging or parallel, but may be slightly tilted or have deviations. For example, "horizontal" only means that its direction is more horizontal than "vertical", and does not mean that the structure must be completely horizontal, but may be slightly tilted. Alternatively, it can be simplified to mean that the corresponding devices / components / elements are set in directions such as "horizontal", "vertical", "overhanging", "parallel", etc., and can have an error / deviation of ±10% relative to the corresponding direction setting, more preferably an error / deviation within ±8%, more preferably an error / deviation within ±6%, more preferably an error / deviation within ±5%, and more preferably an error / deviation within ±4%. As long as the corresponding device / component / element is within the error / deviation range, it can still achieve its role in the scheme of the present invention.
[0022] In addition, the expressions “first”, “second”, “third”, etc. that appear in the terms are merely descriptions used to distinguish the same or similar components and should not be understood as emphasizing or implying the relative importance of specific components.
[0023] In addition, in the description of the embodiments of the present invention, "several", "plurality" and "a number" represent at least 2. It can be any number such as 2, 3, 4, 5, 6, 7, 8, 9, and even more than 9.
[0024] In addition, in the description of the technical solution of the present invention, unless otherwise clearly specified / defined / restricted, the terms "set", "install", "connect", "connected", "provided with", "laid", and "arranged" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection, and can be welding, riveting, bolting, threading, and other commonly used connection means in the field. This connection can be a mechanical connection, an electrical connection, or a communication connection; it can be a direct connection, or an indirect connection through an intermediate medium, and it can be the internal connection of two elements.
[0025] Example 1 This embodiment provides a method for estimating face depth data guided by semantic information, comprising the following steps: S1, Generate pseudo-label data using unlabeled data; S2, perform data enhancement on unlabeled data; S3, using the enhanced data to train the face depth estimation model, and using the face parsing model for supervision during the training process; S4. Perform face depth estimation using the trained face depth estimation model.
[0026] The present invention generates pseudo-label data from unlabeled data and performs data enhancement on the unlabeled data, thereby providing a wider range of training samples for the face depth estimation model and solving the problem of the small number of face depth data sets. In addition, the present invention adds face semantic information, namely face parsing features, to the face depth estimation model. This feature can effectively provide the depth estimation model with accurate pixel classification information of any face image, which is conducive to the depth estimation model to estimate the depth more accurately. And this additional information can also be used as a supplement to the pseudo-label data, so that the model has auxiliary information to help it predict the depth of the face when facing the pseudo-label data that has been data enhanced.
[0027] Example 2 This embodiment is a further improvement of embodiment 1, and the repeated contents are not repeated here. According to a preferred implementation, step S1 at least includes: S11, training a large depth estimation model using a data set synthesized from a three-dimensional face model, so that the large depth estimation model has the absolute depth estimation capability for the face model. According to a preferred implementation, According to a preferred implementation, step S1 further includes: using a trained large depth estimation model to perform depth estimation on unlabeled data, generating a pseudo depth map, and forming pseudo label data.
[0028] According to a preferred implementation, step S2 at least includes: performing data enhancement on the unlabeled data to expand more face RGB images. The unlabeled data at least includes an open source real face data set. Strong data enhancement is divided into high-intensity color perturbation and spatial perturbation. High-intensity color perturbation includes color jittering, Gaussian blur, and large-scale downsampling. Spatial perturbation is to randomly splice two pictures.
[0029] According to a preferred implementation, step S3 includes: first training a face parsing model including a face parsing encoder and a decoder, and then using the trained face parsing encoder to train a face depth estimation model. The face depth estimation model includes a face depth estimation encoder and a decoder. When training the face depth estimation model, the cosine similarity of the output features of the face depth estimation encoder and the features output by the face parsing encoder is calculated. A threshold is set, and when the cosine similarity is greater than the threshold, the face depth estimation model will not update the network parameters.
[0030] According to a preferred implementation, the data processing architecture of the face parsing encoder is: using a convolution-based block embedding layer to perform convolution batch embedding on the face image, and then using 2 Transformer Blocks to perform the first feature extraction; using the first downsampling layer to process the data of the first feature extraction, and then using 2 TransformerBlocks to perform the second feature extraction; using the second downsampling layer to process the data of the second feature extraction, and then using 2 Transformer Blocks to perform the third feature extraction to obtain face parsing features.
[0031] According to a preferred implementation, the Transformer Block includes two branches, one is a multi-head self-attention branch, and the other is a convolution branch; wherein the structure of the convolution branch is an hourglass-shaped convolution layer.
[0032] According to a preferred implementation, the architecture of the face depth estimation encoder-decoder is as follows: a depth estimation encoder, which obtains a feature map through a convolution operation, and then performs downsampling twice. Two ViT Blocks are used for feature extraction before and after sampling. Each stage is connected by a maximum pooling and downsampled by 2 times; a depth estimation decoder, which obtains the output of the depth estimation encoder and performs upsampling twice. Two ViT Blocks are used for feature extraction before and after sampling, and then the depth map is output through the output convolution layer.
[0033] Example 3 This embodiment is a further improvement of Embodiment 1 and Embodiment 2, and the repeated contents are not repeated here. This embodiment provides a method for estimating face depth data guided by semantic information.
[0034] The method for estimating face depth data provided in this embodiment may include: S1. Generate pseudo-label data.
[0035] S11. Use a data set synthesized from a three-dimensional face model to train a large depth estimation model, so that the large depth estimation model has the ability to estimate the absolute depth of the face model.
[0036] Preferably, the data sets for synthesizing the three-dimensional face model are Diverse_Human_Faces_Dataset and C3IFace Depth Datasets.
[0037] Preferably, the depth estimation large model can be an existing depth estimation large model, for example, the depth estimation large model involved in Yang, L., Kang, B., Huang, Z., Xu, X., Feng, J., & Zhao, H. (2024). Depth anything: Unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp.10371-10381).
[0038] Preferably, the specific method of training the large depth estimation model may be: The depth estimation model is fine-tuned using Diverse_Human_Faces_Dataset and C3I Face Depth Datasets, two datasets synthesized from 3D face models. Specifically, the depth estimation model is trained for 10 rounds using Diverse_Human_Faces_Dataset and C3I Face Depth Datasets, two datasets synthesized from 3D face models. The loss function used in the training is as follows: in, They are the predicted result and the real label respectively, H and W represent the length and width of the image.
[0039] S12. Use the fine-tuned depth estimation large model to perform depth estimation on the unlabeled data, generate a pseudo depth map, and form pseudo labeled data. Preferably, the unlabeled data can be an open source real face dataset, or a collected real face dataset. Preferably, the open source real face dataset can be a face RGB map of LaPa and CelebA.
[0040] Preferably, the fine-tuned depth estimation model is used to perform depth estimation on an open-source real face dataset to generate a pseudo depth map for subsequent training. The generated pseudo depth map is pseudo-label data.
[0041] Preferably, after the trained large depth estimation model performs depth estimation on an open source real face dataset, several pseudo depth maps generated constitute a pseudo label dataset.
[0042] S2. Perform data enhancement The face depth estimation model is trained using the generated pseudo-label dataset and the labeled real-label dataset. The labeled real-label datasets are Diverse_Human_Faces_Dataset and C3I FaceDepth Datasets.
[0043] In order to fully utilize the pseudo-label data, this embodiment uses strong data enhancement to enable the pseudo-label data to provide a wider range of training samples for the face depth estimation model. Preferably, this embodiment performs data enhancement on the unlabeled data to expand more face RGB images.
[0044] Strong data enhancement is divided into high-intensity color perturbation and spatial perturbation.
[0045] Preferably, the high intensity color perturbation includes color jittering, Gaussian blur, and large downsampling.
[0046] The color jittering algorithm includes random brightness, random contrast, and random saturation.
[0047] Random brightness is as shown in the formula: Among them, x, y are the pixel coordinates of the image, I(x, y) is the pixel intensity of the original image, and I ’ (x,y) is the pixel intensity after random brightness adjustment, α B A random brightness coefficient, floating between [0.7, 1.3].
[0048] Random contrast is expressed as follows: in, is the average brightness of the image, α C is the contrast coefficient. The random saturation is as follows: Among them, I gray (x, y) is a grayscale image, α s is the saturation coefficient.
[0049] Gaussian blur is as follows: Among them, G(u,v) is the Gaussian kernel, and u and v represent the spatial coordinates in the image. is the variance, and k is set to 7.
[0050] The formula for large downsampling is as follows: I ’ (x ’ ,y ’ ) = I(r·x ’ ,r·y ’ ) Among them, x ’ ,y ’ is the pixel coordinate after significant downsampling, I ’ (x ’ ,y ’ ) is the pixel intensity after significant downsampling, and r is set to float between [0.5, 1.5].
[0051] Preferably, the spatial perturbation is to randomly splice two pictures. Preferably, the specific spatial perturbation method is shown in the following formula: Where M is a binary mask containing rectangular regions with values of 1. a ,u b are two different pictures with pseudo-label data. Indicates that from u a The selected image area, Indicates u b The image area selected in . ab This is the synthetic image obtained after spatial perturbation.
[0052] Preferably, the enhanced data will enter the face depth estimation model and obtain the corresponding pseudo-label prediction results.
[0053] Calculate the loss between the pseudo-label prediction result and the pseudo-label data. The loss function is calculated as follows: Among them, S() represents the image after strong data enhancement, M represents the binary mask used to select the image area, and T() represents the original image. Represents the calculation from u a The loss function of the image part is taken, Represents the calculation from u b The loss function of the selected image part. The image sample after data enhancement is essentially composed of two pseudo-label data. Therefore, the above loss calculation formula is used to calculate the loss of the pseudo-label prediction result of the combination of two pseudo-label data.
[0054] S3. Use the enhanced data to train the face depth estimation model, and use the face parsing model for supervision during the training process.
[0055] Since face parsing will classify different facial parts, the pixels of facial parts of the same category are likely to be in a similar depth range. Therefore, using face parsing to perform a certain degree of supervision on the depth can speed up the training speed and the accuracy after training.
[0056] Preferably, this embodiment first trains a face parsing model including a face parsing encoder and a decoder, and then uses the trained face parsing encoder to train a face depth estimation model. Preferably, when training the face parsing model, the CelebAMask-HQ and Lapa datasets are used for training, and a 224×224 image is input to the face parsing encoder, and then passed through the face parsing decoder, and then each pixel is classified through the MLP layer to obtain a block image after face parsing. The training loss adopts Focal loss, as follows: in, , P t The probability that the model predicts that the sample is a positive class.
[0057] Preferably, the face parsing model is trained for 50 epochs at a learning rate of 0.001 to complete the training.
[0058] When training the face depth estimation model, the output feature f of the face depth estimation encoder is i And the features output by the face parsing encoder Calculate the cosine similarity, the formula is as follows: At the same time, in order to avoid excessive supervision of the face parsing encoder, this embodiment uses a threshold to control supervision. When the cosine similarity is greater than the threshold, this loss will not be used for network parameter update.
[0059] Preferably, the threshold is set to 0.75. During training, no loss will be calculated for samples with a cosine similarity greater than 0.75, and the network parameters will not be updated based on this sample. On the contrary, for samples with a cosine similarity less than 0.75, loss will be calculated, and the sample will be used as the basis for updating the network parameters.
[0060] Preferably, different from directly using the results of face parsing for supervision, this embodiment constrains the output features of the face depth estimation model using the features output by the face parsing encoder in a continuous space (high-dimensional feature space). This weaker supervision can also provide certain semantic constraints without seriously affecting the depth estimation.
[0061] Preferably, all encoders and decoders are composed of Vision Transformer, and all decoders and encoders are constructed from 6 Transformer Encoders.
[0062] See also Figure 2 ,The data processing architecture of the face parsing encoder is as follows: The convolution-based block embedding layer first performs convolution batch embedding on face images of size H × W. C is the number of channels.
[0063] Face pictures , f() is a convolution operation, where the convolution kernel is s*s and the step size is , the surrounding pixels are filled with After the convolution operation, the feature map is obtained The dimensions are: The convolution-based block embedding layer consists of two convolutions. The first convolution has a kernel size of 3, a stride of 1, and a padding of 1; the second has a kernel size of 4, a stride of 4, and no padding.
[0064] After the image is processed by the block embedding layer, it becomes a feature map of size 56×56, and then two Transformer Blocks are used for the first feature extraction.
[0065] The first downsampling layer is used to process the data of the first feature extraction to obtain a feature map of size 28×28; then two Transformer Blocks are used for the second feature extraction.
[0066] The second downsampling layer is used to process the data of the second feature extraction to obtain a feature map of size 14×14; then two Transformer Blocks are used to perform the third feature extraction to obtain the face parsing features.
[0067] See also Figure 3 , Transformer Block contains two branches, one is the multi-head self-attention branch, and the other is the convolution branch.
[0068] The calculation of the multi-head self-attention branch in the Transformer Block is as follows: where d k is the feature map depth of Q, K, V; are the corresponding learnable parameters. model Represents the feature depth of the input transformer block. Softmax() is the softmax function. For the i-th feature z of the input softmax function i , the formula is as follows: In this embodiment, this d k The sizes are set to 96, 192, and 384 respectively when extracting feature maps of three different sizes. For each size of feature map, self-attention is only calculated within a non-overlapping 7*7 window.
[0069] The structure of the convolution branch is a combination of hourglass convolution layers (bottleneck), which consists of two reshape layers, three convolutions and a layer normalization layer. The number of convolution kernels in the three layers is 1, 3, and 1 respectively. After layer normalization, the convolution is merged with the output of the window multi-head self-attention. The overall calculation of the fusion process of the two branches is as follows: Among them, x represents the features of the input Transformer Block, Refers to the dimension transformation operation of the first reshape layer, which transforms the features from spatial vectors (4 dimensions) to sequence vectors (3 dimensions); Represents the size transformation operation of the second reshape layer, and the size transformation operation of the second reshape layer is the reverse operation of the first reshape layer; bottleneck represents the output of the hourglass convolution layer after processing the features, and LN represents the output of the features after layer normalization. GELU is a Gaussian error linear unit, which is as follows: Φ q ,φ k ,φ v It refers to a projection layer consisting of three convolutional layers with a kernel of 1, and assigns the output to Q, K, and V respectively; in the formula, x represents the feature of the input GELU.
[0070] The final output x of each Transformer Block output as follows: x output =LN d (x conv+attention )+MLP(LN(x conv+attention )) The MLP is composed of two fully connected layers. When the feature map depth is 96, the settings of (input depth, output depth) of the two fully connected layers are (96, 48) and (48, 96), respectively; when the feature map depth is 192, the settings are (192, 96) (96, 192), respectively; when the feature map depth is 384, the settings are (384, 192) (192, 384), respectively.
[0071] The architecture of the encoder-decoder for face depth estimation is as follows Figure 4 As shown. Face picture , f d () is the convolution operation, where the convolution kernel is s d *s d , the step length is , the surrounding pixels are filled with After the convolution operation, the feature map is obtained , Represents the height, width, number of channels of the feature map, and the feature map size is: For the convolution-based block embedding layer, it consists of two convolutions. The first convolution kernel size is 3, the stride is 1, and the padding is 1; the second kernel size is 4, the stride is 4, and no padding is performed. The image passes through the block embedding layer to obtain a feature map of scale 56 and enters the next module.
[0072] The depth estimation encoder uses two ViT Blocks for feature extraction in three stages when the feature map size is 56×56, 28×28, and 14×14. Each stage is connected by maximum pooling and downsampled by 2 times.
[0073] Refers to the projection layer composed of 3 convolutional layers with a kernel of 1, and assigns the output to Q respectively d , K d , V d ; In the formula, x d It is the feature of input Vit Block.
[0074] See also Figure 5 , the calculation of the multi-head self-attention branch of ViT Block is as follows: Among them, d k Q d , K d , V d Depth are the corresponding learnable parameters. model Represents the feature depth of the input Vit Block, softmax() is the softmax function, for the i-th feature z of the input softmax function i , the formula is as follows: In the present invention, this d k The size is set to 96, 192, and 384 respectively when extracting three different size feature maps. For each size feature map, self-attention is only calculated within a non-overlapping 7*7 window. The output x of the entire ViT Block is output as follows x output =LN d (x attention )+MLP d (LN d (x attention )) Among them, x attention is the output after multi-head self-attention, LN d For layer regularization in depth estimation networks, MLP d It is composed of two fully connected layers. k When d is 96, the (input depth, output depth) of the two fully connected layers are set to (96, 48) and (48, 96) respectively. kWhen d is 192, the settings are (192, 96) (96, 192) respectively; k When it is 384, the settings are (384, 192) (192, 384) respectively.
[0075] The depth estimation decoder uses two Vit Blocks for feature extraction in three stages with feature sizes of 14×14, 28×28, and 56×56. Each stage is connected by nearest neighbor interpolation. The formula is as follows: I ’ (x ’ ,y ’ ) = I(r·x ’ ,r·y ’ ) where r is set to 2. ’ ,y ’ is the pixel coordinate after the nearest interpolation. ’ ,r·y ’ ) is the eigenvalue of the original image corresponding to the coordinate, I ’ (x ’ ,y ’ ) are the eigenvalues of the new coordinates.
[0076] After the feature extraction of feature size 56×56 is completed, it first passes through a convolution layer with a kernel of 3, an input depth of 96, and an output depth of 48, and then uses pixel shuffle to restore the depth image to a size of 224*224. The pixel shuffle formula is as follows: I'(h·r+i,w·r+j,c)=I(h,w,c·r 2 +i·r+j) Where: h and w are the spatial locations of the input tensor. c is the channel index of the output tensor, set to 48. I, j are the offsets in the r×r grid, with r being 4. ’ is the eigenvalue of the new feature map, and I is the eigenvalue of the original size feature map.
[0077] This embodiment adds face semantic information, namely face parsing features. This feature can effectively provide the depth estimation model with accurate pixel classification information of any face image, which is conducive to the depth estimation model to estimate the depth more accurately. And this additional information can also be used as a supplement to pseudo-label data, so that the model has auxiliary information to help it predict the depth of the face when facing pseudo-label data that has been enhanced by data.
[0078] In view of the shortcomings of small face depth data sets, low depth estimation accuracy and complex models, this embodiment proposes a method for generating a large amount of pseudo-label face depth data and estimating face depth based on face analysis priors.
[0079] S4. Perform face depth estimation using the trained face depth estimation model.
[0080] Preferably, after completing the training of the face depth estimation model in step S3, the picture for which face depth estimation is required is input into the face depth estimation model to obtain the corresponding depth map.
[0081] The method for estimating the depth of a human face provided by this embodiment has obvious advantages over the existing methods for estimating the depth of a human face: J.Cui, H.Zhang, H.Han, S.Shan, and X.Chen, "Improving 2D face recognition via discriminative face depth estimation" in Proc.Int.Conf.Biometrics (ICB), Feb.2018, pp.140–147. A facial recognition system was designed in which a fully convolutional network (FCN) attempted to recover depth from RGB images, while a convolutional neural network (CNN) maintained identity information. However, the depth estimated by this method is relatively coarse, and the predicted face depth map contains a large amount of noise caused by depth loss and estimation failure. The face depth estimation method provided in this embodiment uses high-precision face depth training samples obtained by a high-precision three-dimensional face model, and uses these data to fine-tune the depth prediction model, so that pseudo-label data with relatively high accuracy can be generated, so that there is no noise caused by depth loss and estimation failure.
[0082] KAT Arslan and E. Seke, "Face depth estimation with conditional generative adversarial networks," IEEE Access, vol. 7, pp. 23222–23231, 2019. A facial depth estimator based on conditional generative adversarial networks (GAN) is proposed, which is used to estimate the depth map from a single face image through a GAN-based method. The method also concludes that the conditional Wasserstein GAN structure is the most reliable technology using the GAN network. However, due to the cross-optimization of the generator and the discriminator, the training of GAN is very unstable, and due to the small amount of training data, the generalization ability of this algorithm is not strong. It mainly performs depth estimation on frontal high-definition faces, but cannot cope with the depth estimation task of faces with larger postures and more complex lighting. The face depth estimation method provided in this embodiment can generate a large amount of pseudo-label data for each scene and each posture, thereby improving the coverage of training samples, so that the depth estimation method has better generalization performance.
[0083] The authors of JRAMoniz, C.Beckham, S.Rajotte, S.Honari, and C.Pal, ''Unsupervised depth estimation, 3D face rotation and replacement,'' in Proc.Adv.NeuralInf.Process.Syst., vol.31, 2018, pp.1–14 used an unsupervised method to estimate the depth of 3D facial rotation and replacement by implicitly implying the depth of facial key points of the input image. However, this algorithm only focuses on the depth of 68 facial feature points, rather than the depth estimation of the entire face. It provides limited help for other tasks such as face reconstruction and editing after depth estimation. The face depth estimation method provided in this embodiment predicts the depth value corresponding to each face pixel.
[0084] AT Baby, A. Andrews, A. Dinesh, A. Joseph, and VK Anjusree, "Face depth estimation and 3D reconstruction," in Proc. Adv. Comput. Commun. Technol. High Perform. Appl. (ACCTHPA), Jul. 2020, pp. 125–132. A GAN-based technical solution is proposed to generate robust facial depth estimation. However, this technology focuses on high-precision frontal face images, and there is not enough experimental verification for general face images and faces with large postures. The face depth estimation method provided in this embodiment: first, the labeled data set used is a multi-pose projection of a three-dimensional face model, which itself has samples of various postures; secondly, the generated pseudo-label data also contains face depth data of various environments and various postures.
[0085] F. Zhang, N. Liu, Y. Hu, and F. Duan, "MFFNet: Single facial depth map refinement using multi-level feature fusion," Signal Process., Image Commun., vol. 103, Apr. 2022, Art. no. 116649. A GAN technology for facial depth estimation using a segmentation and mask-guided attention network is proposed. However, the application scenario of this technical solution is high-definition face images with a relatively simple background and a single lighting environment, and it cannot be extended to general face images with more complex backgrounds and more complex scenes. The large depth estimation model used in the face depth estimation method provided in this embodiment has extremely strong generalization. We fine-tune the face depth estimation dataset based on the large depth estimation model, which can inherit the generalization of the model to cope with face depth estimation in complex backgrounds. At the same time, the labeled face depth dataset used also contains many complex simulated backgrounds. Rich data can enable the face depth estimation model of this embodiment to cope with complex background environments.
[0086] Example 4 This embodiment provides a semantic information-guided face depth data estimation model. The face depth data estimation model provided by this embodiment can implement the semantic information-guided face depth data estimation method involved in Embodiment 1, Embodiment 2, and Embodiment 3.
[0087] Example 5 This embodiment provides a semantic information-guided face depth data estimation system. The face depth data estimation system provided by this embodiment can implement the semantic information-guided face depth data estimation method involved in Embodiment 1, Embodiment 2, and Embodiment 3.
[0088] The face depth data estimation system includes an input unit, an output unit and a processing unit.
[0089] The input unit is used to input a face picture and transmit the face picture to the processing unit. The processing unit runs a preset program to execute the face depth data estimation method guided by semantic information involved in Embodiment 1, Embodiment 2, and Embodiment 3 to obtain a depth estimation map corresponding to the face picture, and transmits the depth estimation map to the output unit. The output unit is used to output the depth estimation result of the face depth data estimation system on the face picture.
[0090] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A semantic information-guided face depth data estimation method, characterized in that: The following steps are involved: S1, Generate pseudo-label data using unlabeled data; S2, perform data enhancement on unlabeled data; S3, using the enhanced data to train the face depth estimation model, and using the face parsing model for supervision during the training process; S4. Perform face depth estimation using the trained face depth estimation model.
2. The method for estimating face depth data guided by semantic information according to claim 1, characterized in that: Step S1 at least includes: S11. Use a data set synthesized from a three-dimensional face model to train a large depth estimation model, so that the large depth estimation model has the ability to estimate the absolute depth of the face model.
3. The method for estimating face depth data guided by semantic information according to claim 2, characterized in that: Step S1 also includes: The trained depth estimation model is used to perform depth estimation on the unlabeled data to generate a pseudo depth map and form pseudo labeled data.
4. The method for estimating face depth data guided by semantic information according to claim 1, characterized in that: Step S2 at least includes: Performing data enhancement on the unlabeled data to expand more face RGB images; wherein the unlabeled data at least includes an open source real face data set; The strong data enhancement is divided into high-intensity color perturbation and spatial perturbation; The high intensity color disturbance includes color jittering, Gaussian blur, and large-scale downsampling; The spatial perturbation is to randomly splice two pictures.
5. The method for estimating face depth data guided by semantic information according to claim 1, characterized in that: Step S3 includes: First, a face parsing model including a face parsing encoder and a decoder is trained, and then the trained face parsing encoder is used to train a face depth estimation model; the face depth estimation model includes a face depth estimation encoder and a decoder; When training the face depth estimation model, the cosine similarity of the output features of the face depth estimation encoder and the features output by the face parsing encoder is calculated; Set a threshold. When the cosine similarity is greater than the threshold, the face depth estimation model will not update the network parameters.
6. The method for estimating face depth data guided by semantic information according to claim 5, characterized in that: The data processing architecture of the face parsing encoder is: The convolution-based block embedding layer is used to perform convolution batch embedding on the face image, and then two TransformerBlocks are used for the first feature extraction; The first downsampling layer is used to process the data of the first feature extraction, and then two TransformerBlocks are used for the second feature extraction; The second downsampling layer is used to process the data of the second feature extraction, and then two TransformerBlocks are used to perform the third feature extraction to obtain the face parsing features.
7. The method for estimating face depth data guided by semantic information according to claim 6, characterized in that: The Transformer Block includes two branches, one is a multi-head self-attention branch, and the other is a convolution branch; wherein the structure of the convolution branch is an hourglass-shaped convolution layer.
8. The method for estimating face depth data guided by semantic information according to claim 5, characterized in that: The architecture of the face depth estimation encoder-decoder is: The depth estimation encoder obtains the feature map through convolution operation, and then downsamples it twice. Two ViT Blocks are used for feature extraction before and after sampling. Each stage is connected by maximum pooling and downsampled by 2 times. The depth estimation decoder obtains the output of the depth estimation encoder and performs two upsamplings. Two ViT Blocks are used for feature extraction before and after sampling, and then the depth map is output through the output convolution layer.
9. A semantic information-guided face depth data estimation model, characterized in that: Used to implement the steps involved in the semantic information guided facial depth data estimation method as described in claims 1 to 8.
10. A semantic information guided face depth data estimation system, characterized in that: Used to implement a semantic information guided facial depth data estimation method as described in claims 1 to 8.
Citation Information
Patent Citations
Cross-domain facial expression recognition method, system and device and storage medium
CN117152806A