A method, system and readable storage medium for training a three-dimensional face reconstruction model

By embedding the channel-space attention mechanism in the Encoder-Decoder network, the accuracy problem of three-dimensional face reconstruction and dense point alignment under occlusion is solved, and high-precision face reconstruction and alignment in occlusion is achieved.

CN115115784BActive Publication Date: 2025-07-11SHANGHAI JIELU INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210852098.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-20
Publication Date
2025-07-11
Estimated Expiration
2042-07-20

AI Technical Summary

Technical Problem

The existing three-dimensional face reconstruction and dense point alignment methods are insufficient in the presence of occlusion, especially when facing self-occlusion and external occlusion, it is impossible to accurately complete the tasks of dense face alignment and three-dimensional face reconstruction.

Method used

The Encoder-Decoder network structure is adopted, combined with the channel-space attention perception mechanism, and the initial three-dimensional face reconstruction model is constructed by obtaining the face data set and feature point annotation information, and the face area mask and target loss function are used for model training to improve the sensitivity to the face visibility area.

Benefits of technology

In the presence of occlusion, the accuracy of dense face alignment and three-dimensional face reconstruction is improved, the complexity of network learning is reduced, and the accuracy of reconstruction results is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115784B_ABST
    Figure CN115115784B_ABST
Patent Text Reader

Abstract

A method, system and readable storage medium for training a three-dimensional face reconstruction model provided by an embodiment of the present application. The method includes obtaining a face dataset containing multiple face images and the feature point annotation information of each face image; processing each face image according to the feature point annotation information to obtain training supervision data, where the training supervision data includes a face region mask, a projected standard average face model, and standard face deformation information; constructing an initial three-dimensional face reconstruction model for combining the predicted average face model and the face deformation information for three-dimensional face reconstruction. The initial three-dimensional face reconstruction model is composed of multiple Encoder-Decoder networks with the same structure, and a channel-spatial attention perception mechanism is added to the skip connections of the encoding layer and the decoding layer of the Encoder-Decoder network; training the model based on the face image and the corresponding face region mask, and obtaining the target three-dimensional face reconstruction model when the training end condition is reached.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology. Specifically, it relates to a method, system, and readable storage medium for training a three-dimensional face reconstruction model. Background Art

[0002] The two tasks of reconstructing a three-dimensional face from a single input face image and achieving dense point alignment of the three-dimensional face are actually two closely related tasks, and they both play a huge role in many fields such as face recognition, face tracking, face animation, and human-computer interaction.

[0003] Existing three-dimensional face reconstruction / alignment methods include model-based methods with the 3D Morphable Model (3DMM) as the foundation. Among them, this method is based on the idea of reconstructing a three-dimensional face from a single image using 3DMM, and optimizes the computational loss between its three-dimensional face template and the input face image through a synthesis analysis method to adjust the corresponding face shape parameters and face texture parameters. However, in daily information interaction, the obtained face images often have varying degrees of occlusion. For example, self-occlusion may be caused by the angle offset of one's own head, or external occlusion may be caused by certain objects such as hair or glasses. These occlusion problems will lead to the loss of some face information, resulting in insufficient accuracy in dense face alignment and three-dimensional face reconstruction in the presence of occlusion. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a method, system, and readable storage medium for training a three-dimensional face reconstruction model, which can improve the accuracy of dense face alignment and three-dimensional face reconstruction in the presence of occlusion.

[0005] The embodiments of this application also provide a method for training a three-dimensional face reconstruction model, including the following steps:

[0006] Obtain a face dataset containing multiple face images and the feature point annotation information of each of the face images;

[0007] According to the feature point annotation information, process each face image in the face dataset to obtain training supervision data, where the training supervision data includes a face region mask, a standard average face model after projective transformation, and standard face deformation information;

[0008] Construct an initial 3D face reconstruction model for combining the predicted average face model and face deformation information for 3D face reconstruction. The initial 3D face reconstruction model is composed of multiple Encoder-Decoder networks with the same structure, and a channel-spatial attention perception mechanism is added to the skip connections between the encoding layer and the decoding layer of the Encoder-Decoder network;

[0009] Based on the face image and the corresponding face region mask, perform model training. During the training process, constrain by combining a target loss function used to reflect the deviation degree between the prediction result and the corresponding standard result, and obtain the target 3D face reconstruction model when the training end condition is reached.

[0010] In a second aspect, an embodiment of the present application further provides a 3D face reconstruction model training system, which includes a data acquisition module, a data processing module, a model construction module, and a model training module, where:

[0011] The data acquisition module is used to acquire a face dataset containing multiple face images and the feature point annotation information of each face image;

[0012] The data processing module is used to process each face image in the face dataset according to the feature point annotation information to obtain training supervision data, where the training supervision data includes a face region mask, a projected standard average face model, and standard face deformation information;

[0013] The model construction module is used to construct an initial 3D face reconstruction model for combining the predicted average face model and face deformation information for 3D face reconstruction. The initial 3D face reconstruction model is composed of multiple Encoder-Decoder networks with the same structure, and a channel-spatial attention perception mechanism is added to the skip connections between the encoding layer and the decoding layer of the Encoder-Decoder network;

[0014] The model training module is used to perform model training based on the face image and the corresponding face region mask. During the training process, constrain by combining a target loss function used to reflect the deviation degree between the prediction result and the corresponding standard result, and obtain the target 3D face reconstruction model when the training end condition is reached.

[0015] In a third aspect, an embodiment of the present application further provides a readable storage medium, which includes a 3D face reconstruction model training method program. When the 3D face reconstruction model training method program is executed by a processor, the steps of a 3D face reconstruction model training method as described in any one of the above are implemented.

[0016] As can be seen from the above, a three-dimensional face reconstruction model training method, system, and readable storage medium provided by the embodiments of the present application appropriately embed the attention mechanism into the Encoder-Decoder structure network, enabling the network to always have higher sensitivity to the visible area of the face during the encoding and decoding processes. It can accurately complete dense face alignment and three-dimensional face reconstruction tasks even when there are occlusions in the input face image. In addition, the three-dimensional face to be predicted is unwrapped into an average face model after projective transformation and face deformation information after projective transformation, which can reduce the complexity of network learning, improve the accuracy of network derivation results, and improve the accuracy of dense face alignment and three-dimensional face reconstruction in the presence of occlusions.

[0017] Other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification, or can be understood by implementing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a flowchart of a three-dimensional face reconstruction model training method provided by the embodiments of the present application;

[0020] Figure 2 It is a schematic framework diagram of the Encoder-Decoder network provided by the embodiments of the present application;

[0021] Figure 3 It is a schematic overall framework diagram of a three-dimensional face reconstruction model training method provided by the embodiments of the present application;

[0022] Figure 4 It is a schematic structural diagram of a three-dimensional face reconstruction model training system provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0024] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0025] Please refer to Figure 1 , Figure 1 which is a flowchart of a three-dimensional face reconstruction model training method in some embodiments of the present application. Taking the application of this method to a computer device (the computer device can specifically be a terminal or a server, and the terminal can specifically but not limited to be various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be an independent server or a server cluster composed of multiple servers) as an example, it includes the following steps:

[0026] Step S100, obtain a face dataset including multiple face images and the feature point annotation information of each of the face images.

[0027] Step S200, process each face image in the face dataset according to the feature point annotation information to obtain training supervision data, where the training supervision data includes a face region mask, a standard average face model after projective transformation, and standard face deformation information.

[0028] Step S300, construct an initial three-dimensional face reconstruction model for combining the predicted average face model and the face deformation information for three-dimensional face reconstruction. The initial three-dimensional face reconstruction model is composed of multiple Encoder-Decoder networks with the same structure, and a channel-spatial attention perception mechanism is added to the skip connections between the encoding layer and the decoding layer of the Encoder-Decoder network.

[0029] Step S400: Based on the face image and the corresponding face region mask, model training is performed. During the training process, it is constrained by a target loss function used to reflect the deviation degree between the prediction result and the corresponding standard result. When the training end condition is reached, a target 3D face reconstruction model is obtained.

[0030] As can be seen from the above, a 3D face reconstruction model training method disclosed in this application appropriately embeds the attention mechanism into the Encoder-Decoder structure network, enabling the network to always have higher sensitivity to the visible region of the face during the encoding and decoding processes. It can accurately complete the dense face alignment and 3D face reconstruction tasks even when there are occlusions in the input face image. In addition, the 3D face to be predicted is untangled into an average face model after projective transformation and face deformation information after projective transformation, which can reduce the complexity of network learning, improve the accuracy of network derivation results, and improve the accuracy of dense face alignment and 3D face reconstruction in the presence of occlusions.

[0031] In one embodiment, in step S200, when performing face mask processing on each face image in the face dataset, the method includes:

[0032] Step S2001: According to the feature point annotation information, determine the 3D vertex set of each face image in the world coordinate system.

[0033] Specifically, taking the publicly available face dataset 300W-LP as an example, it is also necessary to perform standardization processing on each face image in the face dataset, that is, crop each face image into standard image blocks of the same size according to a preset cropping standard.

[0034] In one embodiment, since in the face dataset 300W-LP, the facial key feature points (i.e., landmark points) of each face image have been artificially annotated in advance based on the 3DMM method. Therefore, in the current embodiment, the computer device can first determine the cropping range according to the landmark points, and then, based on this cropping range, crop all the images into standard image blocks of, for example, 256x256 size.

[0035] Step S2002: According to the 3D vertex set S of each face image in the world coordinate system, calculate the 3D vertex set V of the corresponding face image in the image coordinate system through the following formula:

[0036] V = f·R·S + t; (1)

[0037] Where f represents the scaling factor, and R and t respectively represent the rotation matrix and translation vector calculated from the 3DMM pose parameters in the face dataset.

[0038] Step S2003, combining the three-dimensional vertex sets of each face image in the image coordinate system to construct a face region mask in the image coordinate system.

[0039] Specifically, when the prediction result is expressed in the form of a UV position map, in the current embodiment, the computer device will first save the three-dimensional vertex set corresponding to each face image in the image coordinate system to the UV space to obtain the UV position map corresponding to the face image. Afterwards, because the unfolding information of the face is strictly aligned in the UV space. Therefore, the computer device will design a binary (0 / 255) face mask in the UV space based on the face area information in the UV space, multiply it with the UV position map of the corresponding image, and exclude invisible points. Afterwards, after converting back to the image domain, the face area mask in the image coordinate system can be obtained.

[0040] It should be noted that the computer device will also randomly add some covering graphics to the aforementioned constructed face area masks, and use these disturbed face area masks as training supervision data for the face area attention network.

[0041] In one embodiment, in step S200, when performing projection transformation processing on each face image in the face data set, the method includes:

[0042] Step S2004, obtaining an average face model corresponding to the face image, and performing a projection transformation process on the average face model to obtain a standard average face model after the projection transformation.

[0043] Specifically, when the prediction result is represented in the form of a UV position map, in the current embodiment, the computer device converts the standard average face template after the projection transformation into a UV map to obtain a UV projection transformation map.

[0044] Step S2005, combining the standard average face model with the difference between the corresponding three-dimensional face vertex positions in the image coordinate system to determine the standard face deformation information after the projective transformation.

[0045] Specifically, based on the above embodiment, after obtaining the UV projection transformation map, the computer device will calculate the difference between the UV projection transformation map and the UV position map of the corresponding face image, and based on the obtained difference result, determine the UV deformation map that reflects the face deformation information after the projection transformation.

[0046] In summary, the implementation of steps S2001 - S2005 is to facilitate obtaining the training supervision data required for training the network, including UV position maps, UV projection transformation maps reflecting the average face model after projection transformation, UV deformation maps reflecting the face deformation information after projection transformation, and masked face regions affected by interference. Of course, in different embodiments, according to different training methods and purposes, the types of training supervision data also vary, and the embodiments of the present application do not limit this.

[0047] In one embodiment, please refer to Figure 2 , in the skip connections between the encoding layer and the decoding layer of the Encoder - Decoder network, a face region attention network applying a channel - spatial attention perception mechanism is embedded, where: the face region attention network uses the masked face region as training supervision data. After the encoded feature map output by the encoding layer passes through the face region attention network, a corresponding visibility score feature map will be obtained; after the face region attention network, there is also a Max - pooling layer connected for adjusting the spatial resolution of the visibility score feature map so that the spatial resolution of the visibility score feature map is consistent with that of the encoded feature map.

[0048] It should be noted that skip connection operations will be used in the encoding layer and the decoding layer corresponding to the Encoder - Decoder network. The reason for embedding the face region attention network in this skip connection (specifically, refer to Figure 2 ) is to enhance the sensitivity to the visible face region of the input feature map and improve the recognition accuracy.

[0049] Specifically, in the skip connection between a corresponding encoding layer and decoding layer, the encoded feature map output by the encoding layer will be fed into the face region attention network. Since the face region attention network will use a pre - constructed masked face region covering interference information as training supervision data. Therefore, in the current embodiment, a Max - pooling layer needs to be connected after this face region attention network so that the output visibility score feature map can have the same spatial resolution as the encoded feature map.

[0050] In one embodiment, the face region attention network consists of a fully - connected layer for converting the number of channels of the feature map and multiple standard convolutional blocks applying the channel - spatial attention perception mechanism, and activates the output of the visibility of each pixel point via the sigmoid function. Among them, the feature extraction operation is implemented through the following formula:

[0051]

[0052] Mc F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))); (3)

[0053]

[0054]

[0055] where F ∈ R CxHxW represents the input feature map, M c (*) represents performing channel attention processing on "*", M s (*) represents performing spatial attention processing on "*"; AvgPool(*) represents performing average pooling processing on "*", MaxPool(*) represents performing max pooling processing on "*"; MLP(*) represents shared weight processing, σ represents the sigmoid function; Conv(*) represents a standard convolution operation, F Avg 、F Max distributions represent applying an average pooling operation and a max pooling operation along the channel axis, respectively, to obtain the corresponding 2D maps.

[0056] It should be noted that when performing channel attention processing, it includes: processing the input feature map through an average pooling operation and a max pooling operation respectively, and then sharing the weights of the respective obtained channel descriptors through an MLP layer. Finally, after weighted summing the above two channel descriptors, it is activated with the sigmoid function. Among them, the entire implementation process can refer to the above formula (3).

[0057] On the other hand, the spatial attention processing process also includes these two operations of average pooling operation and max pooling. However, these two operations are applied along the channel axis to further obtain the corresponding 2D map, namely F Avg ∈ R 1×H×W and F max ∈ R 1×H×W . Then, F Avg and F Max are concatenated and passed through a fully connected layer to generate a two-dimensional spatial attention feature map. Among them, the entire implementation process can refer to the above formula (5).

[0058] In one embodiment, the initial 3D face reconstruction model is composed of two Encoder-Decoder networks with the same structure; the encoding network part of the Encoder-Decoder network is composed of multiple residual block modules with a channel attention mechanism added, and the main feature extraction channels of the residual block modules are composed of a 1x1 standard convolution layer, a 3x3 standard convolution layer, a 1x1 standard convolution layer, and a channel attention extraction operation layer.

[0059] It should be noted that the encoding network part of the Encoder-Decoder network consists of 6 residual modules (i.e., SE_Resblock) with channel attention mechanism added.

[0060] In one embodiment, for a face image with an input size of 256×256×3, it will first pass through a fully connected layer to change the number of channels to 16. Subsequently, it is input into the encoding network part of the Encoder network. Among them, each layer of the encoding network consists of a SE_Resblock with a stride of 2 and a SE_Resblock with a stride of 1. And the main feature extraction channels of the SE_Resblock are composed of the following parts:

[0061] 1×1 standard convolution layer, 3×3 standard convolution layer, 1×1 standard convolution layer, and channel attention extraction operation (SE).

[0062] Among them, the image input to the encoding layer will first pass through the first two standard convolution layers, and then, it will be subjected to BN operation through the last 1×1 standard convolution layer. Finally, it is the SE operation. Specifically, the implementation process of the SE operation includes:

[0063] (1) Perform global average pooling operation on an input image with a size of H×W×C to obtain a feature map of 1×1×C.

[0064] (2) Send the 1×1×C feature map into a fully connected layer with the number of neurons being C / 16. After activation by Relu, it is sent into a fully connected layer with the number of neurons being C.

[0065] (3) Finally, after sigmoid activation, obtain the corresponding 1×1×C feature channel descriptor. Multiply the 1×1×C feature channel descriptor by the original input feature map to obtain the corresponding encoded feature map.

[0066] In one embodiment, in step S400, when performing 3D face reconstruction through the Encoder-Decoder network, the method includes:

[0067] Step S4001: Process the input face image by the encoding layer in the skip connection to obtain an encoded feature map.

[0068] Step S4002: Take the encoded feature map as the input of the face region attention network, and process it by the face region attention network to obtain a visibility score feature map.

[0069] Step S4003: Correlate the obtained visibility score feature map and the encoded feature map through the following formula to obtain the corresponding visible region feature map F att :

[0070] F att = F ⊙ (1 + A); (6)

[0071] where A represents the obtained visibility score feature map, and F represents the encoded feature map.

[0072] Step S4004: Concatenate the visible region feature map with the feature map output after the last transposed convolutional block in the decoding layer in the skip connection to obtain the required output feature map.

[0073] Specifically, in the current embodiment, two parallel Encoder-Decoder networks with the same structure are used to predict the average face model after projection transformation and the face deformation information after projection transformation respectively. The final required 3D face reconstruction result can be determined by adding the prediction results of the two parallel Encoder-Decoder networks.

[0074] In one embodiment, the obtained 3D face reconstruction result is represented in the form of a UV position map, and specific reference can be made to Figure 3 .

[0075] In one embodiment, when representing the prediction result in the form of a UV position map, the calculation formula of the target loss function includes:

[0076]

[0077] In the above formula, h and w represent the height and width of the UV position map; N(u, v) represents the prediction result predicted based on the position coordinate point (u, v) in the UV space; represents the training supervision data corresponding to the prediction result, that is, the standard result, and M(u, v) represents the weight value attached to the position coordinate point (u, v) in the UV space; among them, to ensure the accuracy of the prediction result, the following constraint term L for landmark points - key face feature points is imposed on the average face model after projection transformation lrr :

[0078]

[0079] where P(u, v) represents the predicted 3D coordinate information of the landmark points on the average face predicted based on the position coordinate point (u, v) in the UV space, Indicates the standard three-dimensional coordinate information of the landmark points on the average face model corresponding to the position coordinate points (u, v) in the UV space in the training supervision data.

[0080] Specifically, after the entire network is established in the steps described above (for the overall network structure diagram, reference can be made to Figure 3 ), add the results derived from the two parallel Encoder-Decoder networks, namely the projected and transformed average face model and the projected and transformed face deformation information, to obtain the final three-dimensional face reconstruction result.

[0081] In one embodiment, if the result obtained from this operation is a result saved in the form of a UV position texture map, in the current embodiment, the computer device will use the loss function L rec to constrain the final reconstructed face result, use the loss function L r_mean to constrain the predicted projected and transformed average face model, and use the loss function L r_d to constrain the predicted projected and transformed face deformation information.

[0082] It should be noted that the weight mask M is a mask designed based on different regions of the face with a size ratio of 16:12:3:0. In one embodiment, the corresponding M value of the key feature points on the face can be set to 16, the corresponding M value of the points in the eye, nose, and mouth regions can be set to 12, the corresponding M value of the points in the cheek, chin, and forehead regions can be set to 3, and the corresponding M value of the points in the neck region can be set to 0. The embodiments of the present application are not limited thereto.

[0083] In one embodiment, in order to further improve the performance of the three-dimensional face reconstruction result in the face alignment task. In the current embodiment, a constraint term for the key feature points (i.e., landmark points) of the face is further imposed on the projected and transformed average face model, which is represented by L lrr and is expressed as such.

[0084] In summary, in the case of knowing the loss functions L rec , L r_mean , L r_d , L att , and L lrr , the computer device can determine the required target loss function based on the weighted sum result of the above functions:

[0085] L total = ω rec L rec + ω r_mean L r_mean + ω r_d L r_d + ωatt L att + ω lrr L lrr ; (9)

[0086] Among them, ω rec is the weight value of L rec in the entire objective function. In the current embodiment, its value can be 1. The embodiments of the present application do not limit its specific value; ω r_mean is the weight value of L r_mean in the entire objective function. In the current embodiment, its value can be 0.5. The embodiments of the present application do not limit its specific value; ω r_d is the weight value of L r_d in the entire objective function. In the current embodiment, its value can be 0.5. The embodiments of the present application do not limit its specific value; ω att is the weight value of L att in the entire objective function. In the current embodiment, its value can be 0.1. The embodiments of the present application do not limit its specific value; ω lrr is the weight value of L lrr in the entire objective function. In the current embodiment, its value can be 0.1. The embodiments of the present application do not limit its specific value.

[0087] Please refer to Figure 4 , a three-dimensional face reconstruction model training system 400 disclosed in the present application. The system 400 includes a data acquisition module 401, a data processing module 402, a model construction module 403, and a model training module 404, where:

[0088] The data acquisition module 401 is configured to acquire a face data set including multiple face images and the feature point annotation information of each of the face images.

[0089] The data processing module 402 is configured to process each face image in the face data set according to the feature point annotation information to obtain training supervision data, where the training supervision data includes a face region mask, a standard average face model after projective transformation, and standard face deformation information.

[0090] The model construction module 403 is configured to construct an initial three-dimensional face reconstruction model for combining the predicted average face model and the face deformation information for three-dimensional face reconstruction. The initial three-dimensional face reconstruction model is composed of multiple Encoder-Decoder networks with the same structure, and a channel-spatial attention perception mechanism is added to the skip connection between the encoding layer and the decoding layer of the Encoder-Decoder network.

[0091] The model training module 404 is configured to perform model training based on the face region mask. During the training process, it is constrained by a target loss function used to reflect the deviation degree between the prediction result and the corresponding standard result, and a target three-dimensional face reconstruction model is obtained when the training end condition is reached.

[0092] In one embodiment, each module in the system can execute the method in any optional implementation manner of the above embodiment.

[0093] As can be seen from the above, in a three-dimensional face reconstruction model training system disclosed in the present application, the attention mechanism is appropriately embedded into the Encoder-Decoder structure network, enabling the network to always have higher sensitivity to the visible region of the face during the encoding and decoding processes, and accurately completing the dense face alignment and three-dimensional face reconstruction tasks even when there are occlusions in the input face image. In addition, the three-dimensional face to be predicted is untangled into an average face model after projective transformation and face deformation information after projective transformation, which can reduce the complexity of network learning, improve the accuracy of network derivation results, and improve the accuracy of dense face alignment and three-dimensional face reconstruction in the presence of occlusions.

[0094] The embodiments of the present application provide a readable storage medium. When the computer program is executed by a processor, it executes the method in any optional implementation manner of the above embodiment. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0095] The above-readable storage medium appropriately embeds the attention mechanism into the Encoder-Decoder structure network, enabling the network to always be more sensitive to the visible region of the face during the encoding and decoding processes. It can accurately complete dense face alignment and 3D face reconstruction tasks even when there are occlusions in the input face image. In addition, the predicted 3D face is decomposed into the average face model after projective transformation and the face deformation information after projective transformation, which can reduce the complexity of network learning, improve the accuracy of network derivation results, and enhance the accuracy of dense face alignment and 3D face reconstruction in the presence of occlusions.

[0096] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical functional division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0097] In addition, the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0098] Furthermore, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0099] In this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0100] The above are only the embodiments of the present application and are not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A method for training a three-dimensional face reconstruction model, characterized in that, It includes the following steps: Obtain a face dataset containing multiple face images and the feature point annotation information of each of the face images; Process each face image in the face dataset according to the feature point annotation information to obtain training supervision data, where the training supervision data includes a face region mask, a standard average face model after projective transformation, and standard face deformation information; Construct an average face model for prediction and an initial 3D face reconstruction model for 3D face reconstruction using face deformation information. The initial 3D face reconstruction model is composed of multiple Encoder-Decoder networks with the same structure, and a channel-spatial attention perception mechanism is added to the skip connections between the encoding layer and the decoding layer of the Encoder-Decoder network; Based on the face images and the corresponding face region masks, perform model training. During the training process, it is constrained by a target loss function used to reflect the deviation degree between the prediction result and the corresponding standard result, and when the training end condition is reached, a target 3D face reconstruction model is obtained; The face region attention network is composed of a fully connected layer for converting the number of channels of the feature map and multiple standard convolutional blocks applying the channel-spatial attention perception mechanism, and activates the output of the visibility of the face region of each pixel point through a sigmoid function, where the feature extraction operation is implemented by the following formula: M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))); where, F ∈ R CxHxW denotes the input feature map, M c (*) denotes performing channel attention processing on "*", M s (*) denotes performing spatial attention processing on "*"; AvgPool(*) denotes performing average pooling processing on "*", MaxPool(*) denotes performing max pooling processing on "*"; MLP(*) denotes shared weight processing, σ denotes the sigmoid function; Conv(*) denotes a standard convolution operation, F Avg 、F Max distribution represents the 2D map obtained by applying the average pooling operation and the maximum pooling operation along the channel axis respectively; When representing the prediction result in the form of a UV position map, the calculation formula of the target loss function includes: In the above formula, h and w represent the height and width of the UV position map; N(u, v) represents the prediction result predicted based on the position coordinate points (u, v) in the UV space; represents the training supervision data corresponding to the prediction result, i.e., the standard result, and M(u, v) represents the weight value attached to the position coordinate point (u, v) in the UV space; Among them, to ensure the accuracy of the prediction results, the following constraint term L for landmark points - key facial feature points is imposed on the average face model after projective transformation lrr : Among them, P(u, v) represents the predicted three-dimensional coordinate information of the landmark points on the average human face predicted based on the position coordinate point (u, v) in the UV space. It represents the standard three-dimensional coordinate information of the landmark points on the average face model corresponding to the position coordinate point (u, v) in the UV space in the training supervision data.

2. The method according to claim 1, wherein When performing face mask processing on each face image in the face dataset, the method includes: According to the feature point annotation information, determine the three-dimensional vertex set of each face image in the world coordinate system; According to the three-dimensional vertex set S of each face image in the world coordinate system, calculate the three-dimensional vertex set V of the corresponding face image in the image coordinate system through the following formula: V = f·R·S + t; where f represents a scaling factor, and R and t respectively represent a rotation matrix and a translation vector calculated from the 3DMM pose parameters in the face dataset; Combine the three-dimensional vertex sets of each face image in the image coordinate system to construct a face region mask in the image coordinate system.

3. The method according to claim 1, wherein When performing projective transformation processing on each face image in the face dataset, the method includes: Obtain the average face model corresponding to the face image and perform projective transformation processing on the average face model to obtain a standard average face model after projective transformation; Combine the difference between the standard average face model and the three-dimensional vertex position of the corresponding face in the image coordinate system to determine the standard face deformation information after projective transformation.

4. The method according to claim 1, wherein The Encoder-Decoder network embeds a face region attention network applying the channel-spatial attention perception mechanism in the skip connections between the encoding layer and the decoding layer, where: The face region attention network uses the face region mask as training supervision data. After the encoded feature map output by the encoding layer passes through the face region attention network, a corresponding visibility score feature map will be obtained; After the face region attention network, there is also a Max-pooling layer connected, which is used to adjust the spatial resolution of the visibility score feature map so that the spatial resolution of the visibility score feature map is consistent with that of the encoded feature map.

5. The method according to claim 4, wherein The initial 3D face reconstruction model is composed of two Encoder-Decoder networks with the same structure; The encoding network part of the Encoder-Decoder network is composed of multiple residual block modules with channel attention mechanisms added. The main feature extraction channels of the residual block modules are composed of a 1x1 standard convolutional layer, a 3x3 standard convolutional layer, a 1x1 standard convolutional layer, and a channel attention extraction operation layer.

6. The method according to claim 4, characterized in that When performing 3D face reconstruction through the Encoder-Decoder network, the method includes: The input face image is processed by the encoding layer in the skip connection to obtain an encoded feature map; The encoded feature map is used as the input of the face region attention network, and is processed by the face region attention network to obtain a visibility score feature map; The obtained visibility score feature map and the encoded feature map are correlated through the following formula to obtain the corresponding visible region feature map F att : F att = F?(1 + A); Among them, A represents the obtained visibility score feature map, and F represents the encoded feature map; The visible region feature map is connected to the feature map output after the last transposed convolutional block of the decoding layer in the skip connection to obtain the required output feature map.

7. A three-dimensional face reconstruction model training system, characterized in that The system includes a data acquisition module, a data processing module, a model construction module, and a model training module, where: Obtain a face dataset containing multiple face images and the feature point annotation information of each face image; According to the feature point annotation information, process each face image in the face dataset to obtain training supervision data, where the training supervision data includes a face region mask, a standard average face model after projection transformation, and standard face deformation information; Construct an initial 3D face reconstruction model for predicting the average face model and performing 3D face reconstruction based on face deformation information. The initial 3D face reconstruction model is composed of multiple Encoder-Decoder networks with the same structure, and the Encoder-Decoder network adds a channel-spatial attention perception mechanism in the skip connection between the encoding layer and the decoding layer; Based on the face image and the corresponding face region mask, model training is performed. During the training process, it is constrained by a target loss function used to reflect the deviation degree between the prediction result and the corresponding standard result, and when the training end condition is reached, a target 3D face reconstruction model is obtained; The face region attention network consists of a fully connected layer for converting the number of channels of the feature map and multiple standard convolutional blocks applying the channel-spatial attention perception mechanism, and activates through the sigmoid function to output the visibility of the face region at each pixel point. Among them, the feature extraction operation is implemented by the following formula: M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))); where, F ∈ R CxHxW denotes the input feature map, M c (*) denotes performing channel attention processing on "*", M s (*) denotes performing spatial attention processing on "*"; AvgPool(*) denotes performing average pooling processing on "*", MaxPool(*) denotes performing max pooling processing on "*"; MLP(*) denotes shared weight processing, σ denotes the sigmoid function; Conv(*) denotes a standard convolution operation, F Avg 、F Max distribution represents the 2D map obtained by applying the average pooling operation and the maximum pooling operation along the channel axis respectively; When representing the prediction result in the form of a UV position map, the calculation formula of the target loss function includes: In the above formula, h and w represent the height and width of the UV position map; N(u, v) represents the prediction result predicted based on the position coordinate points (u, v) in the UV space; Denote the training supervision data corresponding to the prediction result, i.e., the standard result, and M(u, v) denotes the weight value attached to the position coordinate point (u, v) in the UV space; Among them, to ensure the accuracy of the prediction results, the following constraint term L for the landmark points - key facial feature points is imposed on the average face model after projective transformation lrr : Among them, P(u, v) represents the predicted three-dimensional coordinate information of the landmark points on the average human face predicted based on the position coordinate points (u, v) in the UV space. It represents the standard three-dimensional coordinate information of the landmark points on the average face model corresponding to the position coordinate points (u, v) in the UV space in the training supervision data.

8. A readable storage medium, characterized in that, The readable storage medium includes a three-dimensional face reconstruction model training method program. When the three-dimensional face reconstruction model training method program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Image conversion method and device, computer equipment and storage medium

    CN111489287A

  • Image target area acquisition method and device, equipment, medium and program product

    CN114299101A