Robust three-dimensional face reconstruction method of deep guidance double-flow network

By adopting a deep-guided dual-stream network architecture and a two-way cross attention module in three-dimensional face reconstruction, integrating RGB images and depth information, the pose uncertainty and scale ambiguity problems of three-dimensional face reconstruction in complex scenarios are solved, and higher robustness and accuracy are achieved.

CN120147557AActive Publication Date: 2025-06-13WENZHOU UNIV METAVERSE & ARTIFICIAL INTELLIGENCE RES INST
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510615584.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing three-dimensional deformable model (3DMM) method faces problems such as complex lighting conditions, severe occlusion and large pose changes in complex poses in complex scenarios, especially in outdoor scenarios, resulting in pose uncertainty and scale ambiguity in three-dimensional face reconstruction.

Method used

Using a depth-guided dual-stream network architecture, features are extracted from RGB images and depth maps through independent encoders, and a two-way cross attention module and facial geometry perception module are designed to integrate image and depth information to enhance the robustness and accuracy of three-dimensional face reconstruction.

Benefits of technology

Effectively suppress background interference, improve the accuracy of three-dimensional face position and scale estimation, and significantly improve the robustness and accuracy of three-dimensional face reconstruction, especially in complex lighting and large posture changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147557A_ABST
    Figure CN120147557A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional face reconstruction method based on a deep guidance double-flow network, constructs an end-to-end deep guidance double-flow network framework, and aims to improve the robustness and generalization ability of three-dimensional face reconstruction in a field complex environment. According to the method, image information and depth prior information are fused, so that the performance of a reconstruction system under the conditions of complex background, extreme illumination and large attitude change is remarkably enhanced. Specifically, the method comprises the following steps: firstly, designing a facial geometric perception module, fully utilizing semantic features and depth priori of an image, generating an accurate facial mask, and effectively inhibiting negative effects of background interference on reconstruction quality; on the basis, a two-way cross attention module is introduced, efficient interaction and fusion between RGB information and depth information are achieved, and the feature representation capacity is further enhanced. Through the synergistic effect of the key modules, the reconstruction network effectively relieves the problem of depth fuzziness, then a decoder generates a high-quality UV position map and a high-quality UV texture map, and finally three-dimensional face representation rich in details and high in precision is obtained. Experimental results show that the method is superior to the existing mainstream method in multiple standard evaluation indexes or achieves the current optimal performance, especially has excellent performance in high-difficulty scenes such as complex backgrounds and the like, and verifies the effectiveness, robustness and wide application potential of the proposed framework. Technical innovation and practical application value are provided for solving the problem of face reconstruction under the extreme environment condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to technical fields such as deep learning, intelligent image processing, convolutional neural network, and three-dimensional reconstruction. By utilizing appropriate image depth prior information, it particularly relates to a single-view face three-dimensional reconstruction method based on deep learning. Background Art

[0002] With the rapid development of computer technology and digital media technology, images of objects have become easier to obtain. However, an image is simply two-dimensional information with limited information conveyed. Therefore, how to obtain more information about an object has gradually become the focus of current research. Three-dimensional face reconstruction technology aims to accurately restore the three-dimensional structure of a face from a two-dimensional image and has important application values in fields such as augmented reality (AR), virtual reality (VR), and virtual character generation. Existing three-dimensional deformable model (3DMM) methods usually rely on extracting features from images to infer three-dimensional geometric information. However, the performance of these methods is often significantly limited in complex scenarios, such as in cases involving shadows and large pose changes. This limitation is mainly attributed to the lack of clear depth information constraints, resulting in pose uncertainty and scale ambiguity in the generated three-dimensional face. Researchers have proposed a large number of face three-dimensional reconstruction algorithms to solve the above problems.

[0003] Specifically, the three-dimensional deformable model (3DMM) has become one of the key methods to solve this problem. The three-dimensional deformable model (3DMM) constructs a three-dimensional face representation through statistical methods, which can significantly improve the fidelity and interpretability of the generated three-dimensional face, mainly due to the precise control of coefficients such as expression and shape. The application of existing three-dimensional deformable model (3DMM) methods in complex scenarios still faces many challenges. For example, in outdoor scenarios, factors such as complex lighting conditions, severe occlusion, and large pose changes will significantly affect the reconstruction effect. The root cause of these problems lies in the lack of clear depth information constraints, resulting in depth ambiguity, and further making it difficult to accurately estimate the pose uncertainty and scale of the three-dimensional face. Specifically, when relying only on RGB images, strong light will produce shadows on facial features, interfering with the extraction of geometric features and making it difficult to accurately estimate the facial orientation and scale. Similarly, at extreme viewpoints, due to the lack of reliable depth information as a geometric constraint, the direction error of the reconstructed face is often large. There are obvious limitations in the existing technology for three-dimensional face reconstruction in complex scenarios, and there is an urgent need for a method that can effectively integrate depth information to improve the robustness and accuracy of reconstruction. Summary of the Invention

[0004] To solve the above problems, the present invention proposes a robust 3D face reconstruction method based on a depth-guided two-stream network. By combining RGB information with depth prior information and making full use of their respective advantages, better 3D reconstruction effects can be achieved. On the one hand, an innovative bidirectional cross-attention module is adopted to capture complementary information between the image and depth modalities, improve the understanding of complex scenes, and integrate image and depth information to enhance 3D face reconstruction. On the other hand, to further improve the accuracy of 3D facial position and scale estimation, an accurate facial visibility mask is introduced to guide the network to focus on facial features. Finally, through the synergistic effect of two key modules, background interference is effectively suppressed, and an accurate 3D reconstruction model is obtained.

[0005] The technical solution adopted by the present invention is a robust 3D face reconstruction method based on a depth-guided two-stream network architecture. It is specifically implemented according to the following steps:

[0006] Step 1: Construction of a two-stream network guided by image and depth information. Provide a depth-guided two-stream network that extracts features from images and depth maps respectively through independent encoders, enabling full exploration of the independent features of the two modalities.

[0007] Step 2: Facial geometry perception module. To alleviate the problems of background interference and detail loss in facial image feature extraction, a facial geometry perception module is designed. This module combines the semantic information of the image with depth prior to generate an accurate facial mask, effectively suppressing background interference and improving the accuracy of 3D face position and scale estimation.

[0008] Step 3: Construction of a bidirectional cross-attention module. Through a bidirectional feature interaction mechanism, the previously extracted image and depth features are effectively interacted to promote the bidirectional information flow between RGB features and depth features. This mechanism not only realizes the complementary cooperation of image and depth information but also significantly enhances the feature representation ability, thereby improving the robustness and accuracy of the model in 3D face reconstruction tasks.

[0009] The features of the present invention also lie in:

[0010] The specific implementation process of Step 1 is as follows: Given an RGB image and a depth map , independent encoders and extract features from the RGB and depth images respectively to obtain feature representations and .

[0011] The specific implementation process of step 2 is as follows: To improve the accuracy of 3D face pose and scale estimation, the facial geometry perception module quickly combines the RGB information of the image with the depth prior to generate an accurate facial visibility mask, guiding the network to focus on key facial regions. First, the facial geometry perception module takes the image features and depth features and as inputs. Then, feature-level fusion is performed, expressed as:

[0012] Next, channel attention and spatial attention are used to model the feature dependency relationships:

[0013] To enhance the adaptive fusion between modalities, a confidence estimation module is introduced to generate a confidence map by calculating the per-pixel differences between the modality features:

[0014] In addition, a pixel attention module is used to achieve fine-grained feature fusion, where the inputs to the pixel attention are and , and the Hadamard product is denoted by :

[0015] Subsequently, is used to adjust the number of channels, and then a facial mask is generated through the activation function. Finally, the enhanced features are obtained by element-wise multiplication with and : and

[0016] The specific implementation process in step 3 is as follows: The specific implementation process in step 3 is as follows: To effectively utilize the complementary information between the image and the depth image, we design a bidirectional cross-attention module to enhance the multi-modal feature representation through a bidirectional feature interaction mechanism. First, taking the depth stream as an example, the above features are first subjected to a convolution operation to obtain and as inputs. Subsequently, we obtain the query, key, and value vectors by applying a linear transformation to the enhanced input features: , , where , and are learnable linear transformation matrices.

[0017] Next, the interaction between RGB and depth features is calculated through the attention mechanism, aiming to fully exploit the complementary information between these two modalities. Image features and depth features usually contain different types of spatial information, and using either modality alone may not be able to effectively capture all the key information. Through the attention mechanism, the model can pay more attention to the correlation between these two features when fusing them, thereby enhancing the final feature representation ability. Adding the image query to the attention-weighted depth features allows the model to introduce depth information while maintaining the original image features, enhancing the understanding of the scene structure. The specific formula is as follows: Further, to further enhance feature fusion and retain depth information, we fuse the features output by the attention mechanism with the input depth features through an addition operation. Then, we use a non-linear activation function to introduce non-linear feature representation, thereby improving the expression ability and robustness of the model: Similarly, the RGB stream undergoes the same operations. Finally, the fused features of the two branches are input into the fusion module after one round of convolution respectively to generate the final feature representation: It should be noted that the overall architecture of this fusion module is very similar to the facial geometry perception module, but the main difference is the removal of the function in the final layer.

[0018] The bidirectional cross-attention module is the main feature extraction module of the 3D face reconstruction model. The reconstruction network fuses the attention feature map from the face attention supervision network and the depth detail supplementary information. Among them, the following two problems need to be solved: 1) Since the dimensions of image and depth information are different in space, the area and degree of occlusion of the occluded pixels in the image are unknown; 2) Since the depth channel prior contains inconsistent feature information compared with the original image, how to effectively fuse them together and make full use of the detail prior is the key to improving the quality of the 3D reconstruction model. Through the bidirectional feature interaction mechanism, the information flow between image features and depth features is promoted, and the fused features are obtained, realizing the complementary cooperation of image and depth information, significantly enhancing the feature representation ability of the model, and improving the robustness and accuracy of the final 3D face reconstruction.

[0019] In summary, the main contributions of the present invention are as follows:

[0020] (1) The present invention provides a robust 3D face reconstruction method based on a depth information-guided two-stream network. By using independent encoders to extract the features of images and depth respectively, a depth information-guided two-stream network framework is designed. The interaction between features is enhanced through a bidirectional cross-attention module, and the complementary information across modalities is optimized. The image input of the cross-attention supervised network is coordinated by depth information and image information, and an attention feature map that senses the uneven feature distribution information of complex backgrounds and large poses is output. The generated attention feature map will guide the main network to perform feature screening, focus on more important feature regions, and significantly improve the robustness and accuracy of 3D face reconstruction.

[0021] (2) The present invention provides a robust 3D face reconstruction method based on a depth information-guided two-stream network. To alleviate the problems of background interference and detail loss in face image feature extraction, a supplementary face geometry perception module is introduced. This module combines the semantic information of the image and the depth prior to generate an accurate face mask, effectively suppressing background interference, enhancing the feature representation ability, and improving the accuracy of 3D face position and scale estimation.

[0022] (3) The present invention provides a robust 3D face reconstruction method based on a depth information-guided two-stream network. To enhance the model's feature representation ability, a bidirectional cross-attention module is introduced. As the main feature extraction module of the 3D face reconstruction model, the reconstruction network fuses the attention feature map from the face attention supervised network and the depth detail supplementary information. Through the bidirectional feature interaction mechanism, the information flow between image features and depth features is promoted, and the fused features are obtained, realizing the complementary cooperation of image and depth information, significantly enhancing the model's feature representation ability, and introducing depth information as a geometric constraint to solve the problems of pose uncertainty and scale ambiguity of traditional methods under complex lighting, shadows, and large pose changes. Description of the Drawings

[0023] FIG. 1 is a schematic diagram of the overall structure of a robust 3D face reconstruction method based on a depth-guided two-stream network proposed by the present invention;

[0024] FIG. 2 is a schematic diagram for comparing the 3D face reconstruction effects proposed by the present invention. Detailed Embodiments

[0025] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the described embodiments are only for facilitating the understanding of the present invention and do not impose any limitation on it. The accompanying drawings are all in a very simplified form and use non-precise scales, only for conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention. The structures shown in the drawings are part of the actual structures. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0026] The present invention proposes a three-dimensional face reconstruction method based on a depth information-guided two-stream network framework. First, the network extracts the features of the image and the depth map through independent encoders respectively. Secondly, by using the facial geometry perception module, by combining the semantic information of the image and the depth prior, a refinement network is constructed to make up for the background interference problem in the reconstruction process. Finally, through the bidirectional cross-attention module, the attention feature map and the depth detail supplementary information are fused to fuse the features, realize the three-dimensional face reconstruction with enhanced representation ability, and effectively improve the accuracy of three-dimensional face position and scale estimation.

[0027] The overall model is shown in Figure 1. It includes three modules, namely, the depth-guided multi-modal feature extraction two-stream network, the facial geometry perception module, and the bidirectional cross-attention module.

[0028] The purpose of the two-stream feature extraction network is to generate an image semantic and depth information feature map with non-uniform feature distribution information, which serves as the attention basis for subsequent main network feature selection.

[0029] First, the depth-guided multi-modal feature extraction two-stream supervised network proposed by the present invention follows the classic encoding-decoding structure. Taking the RGB image and the depth map as the inputs of the two-stream network, the image and depth features are extracted through independent encoders respectively, and an image semantic feature map and a depth feature map with the same size are obtained.

[0030] Next, in order to improve the accuracy of the network in three-dimensional face estimation, the present invention introduces a facial geometry perception module. This module combines the RGB information of the image with the depth prior information to generate an accurate facial visibility mask. This mask can guide the network to focus on the key facial regions and suppress the interference of irrelevant regions, thereby improving the accuracy of pose and scale estimation. Specifically, the facial geometry perception module performs feature-level fusion on the image features and the depth features, and uses channel and spatial attention mechanisms to model the dependencies between the features, further enhancing the adaptive fusion between the modalities.

[0031] The facial geometry perception module receives image features and depth features as inputs. To model the dependencies between features, channel attention and spatial attention mechanisms are adopted. To enhance the adaptive fusion between modalities, a confidence estimation module is introduced, which generates a confidence map by calculating the per-pixel differences between RGB and depth modality features and processes it through an activation function. In addition, the pixel attention module further improves the fusion effect through fine-grained feature fusion. Finally, the number of channels is adjusted and an activation function is used to generate a facial mask , and then the enhanced feature representation is obtained by element-wise multiplication with .

[0032] To effectively utilize the complementary information between the image and the depth image, the present invention designs a bidirectional cross-attention module. This module enhances the multi-modal feature representation of RGB and depth features through a bidirectional feature interaction mechanism. In this process, the network first performs a convolution operation on the input features to obtain enhanced input features, and further calculates the interaction between RGB and depth features through an attention mechanism. This process can fully exploit the correlation between the two modalities, thereby improving the estimation accuracy of the model for 3D face pose and scale.

[0033] Finally, a decoder network is used to perform regression prediction on the reconstructed face image. The purpose of the decoder is to integrate the features of the network and the detail refinement network to achieve the final 3D face reconstruction work.

[0034] The decoder network gradually restores the low-dimensional feature map to a higher-dimensional spatial resolution through a series of gradually upsampling convolution operations. The input features first pass through a preliminary decoding module, which consists of multiple transposed convolution layers. Each convolution operation can effectively expand the spatial resolution of the feature map and gradually increase the number of channels. After being processed by these transposed convolution layers, the network can gradually restore the detail information of the image, ensuring that the final output has sufficient spatial information and feature expression ability.

[0035] Based on the feature map after preliminary decoding, the network is divided into two independent branches for further feature generation. The final output of the first branch is a three-channel texture coordinate position map. The second branch is responsible for generating the horizontal and vertical coordinate maps. The two branches output the key spatial information related to face reconstruction through different decoding paths.

[0036] Finally, these outputs work together to provide complete spatial coordinates for the reconstructed face, and finally restore the structure and details of the original face. This process not only accurately captures the position and shape of facial features, but also can reconstruct facial textures at a fine-grained level, providing rich feature information for subsequent face analysis tasks.

[0037] In summary, the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A robust 3D face reconstruction method based on a deep guided two-stream network, characterized by: Step 1: The dual-stream network guided by image and depth information is constructed. The preliminary features of the image and depth map are extracted through independent encoders. Then, the facial geometry perception module is used to combine the semantic information of the image and the depth prior to guide the network to focus on the features of the face area. The interaction between features is enhanced by using and through the bidirectional cross-attention module to optimize the complementary information across modalities. Finally, the decoder network is used to reconstruct the 3D face. Step 2: Construction of facial geometry perception module. To alleviate the problems of background interference and detail loss in facial image feature extraction, a facial geometry perception module is designed. This module combines the semantic information of the image with the depth prior to generate an accurate facial mask, effectively suppress background interference, and improve the accuracy of 3D face position and scale estimation; Step 3: Construction of the bidirectional cross-attention module, as the main feature extraction module of the face 3D reconstruction model, the network fuses the aforementioned feature maps and promotes the information flow between image features and depth features through a bidirectional feature interaction mechanism to obtain the fused features, thereby achieving complementary synergy between image and depth information, significantly enhancing the feature representation capability of the model, and being responsible for improving the robustness and accuracy of the final 3D face reconstruction.

2. The method for robust 3D face reconstruction using a deep-guided two-stream network as claimed in claim 1, characterized in that: The construction of the dual-stream network guided by the image and depth information in step 1 adopts the classic encoding and decoding structure as the main network structure. At the same time, the hybrid feature extraction module containing image information and depth information is used as the core of the network extraction to predict and regress the feature map of the feature image semantics and depth position.

3. The method for robust 3D face reconstruction using a deep-guided two-stream network as claimed in claim 1, characterized in that: The construction of the bidirectional cross attention module in step 3 uses the bidirectional cross attention feature module as its feature extraction core, which can effectively obtain feature maps of different modalities and enrich the feature representation capability of the overall network.

Citation Information

Patent Citations

  • Three-dimensional face modeling method based on double-tributary network

    CN112288851A

  • 360-degree environment depth completion and map reconstruction method based on cross-modal fusion

    CN114119889A

  • Image semantic segmentation method and device based on attention guidance multi-modal feature fusion

    CN114372986A

  • Multi-view three-dimensional reconstruction method based on attention mechanism and variable convolutional depth network

    CN116310098A

  • Double-stage face analysis method based on depth estimation and cross-modal feature sharing

    CN116778552A