The present invention proposes a three-dimensional face
reconstruction method based on a depth-guided two-
stream network, constructs an end-to-end depth-guided two-
stream network, and improves the robustness and generalization in complex environments. This method integrates image and depth
prior information to enhance the performance of the reconstruction
system under complex backgrounds, extreme lighting and large postures. Specifically: a
facial geometry perception module is designed to generate accurate facial masks using image semantic features and depth priors to suppress background interference; a bidirectional cross-attention module is introduced to achieve efficient fusion of RGB and depth information and enhance feature representation capabilities; key modules work together to alleviate depth
ambiguity problems, generate high-quality UV position maps and texture maps, and obtain high-precision three-dimensional face representations. This method outperforms mainstream methods in multiple evaluation indicators, especially in difficult scenes such as complex backgrounds, verifying the effectiveness, robustness and application potential of the framework, and providing
technical innovation and practical value for face reconstruction in extreme environments.