Unsupervised multi-view stereo reconstruction method based on bidirectional attention mechanism
The unsupervised multi-view stereo reconstruction method using a bidirectional attention mechanism solves the problem of low 3D reconstruction quality in weak texture regions and complex occluded scenes, achieving more accurate 3D reconstruction results, especially in the preservation of details at edges and in areas of abrupt depth changes.
Patent Information
- Application Number
- CN202511038633.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-18
AI Technical Summary
Existing unsupervised multi-view stereo reconstruction methods have low 3D reconstruction quality in weakly textured regions and complex occluded scenes. They have difficulty distinguishing geometric surfaces with similar appearance features but separate spatial locations, and lack explicit modeling capabilities for the photometric properties of complex materials, resulting in distortion of the reconstructed point cloud structure and loss of details.
An unsupervised multi-view stereo reconstruction method based on bidirectional attention mechanism is adopted. Cross-scale connections are constructed through the bidirectional feature pyramid network Bi-FPN. Combined with attention-guided feature fusion module and multi-scale attention mechanism, multi-stage depth estimation and data augmentation are performed to optimize the initial depth map to generate a refined 3D point cloud.
It significantly improves the model's robustness to changes in lighting and occlusion interference, effectively preserves low-level geometric details and enhances high-level semantic representation, generating more accurate 3D reconstruction results, especially performing well in edge and depth abrupt change regions.
Smart Images

Figure CN120976423A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and deep learning technology, and further relates to multi-view stereo vision technology, specifically an unsupervised multi-view stereo reconstruction method based on a bidirectional attention mechanism, which can be used for 3D reconstruction of complex occluded scenes. Background Technology
[0002] Multi-view stereo (MVS) systematically constructs a complete technological chain from two-dimensional visual signals to three-dimensional geometric representations, playing a pivotal role in the virtual-real mapping between the physical world and digital space. In recent years, researchers have gradually introduced deep learning into the field of 3D reconstruction, developing two main technical routes: supervised MVS methods and unsupervised MVS methods.
[0003] The rapid development of unsupervised MVS methods is driven by the core innovation of constructing a self-consistent geometry-photometric optimization framework. This framework establishes reprojection constraints between multiple views through differentiable homography transformations and utilizes a photometric consistency loss function to drive the network to learn scene geometric priors. For example, patent application CN119417995A, entitled "Unsupervised Multi-View Stereo Reconstruction Method for Anti-Occlusion Regions," discloses an unsupervised multi-view stereo reconstruction method for anti-occlusion regions. This invention effectively mines the features of the input image itself by introducing a structured occlusion module and a perceptual consistency module, estimating a high-quality depth map, and then calculating an accurate point cloud model. However, traditional unsupervised methods struggle to distinguish geometric surfaces with similar appearance features but spatially separated, leading to noise accumulation in high-frequency details such as leaf edges and building outlines. Furthermore, existing methods lack explicit modeling capabilities for the photometric properties of complex materials. In nonlinear lighting interaction scenarios such as specular reflection and semi-transparent media, the reconstructed point cloud is prone to structural distortion and loss of detail. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of the prior art by proposing an unsupervised multi-view stereo reconstruction method based on a bidirectional attention mechanism, which solves the technical problem of low 3D reconstruction quality in weak texture regions and complex occluded scenes in the prior art.
[0005] To achieve the above objectives, the technical solution adopted by the present invention includes the following steps:
[0006] (1) Input multi-view images and camera parameters:
[0007] For a real-world scene requiring 3D reconstruction, input images I = {I1,…,I2} taken by the same camera from N viewpoints. i ,…,I N The pose of the camera when capturing each image is T = {T1, ..., T2}.i ,…,T N}, Camera internal parameters K, selecting I1 as the reference image, I2~I N Source image;
[0008] (2) Extract the original feature map of the input image:
[0009] The input image I is initially extracted using the Bi-FPN (Bi-Pyramid Feature Network). i The multi-scale features, namely the Q-stage basic feature maps F′1, F′2...F′ Q And Q is not less than 2; the resolution levels correspond sequentially to the input image I. i of ...original resolution;
[0010] (3) Calculate the fusion feature map of the input image:
[0011] Input image I i The original feature map F s As input data, the enhanced feature map F is obtained through an attention mechanism. s "; Subsequently, the original feature map F is fused together using an adaptive fusion network." s ′ and enhanced feature map F s "The images are fused to obtain the input image I." i fusion feature map F s Where s = 1, 2, ..., Q;
[0012] (4) Perform multi-stage depth estimation from coarse to fine:
[0013] Each input image I i The fusion feature maps F1, F2, ..., F Q The data are used as input data for the first stage, the second stage, ..., the Qth stage of depth estimation. Each stage goes through five sub-steps in sequence: depth sampling, projecting feature map, generating cost volume, cost volume regularization, and cost volume regression. Finally, the initial depth map of each stage is generated. The depth map of the reference image I1 output in the Qth stage is used as the initial depth map estimated by the algorithm.
[0014] (5) Calculate the initial depth map for the data augmentation branch:
[0015] The input image I is enhanced by improving its brightness, contrast, and occlusion simulation to obtain the enhanced image I′. The enhanced image I′ is then processed in the same way as the input image I to obtain the initial depth map for data augmentation branch prediction.
[0016] (6) Integrate multi-source information to optimize the initial depth map:
[0017] The initial depth map of the reference image I1, the initial depth map predicted from the original input image, and the initial depth map predicted by the data augmentation branch are used as multi-source optimization information to optimize the initial depth map of the reference image I1 and generate the final refined depth map.
[0018] (7) Convert the refined depth map into a point cloud model:
[0019] Based on the camera pose T1 and camera intrinsic parameters K when the reference image I1 is captured, the refined depth map is converted into 3D point cloud data to obtain the final 3D reconstruction result.
[0020] Compared with the prior art, the present invention has the following advantages:
[0021] First, this invention constructs bidirectional cross-scale connections through a bidirectional feature pyramid network (Bi-FPN), integrating top-down and bottom-up multi-level feature flows. While breaking through the information transmission bottleneck of traditional unidirectional feature pyramids, it effectively preserves the geometric details of the lower layers and enhances the semantic representation of the higher layers, providing robust contextual support for multi-scale depth estimation.
[0022] Secondly, because the present invention designs an attention-guided feature fusion module AFFM, combined with an efficient multi-scale attention mechanism EMA, it enhances the feature response of key regions (such as edge and depth abrupt change regions) through attention, thereby significantly improving the robustness of the model to illumination changes and occlusion interference. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the implementation of the present invention;
[0024] Figure 2 This is a schematic diagram of the overall architecture of the network model of the present invention;
[0025] Figure 3 This is a schematic diagram of the Bi-FPN network model provided in Embodiment 2 of the present invention;
[0026] Figure 4 This is a schematic diagram of the EMA network structure provided in Embodiment 2 of the present invention;
[0027] Figure 5 This is a data augmentation effect diagram provided in Embodiment 2 of the present invention;
[0028] Figure 6 This is a magnified view of the reconstructed point cloud details of the T&T dataset provided in the experimental section of this invention.
[0029] Figure 7 This is a Family scene reconstruction error map provided in the experimental section of this invention. Detailed Implementation
[0030] The present invention will be further described in detail below with reference to the accompanying drawings.
[0031] Example 1: Refer to Figure 1-2 The present invention proposes an unsupervised multi-view stereo reconstruction method based on a bidirectional attention mechanism, which specifically includes the following steps.
[0032] Step 1. Input multi-view images and camera parameters:
[0033] For a real-world scene requiring 3D reconstruction, input images I = {I1,…,I2} taken by the same camera from N viewpoints. i ,…,I N The pose of the camera when capturing each image is T = {T1, ..., T2}. i ,…,T N}, Camera internal parameters K, selecting I1 as the reference image, I2~I N Source image;
[0034] Step 2. Extract the original feature map of the input image:
[0035] The input image I is initially extracted using the Bi-FPN (Bi-Pyramid Feature Network). i The multi-scale features, namely the Q-stage basic feature maps F′1, F′2...F′ Q And Q is not less than 2; the resolution levels correspond sequentially to the input image I. i of ...original resolution;
[0036] Step 3. Calculate the fusion feature map of the input image:
[0037] Input image I i The original feature map F s As input data, the enhanced feature map F is obtained through an attention mechanism. s "; Subsequently, the original feature map F is fused together using an adaptive fusion network." s ′ and enhanced feature map F s "The images are fused to obtain the input image I." i fusion feature map F s Where s = 1, 2, ..., Q. The attention mechanism described above in this embodiment includes at least the efficient multi-scale attention mechanism EMA, the compressed excitation network SENet, the bottleneck attention module BAM, and the convolutional attention module CBAM.
[0038] In this embodiment, the adaptive fusion network in this step specifically utilizes a weight map W of the same size as the feature map to drive the network and realize the original feature map F. s ′ and enhanced feature map F sThe fusion operation involves the weights of each spatial location in the weight map W, which range from 0 to 1. The fusion network automatically adjusts the weight allocation strategy through backpropagation, enabling the deep learning model to adaptively adjust the weight values stored in the weight map W for different task requirements. The input image I... i fusion feature map F s Calculate according to the following formula:
[0039] F s =W⊙F s "+(1-W)⊙F s ′
[0040] Here, ⊙ represents element-wise multiplication.
[0041] Step 4. Perform multi-stage depth estimation from coarse to fine:
[0042] Each input image I i The fusion feature maps F1, F2, ..., F Q The data is sequentially used as input for stages 1, 2, ..., Q of depth estimation. Each stage sequentially involves five sub-steps: depth sampling, projecting feature maps, generating cost volumes, cost volume regularization, and cost volume regression. This process ultimately generates the initial depth map for each stage. The depth map of the reference image I1 output from stage Q is used as the initial depth map estimated by the algorithm. The five sub-steps are as follows:
[0043] (4.1) Sample the depth values within the depth assumption interval of the current stage to obtain multiple depth assumption values for each pixel in the input image I; the specific operation of this step in this embodiment is as follows: if the current stage is the 1st stage, then uniform sampling is performed; if the current stage is the qth stage, then sampling is performed near the depth value predicted in the q-1th stage; q = 2, 3, ..., Q.
[0044] (4.2) Based on multiple depth assumptions for each pixel, project the feature maps corresponding to all source images onto the viewpoint of the reference image;
[0045] (4.3) Use variance metric to aggregate the features of the corresponding projections of multiple source images to form a cost volume;
[0046] (4.4) Use 3D convolution to perform regularization on the cost volume to generate the probability volume;
[0047] (4.5) Perform Soft Argmin calculation on the probability volume pixel by pixel to generate the initial depth map for the current stage.
[0048] Step 5. Calculate the initial depth map for the data augmentation branch:
[0049] The input image I is enhanced by improving its brightness, contrast, and occlusion simulation to obtain the enhanced image I′. The enhanced image I′ is then processed in the same way as the input image I to obtain the initial depth map for data augmentation branch prediction.
[0050] Step 6. Integrate multi-source information to optimize the initial depth map:
[0051] The initial depth map of the reference image I1, the initial depth map predicted from the original input image, and the initial depth map predicted by the data augmentation branch are used as multi-source optimization information to optimize the initial depth map of the reference image I1 and generate the final refined depth map.
[0052] Step 7. Convert the depth map into a point cloud model:
[0053] Based on the camera pose T1 and camera intrinsic parameters K when the reference image I1 is captured, the refined depth map is converted into 3D point cloud data to obtain the final 3D reconstruction result.
[0054] Example 2: Refer to Figure 3-5 The overall implementation steps of the multi-view stereo reconstruction method proposed in this embodiment are the same as those in Embodiment 1. The parameter settings are given below, and the implementation process of this invention is further described in detail using three-stage feature extraction as an example:
[0055] Step 1) Input multi-view images and camera parameters:
[0056] For a real-world scene requiring 3D reconstruction, input images I = {I1,…,I2} taken by the same camera from N viewpoints. i ,…,I N The pose of the camera when capturing each image is T = {T1, ..., T2}. i ,…,T N}, Camera internal parameters K, selecting I1 as the reference image, I2~I N This is the source image.
[0057] Step 2) Extract the original feature map of the input image:
[0058] The input image I is initially extracted using a Bidirectional Feature Pyramid Network (Bi-FPN). i Multi-scale features are used to obtain the input image I. i The basic feature maps F′1, F′2, and F′3 for the first, second, and third stages, respectively, correspond to the input image I at different resolution levels. i 1 / 4, 1 / 2, and original resolutions; the Bi-FPN network model used in this embodiment is as follows: Figure 3 As shown.
[0059] Step 3) Calculate the fusion feature map of the input image:
[0060] Input image I i The original feature map F s As input data, the enhanced feature map F is obtained through an attention mechanism. s In this embodiment, it is preferable to use an efficient multi-scale attention mechanism (EMA) network to obtain enhanced feature maps. The EMA network structure is as follows: Figure 4 As shown; subsequently, a network-driven generation method is used with a weight map W of the same size as the feature map to construct a learnable weight map W, thereby realizing the original feature map F. s ′ and enhanced feature map F s The fusion operation involves fusing the feature maps F and F. The weights at each spatial location in the weight map W range from 0 to 1. The fusion network automatically adjusts the weight allocation strategy through backpropagation, enabling the model to adaptively adjust the proportions to meet different task requirements. s The calculation formula is as follows:
[0061] F s =W⊙F s "+(1-W)⊙F s ′
[0062] Where ⊙ represents element-wise multiplication, and s = 1, 2, 3.
[0063] Step 4) Perform multi-stage depth estimation from coarse to fine:
[0064] Each input image I i The fused feature maps F1, F2, and F3 are used as input data for the first, second, and third stages of depth estimation, respectively. Each stage contains five sub-steps:
[0065] Sub-step 1: Sample the depth value within the depth assumption interval of each stage. If it is the second or third stage, sample near the predicted depth value of the previous stage; otherwise, sample uniformly.
[0066] Sub-step 2: Based on several depth assumptions for each pixel, project the feature maps corresponding to all source images onto the viewpoint of the reference image;
[0067] Sub-step 3: Use variance metric to aggregate the features projected from several source images to form a cost volume;
[0068] Sub-step 4: Use 3D convolution to perform regularization on the cost volume and generate the probability volume;
[0069] Sub-step 5: Perform Soft Argmin calculation on the probability volume pixel by pixel to generate the initial depth map.
[0070] Finally, the depth map of the reference image I1 output from the third stage is used as the initial depth map estimated by the algorithm.
[0071] Step 5) Calculate the initial depth map for the data augmentation branch:
[0072] Enhanced image I′ is obtained by enhancing the input image I in three aspects: brightness, contrast, and occlusion simulation. Brightness adjustment is achieved by adding or subtracting pixel values to make the image brighter or darker overall; contrast adjustment is achieved by changing the distribution range of pixel values to make the contrast between light and dark areas stronger or weaker; occlusion simulation is achieved by randomly selecting a rectangular area in the input image and setting the pixel values of that area to zero or filling it with random noise. The image enhancement effect is as follows. Figure 5 As shown, the augmented image I′ is processed using the same method as the input image I to obtain the initial depth map for data augmentation branch prediction.
[0073] Step 6) Integrate multi-source information to optimize the initial depth map:
[0074] The reference image I1, the initial depth map predicted from the original input image, and the initial depth map predicted by the data augmentation branch are concatenated to form a multi-channel feature map, which is then fed into a convolutional network for processing. The optimized features are then summed pixel-by-pixel with the initial depth map predicted from the original input image to generate the final refined depth map. The training process employs a multi-task joint optimization loss function design, comprising the following four parts:
[0075] Photometric data enhancement loss (L DA ): Enhanced images are generated by random brightness and contrast perturbation and regional occlusion. The difference in depth prediction between the enhanced image and the original image is constrained by the Smooth-L1 loss, thereby enhancing the robustness of the model to illumination changes and occluded scenes.
[0076] Structural similarity loss (L SSIM ): Based on the structural similarity measure of image patches, it captures the structural consistency information of depth prediction by comparing the local mean, variance and covariance of the projected image and the reference image.
[0077] Smoothness loss (L smooth ): Constrains the consistency of depth value changes within the local neighborhood of the depth map, suppressing discontinuity noise in the prediction results by minimizing the depth difference between adjacent pixels;
[0078] Photometric uniformity loss (L) PC ): Generate composite views through inverse projection, jointly optimize pixel color differences and gradient differences between reference views and source views, and enhance cross-view geometric consistency;
[0079] The final loss function uses a weighted sum of the above four terms:
[0080] L=λ1LDA +λ2L SSIM +λ3L smooth +λ4L PC
[0081] λ1 to λ4 are pre-selected parameter values used to balance the contributions of different loss terms. This design ensures the geometric accuracy of depth prediction while also taking into account the model's adaptability to complex scenes.
[0082] Step 7) Convert the depth map into a point cloud model:
[0083] By using the camera pose T1 and camera intrinsic parameters K when capturing the reference image I1, the depth map of image I1 can be converted into 3D point cloud data, thus obtaining the final 3D reconstruction result.
[0084] The technical effects of the present invention will be further explained below with reference to simulation experiments.
[0085] 1. Simulation conditions and content:
[0086] The simulation environment was built on an AutoDL cloud server with the following configuration: an AMD EPYC 9754 CPU, an NVIDIA GeForce RTX 4090D graphics card with 24GB of video memory, an Ubuntu 22.04 operating system, an algorithm framework built using PyTorch, and Python code written in Python.
[0087] The simulation data comes from the T&T dataset, a highly authoritative dataset renowned for its complexity. It utilizes industrial-grade laser scanning to accurately capture point cloud information from real-world scenes, constructing reliable ground truth point clouds to ensure accurate evaluation. This dataset encompasses both outdoor and indoor environments, showcasing a wide range of complex visual characteristics, such as fine textures, occlusion, and even self-occlusion. The scenarios involved are diverse, including individual objects such as tanks and trains.
[0088] The T&T dataset was reconstructed using the method of this invention. A magnified image of some reconstructed point cloud details is shown below. Figure 6 As shown; simultaneously, the 3D reconstruction quality of the method of this invention and existing methods in the T&T dataset Family scene is compared, and the results are as follows. Figure 7 As shown.
[0089] 2. Simulation Result Analysis:
[0090] Reference Figure 6 The method of this invention has been found to accurately depict the complex geometric contours and surface features of reconstructed objects. It also effectively preserves sharp structural continuity in easily distorted areas such as edges and corners.
[0091] Reference Figure 7 This paper compares the 3D reconstruction quality of existing methods (RC-MVSNet, DS-MVSNet) with the method of this invention (MSBA-MVSNet) in the T&T dataset Family scene. The first row is an error graph comparison based on accuracy metrics. Visual analysis shows that the method of this invention has significant advantages in surface geometry reconstruction, with fewer noise points and clearer object contours. The second row is an error graph based on integrity metrics. The point cloud reconstructed by the method of this invention has fewer voids on the curved surface of the sculpture's leg, and the overall reconstructed point cloud is more complete.
[0092] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0093] The above simulation analysis proves the correctness and effectiveness of the method proposed in this invention.
[0094] The parts of this invention not described in detail are common knowledge to those skilled in the art.
[0095] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Obviously, those skilled in the art, after understanding the content and principle of the present invention, may make various modifications and changes in form and detail without departing from the principle and structure of the present invention. However, these modifications and changes based on the concept of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. An unsupervised multi-view stereo reconstruction method based on a bidirectional attention mechanism, characterized in that: Includes the following steps: (1) Input multi-view images and camera parameters: For a real-world scene requiring 3D reconstruction, input images I = {I1,…,I2} taken by the same camera from N viewpoints. i ,…,I N The pose of the camera when capturing each image is T = {T1, ..., T2}. i ,…,T N }, Camera internal parameters K, selecting I1 as the reference image, I2~I M Source image; (2) Extract the original feature map of the input image: The input image I is initially extracted using the Bi-FPN (Bi-Pyramid Feature Network). i The multi-scale features, namely the Q-stage basic feature maps F′1, F′2...F′ Q And Q is not less than 2; the resolution levels correspond sequentially to the input image I. i of Original resolution; (3) Calculate the fusion feature map of the input image: Input image I i The original feature map F′ s As input data, an enhanced feature map F″ is obtained through an attention mechanism. s Subsequently, the original feature map F′ is fused together using an adaptive fusion network. s With enhanced feature map F″ s The input image I is obtained by fusion. i fusion feature map F s ; Where s = 1, 2, ..., Q; (4) Perform multi-stage depth estimation from coarse to fine: Each input image I i The fusion feature maps F1, F2, ..., F Q The data are used as input data for the first stage, the second stage, ..., the Qth stage of depth estimation. Each stage goes through five sub-steps in sequence: depth sampling, projecting feature map, generating cost volume, cost volume regularization, and cost volume regression. Finally, the initial depth map of each stage is generated. The depth map of the reference image I1 output in the Qth stage is used as the initial depth map estimated by the algorithm. (5) Calculate the initial depth map for the data augmentation branch: The input image I is enhanced by improving its brightness, contrast, and occlusion simulation to obtain the enhanced image I′. The enhanced image I′ is then processed in the same way as the input image I to obtain the initial depth map for data augmentation branch prediction. (6) Integrate multi-source information to optimize the initial depth map: The initial depth map of the reference image I1, the initial depth map predicted from the original input image, and the initial depth map predicted by the data augmentation branch are used as multi-source optimization information to optimize the initial depth map of the reference image I1 and generate the final refined depth map. (7) Convert the refined depth map into a point cloud model: Based on the camera pose T1 and camera intrinsic parameters K when the reference image I1 is captured, the refined depth map is converted into 3D point cloud data to obtain the final 3D reconstruction result.
2. The method according to claim 1, characterized in that: The attention mechanisms described in step (3) include the efficient multi-scale attention mechanism EMA, the compressed excitation network SENet, the bottleneck attention module BAM, and the convolutional attention module CBAM.
3. The method according to claim 1, characterized in that: The adaptive fusion network described in step (3) specifically utilizes a weight map W of the same size as the feature map to drive the network and realize the original feature map F. s With enhanced feature map F″ s The fusion operation.
4. The method according to claim 3, characterized in that: The weight value of each spatial location in the weighted graph W is between 0 and 1.
5. The method according to claim 3, characterized in that: The fusion network automatically adjusts the weight allocation strategy through backpropagation, enabling the deep learning model to adaptively adjust the weight values stored in the weight graph W for different task requirements.
6. The method according to claim 3, characterized in that: The input image I i fusion feature map F s Calculate according to the following formula: F s =W⊙F″ s +(1-W)⊙F′ s Here, ⊙ represents element-wise multiplication.
7. The method according to claim 1, characterized in that: The initial depth map for each stage described in step (4) is generated according to the following five sub-steps: (4.1) Sample the depth values within the current stage depth hypothesis interval to obtain multiple depth hypothesis values for each pixel in the input image I; (4.2) Based on multiple depth assumptions for each pixel, project the feature maps corresponding to all source images onto the viewpoint of the reference image; (4.3) Use variance metric to aggregate the features of the corresponding projections of multiple source images to form a cost volume; (4.4) Use 3D convolution to perform regularization on the cost volume to generate the probability volume; (4.5) Perform Soft Argmin calculation on the probability volume pixel by pixel to generate the initial depth map for the current stage.
8. The method according to claim 7, characterized in that: Step (4.1) involves sampling the depth value within the depth assumption interval of the current stage. Specifically, if the current stage is stage 1, then uniform sampling is performed; if the current stage is stage q, then sampling is performed near the depth value predicted in stage q-1; q = 2, 3, ..., Q.
Citation Information
Patent Citations
Unsupervised multi-view three-dimensional reconstruction method of anti-occlusion area
CN119417995A
Cited By
Multi-view reconstruction method and system based on geometric perception and attention fusion
CN121304950A
Three-dimensional reconstruction method and system based on frequency-space double-domain characteristics and cascade optimization
CN122199838A