Multi-view three-dimensional reconstruction method based on multi-scale feature fusion

Through a multi-scale feature fusion network architecture, combined with feature pyramid network and compression-excitation network, the problem of insufficient feature fusion in multi-view three-dimensional reconstruction is solved, and the resolution and accuracy of depth estimation is improved, and it is suitable for real-time or near-real-time application scenarios.

CN120580362APending Publication Date: 2025-09-02BEIHANG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510718764.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing multi-view three-dimensional reconstruction method lacks global multi-scale feature fusion during feature fusion, resulting in low depth estimation resolution and loss of detail information, especially in weak texture areas and complex scenarios.

Method used

A network architecture with multi-scale feature fusion is adopted, combining feature pyramid networks and compression-excitation networks, through iterative depth prediction and feature regularization, the interaction and fusion of multi-level features is achieved, and the resolution and accuracy of depth estimation are improved.

Benefits of technology

It significantly improves the accuracy and efficiency of three-dimensional reconstruction, can better balance global information and local details, and is suitable for real-time or near-real-time application scenarios, enhancing the network's adaptability to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580362A_ABST
    Figure CN120580362A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision and three-dimensional reconstruction, and particularly discloses a multi-view three-dimensional reconstruction method based on multi-scale feature fusion, which adopts a multi-scale feature fusion network architecture and comprises a feature extraction module, a cost body construction module, a cost body regularization network and a deep regression network. The method specifically comprises the following steps: constructing a multi-view three-dimensional network architecture based on multi-scale feature fusion; the method comprises the following steps: setting a feature extraction module, combining an FPN structure with a bidirectional feature fusion strategy, establishing a homography transformation and cost body construction module, generating a cost body, deploying a cost body regularization processing unit, carrying out global-local feature fusion by adopting a 3D UNet architecture, and executing a depth inference and refinement process. And optimizing a depth regression result through multi-stage depth hypothesis and a confidence weighting mechanism. And performing a three-dimensional reconstruction experiment and result analysis, implementing depth map fusion and three-dimensional reconstruction, and generating a dense three-dimensional point cloud model. According to the invention, the precision and integrity of three-dimensional reconstruction are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and three-dimensional reconstruction, and in particular to a multi-view stereo three-dimensional reconstruction method based on multi-scale feature fusion. Background Art

[0002] Multi-view stereo (MVS) vision aims to recover the three-dimensional geometric structure of a scene using images captured from multiple viewpoints. It has long been a key task in computer vision. Leveraging the rich geometric information provided by multiple viewpoints, MVS technology can reconstruct high-precision three-dimensional spatial models, driving innovations in fields such as embodied intelligence, spatial intelligence, autonomous driving, biological sciences, and cultural heritage preservation. Traditional MVS methods typically rely on feature descriptions or similarity metrics, limiting their performance in scenes with weak textures, varying illumination, and occlusions. In recent years, with the rapid development of artificial intelligence (AI), deep learning-driven models have achieved end-to-end network model learning, directly transforming the mapping from two-dimensional images to three-dimensional models. Using methods such as convolutional neural networks (CNNs) to obtain deep information about target objects or scenes not only improves model accuracy but also demonstrates significant advantages in reducing resource consumption, shortening time consumption, and improving reconstruction efficiency. Therefore, this technical path that effectively integrates deep learning and multi-view stereo vision provides a new methodological framework for depth perception and geometric reconstruction of three-dimensional scenes, and has important research value and development prospects for promoting the application of three-dimensional vision technology.

[0003] MVSNet and its variants use shallow 2D convolutional networks to downsample the original image to extract high-level features with semantic information. However, due to computational resource constraints, the cost volume module often relies on high-level features, resulting in low-resolution depth maps. To improve the resolution of depth estimation, the cascaded MVS architecture introduces a Feature Pyramid Network (FPN) to extract multi-scale features. This architecture fuses features through a top-down pathway and lateral connections, enhancing the utilization of multi-scale information. While high-level features have a larger receptive field, they are less able to express details. Low-level features, on the other hand, retain more local information but are susceptible to noise. Existing cascaded networks typically fuse features between adjacent scales and lack deep fusion of global multi-scale features. This can lead to the loss of some details during feature transfer, compromising reconstruction accuracy. Furthermore, cascaded networks tend to rely more on deep-level features for depth estimation, while spatial details in shallower layers are often underutilized. This unbalanced feature utilization can lead to poor reconstruction of boundary regions or small objects, further limiting model performance.

[0004] In areas with rich texture information, the local receptive field can effectively capture details, while areas with weak texture need to be matched in a larger range.

[0005] Based on this, this study proposed a multi-view stereo 3D reconstruction method based on multi-scale feature fusion, which can effectively improve the reliability of feature expression. Summary of the Invention

[0006] In response to the above-mentioned problems in the prior art, the present invention provides a multi-view stereo 3D reconstruction method based on multi-scale feature fusion, which effectively improves the reliability of feature expression, enhances the effect of multi-scale feature fusion, and improves the depth perception capability of the MVS network.

[0007] To achieve the above objectives, the present invention proposes a multi-view stereo 3D reconstruction method based on multi-scale feature fusion, including: the method adopts a multi-scale feature fusion network architecture, and the multi-scale feature fusion network architecture includes a feature extraction module, a cost volume construction module, a cost volume regularization network and a deep regression network.

[0008] Preferably, the multi-view stereo 3D reconstruction method based on multi-scale feature fusion includes the following steps:

[0009] S1. Build a multi-view stereo network architecture based on multi-scale feature fusion, adopt a cascaded network architecture, use a coarse-to-fine depth prediction strategy, and generate a high-resolution depth map by iteratively refining the depth results;

[0010] S2. Set up a feature extraction module, using a feature pyramid network (FPN) structure and a bidirectional feature fusion strategy to extract multi-level features from the input multi-view images. This allows the contextual information between features of different scales to interact and generate a multi-level feature map containing spatial geometry information and texture details.

[0011] S3. Establish a homography transformation and cost volume construction module, define the homography transformation matrix based on the geometric projection relationship, map the source view features to the reference view coordinate system to form a feature volume, and use a feature aggregation method based on variance measurement to generate a cost volume;

[0012] S4. Deploy the cost body regularization processing unit, use the 3D convolutional network to regularize the cost body, integrate high-level semantic information with low-level detail information, and improve feature expression capabilities;

[0013] S5. Execute the depth inference and refinement process, regress the initial depth map through multi-stage depth hypothesis and soft-argmin operation, and perform fine depth estimation for each stage;

[0014] S6. Conduct 3D reconstruction experiments and analyze the results, implement depth map fusion and 3D reconstruction, combine photometric consistency and geometric consistency constraints, and generate a dense 3D point cloud model.

[0015] Preferably, the multi-view stereo network architecture for multi-scale feature fusion further includes a compression-excitation network SENet module.

[0016] Preferably, in the homography transformation and cost volume construction module, the homography transformation matrix is ​​jointly calculated by camera parameters and depth hypothesis.

[0017] Preferably, the cost volume regularization processing unit adopts a 3D UNet architecture and an encoder-decoder structure for global-local feature fusion.

[0018] Preferably, in the depth inference and refinement process, each stage generates a hypothesis set with adaptive depth intervals based on the previous depth map, and optimizes the depth regression result through a confidence weighting mechanism.

[0019] Preferably, the method further includes designing a multi-task loss function, selecting L1 depth loss, cross entropy classification loss and smoothing regularization term, and the calculation formula of the loss function is:

[0020]

[0021] Where p valid Represents the set of valid true value pixels, Represents the true depth value of pixel p in the lth stage, Represents the estimated depth value of pixel p in the lth stage.

[0022] Preferably, the bidirectional feature fusion strategy includes a top-down feature upsampling path and a bottom-up feature downsampling path, and achieves cross-scale information complementarity through feature splicing and attention mechanism.

[0023] Preferably, in S6, the three-dimensional reconstruction experiment is trained on the DTU dataset, and the Adam optimizer is used in the training phase, combined with the cosine annealing learning rate scheduling strategy.

[0024] Preferably, in S6, the three-dimensional point cloud model is processed by depth map fusion and hole filling post-processing, combined with normal vector estimation and mesh reconstruction algorithm to generate a three-dimensional model with complete surface.

[0025] Therefore, the present invention proposes a multi-view stereo 3D reconstruction method based on multi-scale feature fusion, which has the following beneficial effects:

[0026] (1) The present invention exchanges context information between multi-level features through multi-scale feature fusion, constructs a rich feature representation, effectively balances global information and local details, significantly enhances stereo matching performance, and improves the accuracy of 3D reconstruction;

[0027] (2) The SENet channel attention mechanism is introduced to dynamically adjust the feature channel weights, highlight key features, suppress redundant information, further enhance the multi-scale feature fusion effect, and improve the network's ability to understand key areas and depth perception;

[0028] (3) While ensuring high-precision reconstruction, it significantly reduces the inference time and hardware resource requirements, improves the efficiency of 3D reconstruction, and makes it more suitable for real-time or near-real-time application scenarios; it has wide applicability and strong adaptability to complex scenes, providing strong support for the application of 3D reconstruction technology in different fields.

[0029] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 2. It is a multi-view stereo network architecture diagram of the multi-view stereo 3D reconstruction method based on multi-scale feature fusion of the present invention;

[0031] Figure 2 This paper presents a multi-view stereo 3D reconstruction method based on multi-scale feature fusion, wherein (a) is a feature pyramid structure, (b) is a cascaded encoding-decoding structure, and (c) is the proposed multi-scale feature fusion structure.

[0032] Figure 3This is a diagram of the compression-excitation network architecture of the multi-view stereo 3D reconstruction method based on multi-scale feature fusion of the present invention;

[0033] Figure 4 This is a point cloud reconstruction effect diagram of some scenes on the DTU dataset using the multi-scale feature fusion MVS network of the multi-view stereo 3D reconstruction method based on multi-scale feature fusion of the present invention;

[0034] Figure 5 This is a comparison chart of the visualization effects of point cloud reconstruction using different MVS methods based on the multi-view stereo 3D reconstruction method of the present invention on the DTU dataset;

[0035] Figure 6 This is a point cloud reconstruction effect diagram of the Tanks&Temple dataset using the multi-scale feature fusion MVS network of the multi-view stereo 3D reconstruction method based on multi-scale feature fusion of the present invention;

[0036] Figure 7 It is the scene graph, confidence map, depth map and depth filter map effect of the Tanks&Temples dataset of the multi-view stereo 3D reconstruction method based on multi-scale feature fusion of the present invention. DETAILED DESCRIPTION

[0037] To make the technical solutions, advantages, and purposes of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below. The described embodiments are part of the embodiments of the present invention, not all of them. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0038] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0039] Example

[0040] 1. Multi-view Stereo Network with Multi-scale Feature Fusion

[0041] like Figure 1 As shown in the figure, a multi-view stereo 3D reconstruction method based on multi-scale feature fusion adopts a multi-scale feature fusion network architecture, which includes a feature extraction module, a cost volume construction module, a cost volume regularization network, and a depth regression network. The modules used in the present invention fully utilize the advantages of multi-view stereo vision and effectively improve the performance of depth estimation.

[0042] The multi-view stereo network architecture for multi-scale feature fusion also includes a compression-excitation network SENet module.

[0043] The steps of the multi-view stereo 3D reconstruction method based on multi-scale feature fusion include:

[0044] S1. Build a multi-view stereo network architecture based on multi-scale feature fusion, adopt a cascaded network architecture, use a coarse-to-fine depth prediction strategy, and generate a high-resolution depth map by iteratively refining the depth results;

[0045] S2. Set up a feature extraction module, using a feature pyramid network (FPN) structure and a bidirectional feature fusion strategy to extract multi-level features from the input multi-view images. This allows the contextual information between features of different scales to interact and generate a multi-level feature map containing spatial geometry information and texture details.

[0046] The bidirectional feature fusion strategy includes a top-down feature upsampling path and a bottom-up feature downsampling path, and achieves cross-scale information complementarity through feature splicing and attention mechanism.

[0047] S3. Establish a homography transformation and cost volume construction module, define the homography transformation matrix based on the geometric projection relationship, map the source view features to the reference view coordinate system to form a feature volume, and use a feature aggregation method based on variance measurement to generate a cost volume;

[0048] S4. Deploy the cost body regularization processing unit, use the 3D convolutional network to regularize the cost body, integrate high-level semantic information with low-level detail information, and improve feature expression capabilities;

[0049] In the homography and cost volume construction module, the homography matrix is ​​calculated jointly using camera parameters and depth assumptions. The cost volume regularization processing unit adopts a 3D UNet architecture and an encoder-decoder structure for global-local feature fusion.

[0050] S5. Execute the depth inference and refinement process, regress the initial depth map through multi-stage depth hypothesis and soft-argmin operation, and perform fine depth estimation for each stage;

[0051] In the depth inference and refinement process, each stage generates a hypothesis set with adaptive depth intervals based on the previous depth map, and optimizes the depth regression results through a confidence weighting mechanism.

[0052] S6. Conduct 3D reconstruction experiments and analyze the results, implement depth map fusion and 3D reconstruction, combine photometric consistency and geometric consistency constraints, and generate a dense 3D point cloud model.

[0053] The 3D reconstruction experiments were trained on the DTU dataset, using the Adam optimizer with a cosine annealing learning rate scheduling strategy. The 3D point cloud model was post-processed using depth map fusion and hole filling, combined with normal vector estimation and mesh reconstruction algorithms to generate a surface-complete 3D model.

[0054] The method also includes the design of multi-task loss function, using L1 depth loss, cross entropy classification loss and smoothing regularization term. The loss function is calculated as follows:

[0055]

[0056] Where p valid Represents the set of valid true value pixels, Represents the true depth value of pixel p in the lth stage, Represents the estimated depth value of pixel p in the lth stage.

[0057] 2. Feature Extraction Module

[0058] like Figure 2 As shown, extracting key information from the original image can provide efficient and robust feature representation for the network model. Shallow features have a small receptive field and are suitable for detecting and recognizing small objects. Deep feature maps have a large receptive field and are more suitable for processing semantic information in large scenes. To achieve better feature extraction at different scales, the FPN architecture was developed.

[0059] like Figure 2 As shown in (a), the classic FPN architecture extracts image features at different scales by constructing a multi-level pyramid network. The FPN generates high-level semantic features through a bottom-up approach and upsamples these features using a top-down approach to restore spatial resolution. Simultaneously, the upsampled features are fused with shallow features using lateral connections, combining global semantic information with local detail. This allows the network to capture multi-level image information, from large to small, and from detail to global. The FPN's feature representation capabilities play a key role in computer vision tasks such as object detection, semantic segmentation, and instance segmentation.

[0060] like Figure 2As shown in Figure (b), unlike FPN, which uses a lightweight fusion module for object prediction, CEDNet, based on a backbone module, first extracts high-resolution initial features and then gradually generates multi-scale features by combining multiple cascaded stages. The three stages of this network share the same encoder-decoder structure, performing multi-scale feature aggregation within each decoder, effectively enhancing the effect of multi-scale feature fusion. This strategy enables the network to capture high-level semantic features from an early stage and combine them with low-level details in subsequent stages, better guiding feature learning in subsequent stages and improving the multi-level expressiveness of features.

[0061] In multi-scale feature fusion, FPN achieves the concatenation of features at different scales through upsampling and cross-layer connections, but still underutilizes the contextual feature information between scales. Low-layer feature maps contain richer detailed information, but they are not well integrated with high-level features, limiting their expressiveness in complex scenarios. Furthermore, when FPN aggregates low- and high-level features through cross-layer connections, it lacks a dynamic adjustment mechanism tailored to task requirements. Low-level features may overweight high-level features, while high-level features may lack sufficient semantic information when passed to lower layers.

[0062] like Figure 2 As shown in (c), an efficient multi-scale feature fusion architecture is used. This framework performs deeper multi-scale feature fusion by exchanging contextual information between multiple levels. Features at different levels complement each other to construct richer feature representations. The model adopts a bidirectional fusion strategy: in the upsampling path, low-resolution, high-semantic information features are transferred to high-resolution features; in the downsampling path, high-resolution, low-semantic information features are transferred to low-resolution features. This design effectively avoids the problem of insufficient low-level feature expression caused by the one-way flow of information in FPN. The designed network structure can better realize the interaction and information complementation of features at different scales, organically combining global geometric structure and local information details, improving the feature extraction capability in MVS tasks and the accuracy of subsequent depth estimation.

[0063] like Figure 3 As shown in the figure, a compression and excitation network (SENet) model is introduced to improve the expressive ability of multi-scale feature fusion.

[0064] Specifically, SENet is a channel-attention mechanism model. A compression block compresses global spatial information to generate channel-level global feature descriptions. Feature learning is then performed in the channel dimension to generate a weight vector for each channel. Finally, an excitation block assigns different weights to different channels, enhancing important features and suppressing redundant ones. The core idea of ​​SENet is to model the dependencies between feature channels to improve the network's expressiveness and generalization capabilities. Its lightweight nature and small parameter size ensure improved network performance while maintaining manageable computing resources.

[0065] Specifically, for the input feature map, its spatial dimensions H×W are first aggregated through a squeeze operation to generate a channel feature descriptor. This process is achieved by applying global average pooling to each channel, thereby embedding global information into the descriptor. The calculation formula is as follows:

[0066]

[0067] Where z c Represents the global statistical features of the c-th channel.

[0068] The model uses a channel-dependent adaptive gating mechanism, namely the excitation operation, to fully utilize the information aggregated in the compressed block. This mechanism dynamically adjusts the importance of each channel by learning the activation value of a specific sample, thereby highlighting useful information features and suppressing redundant information.

[0069] The excitation operation uses two fully connected layers to perform nonlinear transformation on the compressed features, first performing channel dimensionality reduction and then dimensionality increase. The formula is defined as follows:

[0070] s=F ex (z,W)=σ(g(z,W))=σ(W2δ(W1z));

[0071] Where W1 and W2 represent the matrix parameters of dimensionality reduction and dimensionality increase, respectively, δ is the ReLU activation function, and σ is the Sigmoid activation function to generate the normalized weight vector s.

[0072] Finally, the weight vector s output by the excitation module is multiplied by the original feature map U channel by channel to obtain the enhanced feature recalibration:

[0073]

[0074] SENet improves the effectiveness of multi-scale feature fusion by adjusting the importance of feature channels at different scales and weighting features at each scale. The network learns the importance of each scale feature in different contexts and adaptively focuses on key areas (such as occlusion boundaries or complex geometric structures), improving the network's understanding of valid areas. At the same time, SENet reduces the weight of irrelevant background information or noise, suppressing its proportion in the feature representation, effectively optimizing the overall quality of feature extraction and laying a solid foundation for cost volume measurement and depth estimation in subsequent stages.

[0075] 3. Homography Transformation and Cost Volume Construction Module

[0076] The plane scanning method is introduced to use the front parallel planes of different depths as hypothetical planes to achieve efficient feature alignment and matching in stereo matching tasks. By mapping a differentiable homography transformation, the feature homography of the source view is mapped to the coordinates of the reference view, and feature volume modules can be constructed at multiple levels.

[0077] In detail, given the camera intrinsic parameter matrix K1 of the reference view and the extrinsic parameter parameters [R1, t1] and the camera intrinsic parameter matrix K of the source view i With external parameters [R i ,t i ], the homography transformation matrix can be defined as:

[0078]

[0079] Among them, H i (d) represents the homography matrix between the feature map of the i-th image and the reference feature map at depth d, where I is the identity matrix and n1 is the principal axis of the reference camera.

[0080] According to the calculation formula and based on the geometric projection relationship, the feature points in the source view are mapped to the same depth plane of the reference view to achieve consistent alignment of multi-view features.

[0081] In order to adapt to any number of input view images, a feature aggregation method based on variance measurement is adopted to aggregate multiple feature volumes. Aggregated into cost body modules, Represents the average cost volume metric:

[0082]

[0083] The size of the cost volume can be expressed as B×C×D×H×W, where B is the batch size, C is the number of feature channels, D is the number of depth hypotheses, and H and W are the height and width of the feature map, respectively.

[0084] The cost volume is used to measure the matching degree of each pixel in the reference image with its corresponding pixel in the source view, so as to determine whether the stereo matching effect is good.

[0085] 4. Cost Volume Regularization Module and Deep Inference

[0086] 3D convolution is used to regularize the cost volume, suppressing noise interference and preventing overfitting. The 3D UNet network structure is introduced to fuse the global semantic information extracted by the high-level network with the local information captured by the low-level network. This allows neighborhood information to be aggregated in a larger receptive field, achieving efficient feature expression.

[0087] Based on the cost volume regularization, the cost volume is further filtered. A softmax operation is performed to normalize the matching cost of each pixel under different depth assumptions to obtain its probability distribution at different depths.

[0088] In the first stage of depth inference (l=1), the depth hypothesis is in a predefined range d min ,d max ] uniformly sampled, the depth hypothesis of the m-th plane can be expressed as:

[0089] d(m)=d min +m(d max -d min ) / M,m∈[0,1,…,M-1];

[0090] Where M is the number of depth samples.

[0091] After the cost volume module passes through the regularization network to generate the probability volume module, the soft-argmin operation is applied to regress the initial depth map.

[0092] For each pixel p, its predicted depth value can be calculated:

[0093]

[0094] in, Represents the probability value of the depth hypothesis d(m) corresponding to the pixel point p in the first stage.

[0095] In order to further estimate a finer depth D in the first stage l , based on the previous depth map D l-1 Define a depth hypothesis with refined depth intervals:

[0096] d(m)=D l-1 ↑+mΔd p ,m∈{-M / 2,…,M / 2-1};

[0097] Among them D l-1 ↑ represents D upsampled by bicubic interpolation l-1 , Δd p This method can effectively preserve the spatial details of the depth map and provide accurate depth hypotheses for the next stage of fine depth estimation.

[0098] Project the pixel point p into the source view, and define the depth interval as the distance between two adjacent pixels projected onto the 3D ray along the epipolar line. This method effectively avoids the numerical instability problem caused by the projection points being too close in the image.

[0099] The depth update value of each pixel p in the next stage can be expressed as:

[0100]

[0101] 5. Loss Function

[0102] For deep regression, multiple layers of ground truth depth are constructed As a supervisory signal. The overall loss calculation uses L1 loss, which can effectively measure the absolute difference between the true depth value and the predicted depth value. The formula is as follows:

[0103]

[0104] where p valid Represents the set of valid true value pixels, Represents the true depth value of pixel p in the lth stage, Represents the estimated depth value of pixel p in the lth stage.

[0105] 6. 3D Reconstruction Experiment and Result Analysis

[0106] (1) Experimental implementation details

[0107] The experiment was trained on the DTU dataset, with the number of input images N set to 3 and the image resolution adjusted to 640×512 pixels. The network adopts a three-layer pyramid structure, with the depth assumptions of 48, 32, and 8 layers, respectively, and the corresponding depth intervals set to 4, 2, and 1. The model was trained on six NVIDIA TITAN RTX graphics cards (each with 24GB of video memory) with a batch size of 24. The Adam optimizer was used for parameter learning, with β1 of 0.9 and β2 of 0.999. The training process lasted for a total of 16 epochs, with an initial learning rate of 0.001, which was halved after the 10th, 12th, and 14th epochs, respectively.

[0108] After completing training and testing on the DTU dataset, the model was quantitatively and qualitatively evaluated. To further verify the model's generalization performance, generalization experiments were conducted on the Tanks & Temples dataset using the trained model without fine-tuning.

[0109] Ablation experiments are designed to verify the effectiveness of the proposed method, compare different feature fusion strategies, and further analyze the performance of the multi-scale feature fusion method proposed in this study.

[0110] (2) Benchmark dataset testing

[0111] First, the depth map is filtered based on the depth filtering method, setting the photometric consistency filter value to 0.8, the geometric consistency filter value to 0.01, the depth probability threshold to 0.8, the parallax consistency threshold to 0.2, and the minimum consistent view angle to 3.

[0112] The Gipuma method is used to fuse the filtered depth maps, making full use of multi-view stereo vision information and effectively improving the quality of 3D point cloud reconstruction.

[0113] like Figure 4 As shown in the figure, the experimental results show that the point cloud geometric outlines of each scene are clear, the texture and color rendering are good, and the overall reconstruction quality is high.

[0114] The 3D point clouds for scenes 1, 4, and 24 have dense structures and finely reconstructed details, meeting the requirements for 3D reconstruction in indoor scenes and demonstrating their strengths in processing regular structures and rich textures. However, some scenes still have limitations. For example, the multi-story building in scene 15 lacks some valid point clouds due to reduced available information from multiple viewpoints, indicating that the current method still has room for improvement in the reconstruction accuracy of complex structures.

[0115] (3) Comparative experiments and analysis

[0116] Quantitative evaluation on the DTU dataset shows that the average precision, average completeness and overall score of the proposed method are 0.354mm, 0.341mm and 0.347mm, respectively.

[0117] As shown in Table 1, the deep learning-based method can achieve a lower overall error, which fully demonstrates the advantages of the deep learning model in supporting multi-view stereo 3D reconstruction.

[0118] Table 1 Comparison of point cloud reconstruction results of different MVS methods on the DTU dataset

[0119]

[0120]

[0121] like Figure 5 Figure 2 shows the point cloud reconstruction results of different MVS methods for scenes 10, 29, and 32 of the DTU dataset. The experimental results demonstrate that more point clouds can be reconstructed in edge regions, effectively supplementing some detailed information and improving the completeness of the point cloud reconstruction. However, compared with the ground truth point cloud, the structure of the building area on the left side of scene 29 is still incomplete.

[0122] (4) Generalization Experiment and Analysis

[0123] The trained model was generalized and tested on the Tanks & Temples dataset, and the key parameters for depth filtering were set: the minimum number of consistent views was 6, the depth probability threshold was 0.8, and the disparity consistency threshold was 0.4.

[0124] Table 2 shows the experimental results for the Intermediate subset of the dataset. The experimental results demonstrate superior performance scores across multiple scenarios compared to most deep learning methods. The best scores were achieved on the Francis and Horse scenes, with suboptimal results achieved on the Lighthouse, M60, Panther, and Train scenes. Although the scores in other scenarios were close to those of CVP-MVSNet, this method slightly outperformed CVP-MVSNet in terms of average F-score (54.04). The multi-scale feature fusion strategy can fully integrate global information and local details, enhancing multi-scale feature representation capabilities and also demonstrating a positive impact on 3D reconstruction in larger-scale scenes.

[0125] The point cloud reconstructed on the Tanks&Temples dataset is visualized, and the effect is as follows Figure 6 As shown, this method is able to capture rich details in large outdoor scenes, clearly reconstructing the characters on the Train and Panther scenes. The geometric structure of each point cloud model is complete, the color mapping is excellent, and the overall visual presentation is clear and harmonious. Because some scenes contain a certain amount of noise points, which are mainly concentrated in areas with unstable depth information, the depth filter parameters are adjusted in these areas to filter out unnecessary points and improve the overall cleanliness of the model.

[0126] Figure 7The results of a set of scene graphs, confidence maps, depth maps, and depth filtering maps are shown. For each scene, the confidence map reflects the probability distribution of the depth estimate, where lighter areas indicate higher confidence in the depth estimate, while darker areas indicate greater uncertainty in the depth estimate. In the depth map, the presence of background areas and noise points can affect the quality of subsequent 3D reconstruction. Through the depth filtering operation, this redundant information is effectively removed. The filtered depth map can more clearly reflect the outline structure of objects in the scene, while retaining important geometric features and detailed information, enhancing the representation of key scene features.

[0127] (5) Ablation experiment and analysis

[0128] Table 3 shows the ablation test results of different feature fusion strategies on the DTU dataset. The experiments tested the following architectures: a 2D convolutional neural network, a feature pyramid FPN architecture, a cascaded encoder-decoder CEDNet architecture, a multi-scale feature fusion architecture, and a SENet module. The experimental results demonstrate that the proposed method exhibits excellent overall performance. Although its accuracy error of 0.354mm is higher than the FPN architecture's 0.324mm, its completeness score of 0.341mm is significantly better than FPN's 0.397mm, demonstrating its substantial advantages in scene coverage and geometric detail recovery.

[0129] At the same time, compared with the cascaded encoding-decoding structure, the multi-scale fusion strategy proposed in this method can make full use of the feature information between multiple scales, thereby improving the model's ability to express local information and global scenarios.

[0130] By introducing the channel attention mechanism of the SENet module, the model optimizes the feature fusion effect by dynamically adjusting multi-scale features, reducing the integrity error from 0.437mm to 0.341mm, enhancing the depth perception capability of MVS, and showing sufficient robustness in complex scenarios, ultimately improving the overall performance of point cloud model reconstruction.

[0131] Table 3 Ablation test of different feature fusion strategies on the DTU dataset

[0132]

[0133] 7. 3D Reconstruction Operation Efficiency Analysis

[0134] During the testing phase, we analyzed the depth inference time and video memory consumption for a single image. When testing depth estimation on the DTU dataset, the input image resolution was 1152×864 pixels. The average inference time for each depth map was 0.287 seconds, and the total video memory consumption was 4716MB. Table 4 compares the efficiency of this method with other deep learning methods. The proposed MVS method demonstrates strong competitiveness. Compared to cascaded networks, this method significantly reduces inference time and hardware resource requirements while maintaining high-quality reconstruction accuracy.

[0135] Table 4 Experimental results of the operating efficiency of different MVS methods

[0136]

[0137] Therefore, the present invention provides a multi-view stereo 3D reconstruction method based on multi-scale feature fusion. Through multi-scale feature fusion and bidirectional feature interaction, combined with a global attention mechanism, it effectively improves the accuracy and completeness of 3D reconstruction, efficiently integrates global and local information, and balances features of different scales, thereby improving the reconstruction quality.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-view stereo 3D reconstruction method based on multi-scale feature fusion, characterized in that: include: The method adopts a multi-scale feature fusion network architecture, which includes a feature extraction module, a cost volume construction module, a cost volume regularization network and a deep regression network.

2. The multi-view stereo 3D reconstruction method based on multi-scale feature fusion according to claim 1, characterized in that: The multi-view stereo 3D reconstruction method based on multi-scale feature fusion comprises the following steps: S1. Build a multi-view stereo network architecture based on multi-scale feature fusion, adopt a cascaded network architecture, use a coarse-to-fine depth prediction strategy, and generate a high-resolution depth map by iteratively refining the depth results; S2. Set up a feature extraction module, using a feature pyramid network (FPN) structure and a bidirectional feature fusion strategy to extract multi-level features from the input multi-view images. This allows the contextual information between features of different scales to interact and generate a multi-level feature map containing spatial geometry information and texture details. S3. Establish a homography transformation and cost volume construction module, define the homography transformation matrix based on the geometric projection relationship, map the source view features to the reference view coordinate system to form a feature volume, and use a feature aggregation method based on variance measurement to generate a cost volume; S4. Deploy the cost body regularization processing unit, use the 3D convolutional network to regularize the cost body, integrate high-level semantic information with low-level detail information, and improve feature expression capabilities; S5. Execute the depth inference and refinement process, regress the initial depth map through multi-stage depth hypothesis and soft-argmin operation, and perform fine depth estimation for each stage; S6. Conduct 3D reconstruction experiments and analyze the results, implement depth map fusion and 3D reconstruction, combine photometric consistency and geometric consistency constraints, and generate a dense 3D point cloud model.

3. The multi-view stereo 3D reconstruction method based on multi-scale feature fusion according to claim 1, characterized in that: The multi-view stereo network architecture for multi-scale feature fusion further includes a compression-excitation network SENet module.

4. The multi-view stereo 3D reconstruction method based on multi-scale feature fusion according to claim 2, characterized in that: In the homography transformation and cost volume construction module, the homography transformation matrix is ​​jointly calculated by camera parameters and depth assumptions.

5. The multi-view stereo 3D reconstruction method based on multi-scale feature fusion according to claim 2, characterized in that: The cost volume regularization processing unit adopts a 3DUNet architecture and an encoder-decoder structure for global-local feature fusion.

6. The multi-view stereo 3D reconstruction method based on multi-scale feature fusion according to claim 2, characterized in that: In the depth inference and refinement process, each stage generates a hypothesis set with adaptive depth intervals based on the previous depth map, and optimizes the depth regression results through a confidence weighting mechanism.

7. The multi-view stereo 3D reconstruction method based on multi-scale feature fusion according to claim 2, characterized in that: The method also includes the design of a multi-task loss function, which uses L1 depth loss, cross entropy classification loss and smoothing regularization term. The loss function is calculated as follows: Where p valid Represents the set of valid true value pixels, Represents the true depth value of pixel p in the lth stage, Represents the estimated depth value of pixel p in the lth stage.

8. The multi-view stereo 3D reconstruction method based on multi-scale feature fusion according to claim 2, characterized in that: The bidirectional feature fusion strategy includes a top-down feature upsampling path and a bottom-up feature downsampling path, and realizes cross-scale information complementarity through feature splicing and attention mechanism.

9. The multi-view stereo 3D reconstruction method based on multi-scale feature fusion according to claim 2, characterized in that: In S6, the 3D reconstruction experiment is trained on the DTU dataset, and the Adam optimizer is used in the training phase, combined with the cosine annealing learning rate scheduling strategy.

10. The multi-view stereo 3D reconstruction method based on multi-scale feature fusion according to claim 2, characterized in that: In S6, the three-dimensional point cloud model is processed by depth map fusion and hole filling, combined with normal vector estimation and mesh reconstruction algorithms to generate a three-dimensional model with a complete surface.

Citation Information

Cited By

  • Three-dimensional reconstruction method based on guided filtering and Mama geometric feature fusion

    CN120912796A

  • Depth estimation method, device and equipment based on binocular camera

    CN121280502A

  • Depth estimation method, device and equipment based on binocular camera

    CN121280502B

  • Three-dimensional reconstruction method and system based on frequency-space double-domain characteristics and cascade optimization

    CN122199838A