Multi-view three-dimensional reconstruction method of MVSNet network based on quadruple depth and wave type depth geometric enhancement
By introducing a quadruple depth and wave depth geometric enhancement MVSNet network into the multi-view 3D reconstruction method, using an incremental refinement architecture and efficient feature extractor, the limitations of existing methods for reconstruction in non-ideal scenarios and invisible areas are solved, and the reconstruction quality and efficiency are significantly improved.
Patent Information
- Application Number
- CN202510171749.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Existing learning-based multi-view 3D reconstruction methods have limitations in reconstructing non-ideal scenes and invisible areas, and complex network structures lead to prolonged runtime and increased memory burden, and error propagation problems also affect the reconstruction quality.
A multi-view three-dimensional reconstruction method for MVSNet network based on quadruple depth and wave depth geometric enhancement is proposed. The progressive refinement architecture is adopted to enrich the feature extractor through efficient multi-scale information and the visibility-aware view aggregation module based on external polar line transformers to improve the reconstruction quality, and ensure the efficient network running time through quadruple depth refinement and efficient Gaussian-Newtonian module depth optimization.
It significantly improves the accuracy and completeness of rebuilding point clouds, reduces memory consumption and runtime, enhances the reconstruction capability of challenging scenarios, and improves the reconstruction quality of invisible areas.
Smart Images

Figure CN120014174A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to multi-viewing in computer vision Figure 3 The present invention relates to a multi-view reconstruction technology based on a quadruple depth and wave-type depth geometric enhancement MVSNet network. Figure 3 Dimensional reconstruction method. Background Art
[0002] Multi-View Stereo (MVS), as a key basic task in the field of computer vision, has undergone extensive and in-depth research in the past decades. The MVS method aims to reconstruct the three-dimensional dense geometric structure of the target scene from multiple two-dimensional RGB images with known camera parameters. Its underlying logic is to build a pixel-level correspondence between the reference view and the source view. Humans have a more intuitive understanding of the three-dimensional world, which makes the MVS method play an indispensable and key role in many fields such as national defense construction, smart cities, smart transportation, cultural relics protection, disaster relief, agriculture and forestry, virtual / augmented reality, etc.
[0003] According to the different forms of expression of the reconstructed 3D scene, MVS methods can be divided into voxel-based, point diffusion-based, patch-based and depth map-based methods. Among them, the last category of depth map-based methods decouples the MVS task into two parts: depth prediction and depth fusion. It is suitable for large-scale scene reconstruction and is therefore favored for its flexibility and efficiency. Recently, some excellent tools and algorithms have emerged in this field, such as Gipuma, COLMAP, ACMM, etc.
[0004] However, the traditional MVS method based on depth maps has obvious shortcomings. When reconstructing target scenes with illumination changes, weakly textured or textureless areas, reflective surfaces or repetitive patterns (collectively referred to as challenging scenes), the traditional MVS method uses a manually designed similarity metric, which can easily lead to point cloud holes and blurring. At the same time, the engineered regularization method used by traditional MVS cannot eliminate the negative impact of invisible pixels (such as pixels in occluded areas and areas outside the field of view), thereby reducing the point cloud quality in invisible areas.
[0005] With the development of deep learning, more and more scholars have turned their attention to using convolutional neural networks (CNN) to solve the MVS problem. In 2018, Yao et al. first proposed an end-to-end trainable MVS network MVSNet to predict depth, and then used general technology for deep fusion. MVSNet has become the basic pipeline that most MVS networks actually follow, including four key steps: feature extraction using 2D CNN, cost aggregation including differentiable homography transformation steps, cost volume regularization using 3D CNN, and deep regression including soft-argmin steps. Among them, the first and third steps respectively utilize the powerful semantic information enrichment and noise information filtering capabilities of CNN, so that the learning-based MVS method significantly improves the accuracy and completeness of the 3D reconstruction results compared to the traditional MVS method. In addition, with the efficient computing power of the graphics processing unit (GPU) in computer graphics hardware, the reconstruction efficiency of the learning-based MVS method is greatly improved compared to the traditional method. Due to the above advantages, the learning-based MVS method has quickly become a multi-viewing Figure 3 It is a hot research topic in the field of dimensional reconstruction, and related research is committed to solving various problems existing in the MVSNet pipeline.
[0006] For example, in solving the problem of excessive memory and computational burden of 3D CNN, the research mainly develops from two directions. On the one hand, recursive methods (such as R-MVSNet, D2HC-RMVSNet, AA-RMVSNet, etc.) propose to use recursive convolutional neural networks to replace 3D CNN to reduce the dimensionality of three-dimensional processing to two dimensions. Although it comes at the cost of increased running time, it successfully breaks through the memory bottleneck in high-resolution reconstruction; on the other hand, multi-stage cascade methods represented by multi-stage cascade methods (such as CasMVSNet, CVP-MVSNet, UCS-MVSNet, etc.) reason about depth maps in a coarse-to-fine manner based on the idea of residual estimation. Although various cascade frameworks differ in the principles of determining the overall interval and sampling interval of the depth hypothesis in the refinement stage, they all significantly reduce the memory and time costs required by the learning-based MVS method. In addition, there are many targeted and representative research works in enhancing the difference and expressiveness of 2D features, promoting the cost aggregation of visibility perception, and designing different "probability-depth" derivation algorithms.
[0007] Nevertheless, the current learning-based MVS methods still face some challenges in achieving high-precision, high-completeness, and high-efficiency 3D reconstruction: (1) There are still obvious limitations when reconstructing non-ideal scenes and invisible areas; (2) Some complex network structures are introduced to improve prediction accuracy, which leads to longer running time and increased memory burden; (3) Most current methods adopt a multi-stage cascade framework to improve efficiency, but the error propagation problem inherent in this framework may have a potential negative impact on the reconstruction quality. Summary of the invention
[0008] In view of this, the present invention proposes a multi-view MVSNet network based on quadruple depth and wave-shaped depth geometric enhancement. Figure 3 The basic network process is four-fold depth rough estimation, four-fold depth refinement, and deep optimization based on efficient Gauss-Newton modules. This method aims to solve the problems of existing learning-based MVS methods, improve the reconstruction quality of deep learning MVS methods based on progressive refinement architecture from the perspective of deep fusion, and ensure the efficiency of network running time through careful design, thus providing a practical technical solution for achieving high-precision, high-integrity and high-efficiency 3D reconstruction.
[0009] A multi-view MVSNet network based on quadruple depth and wave-shaped deep geometric enhancement Figure 3 The reconstruction method comprises the following steps:
[0010] Step S1: Obtain a set of multi-view images of a scene, select one of them as a reference image each time and select (N-1) source images with the smallest image angle for it, and use these N images as input; construct an efficient multi-scale information-rich feature extractor, use the full-resolution reference image and its corresponding source image as input, and obtain two sets of different quarter-resolution scale feature maps and And a set of half-resolution scale feature maps They are used for quadruple depth coarse estimation, quadruple depth refinement, and depth optimization based on efficient Gauss-Newton modules;
[0011] Step S2: Through the differentiable homography transformation, a reference feature map is included and N-1 source feature maps N feature maps of Mapping onto the forward parallel planes in the reference camera frustum divided based on several depth assumptions, thereby generating N feature volumes accordingly; calculating the inter-group correlation similarity between the reference feature volume and the N-1 source feature volumes, and constructing N-1 similarity volumes; feeding the similarity volumes into the visibility-aware view aggregation module based on the epipolar transformer to generate a unified matching cost volume;
[0012] Step S3: using a double 3D CNN regularized cost volume to obtain a double two-channel probability volume; calculating a quadruple depth map and a quadruple reset confidence map from the double two-channel probability volume; using multiple loss functions to constrain the distribution of the quadruple depth map near the ground truth; using a wavy selection strategy, calculating a rough wavy depth map with wavy depth geometry from the quadruple depth map, and calculating a wavy confidence map corresponding to the rough wavy depth map from the quadruple reset confidence map, thereby completing a rough estimation of the quadruple depth;
[0013] Step S4: Construct a quadruple depth refinement module, determine the uniform initial depth range of the refinement module based on the rough wave-shaped depth map, and use N feature maps Repeat steps S2 and S3 to obtain a refined wave-shaped depth map and a confidence map;
[0014] Step S5: Using the efficient multi-scale information-rich feature extractor described in step S1 to extract the Direct calculation removes the extra feature extraction network used in the traditional Gauss-Newton module, thereby building an efficient Gauss-Newton module, which is used to upsample and optimize the quarter-resolution refined wave-shaped depth map to obtain a half-resolution final depth map;
[0015] Step S6: Replace the reference image and repeat the above steps until the depth maps corresponding to all images of the target scene are obtained.
[0016] Furthermore, the efficient multi-scale information-rich feature extractor described in step S1 can extract information-rich features, and the specific process is as follows:
[0017] First, multiple groups of convolutional layers are used to gradually collect semantic information from the input image, where each group of convolutional layers contains a convolution operation with a stride of 2 to reduce the resolution level; then, bilinear interpolation is used to upsample the low-resolution intermediate features containing semantic information, and these upsampled features are added to the original high-resolution features containing spatial information; finally, several convolutional layers are applied again to filter the features while changing the feature scale; through this iterative "top-down-bottom-up" processing, both high-level semantic information and low-level spatial information are gathered into the information-rich features output by the feature extraction network.
[0018] Furthermore, the efficient multi-scale information-rich feature extractor described in step S1 is highly efficient, and is specifically implemented as follows:
[0019] First, the multi-scale filter convolution block extracts features at different resolution levels for different subsequent stages, so the additional feature network originally required for the subsequent stage can be removed, which improves the computational efficiency and reduces the memory burden from the overall network reasoning level. The efficient multi-scale information-rich feature extractor itself also adopts an efficient image processing method. Most previous MVS methods used a feature pyramid network to process images of different views in sequence. The efficient multi-scale information-rich feature extractor uses packaging processing to merge different views of the same batch of data into the batch size, thereby extracting information-rich features of all input images at one time. This efficient image processing method reduces the network running time at the cost of a certain amount of memory consumption.
[0020] Furthermore, the visibility-aware view aggregation module based on the epipolar transformer described in step S2 calculates the visibility-aware weight by combining the cross-attention idea of the transformer with the filtering network. The specific process is as follows:
[0021] Reference feature As the query vector, the source feature The source feature volume W obtained by differentiable homography transformation i is regarded as a key vector, then assigned to the similarity volume S as a value vector i The visibility-aware weights of can be calculated by the cross-attention mechanism. The intermediate results are further filtered using 2D CNN to enhance robustness. The visibility-aware weights w i The calculation is as follows:
[0022]
[0023] Where i=1,…,N-1 represents the corresponding source view number, Φ represents the 2D CNN filter network, and C represents the number of feature channels. This view weight is also directly used for quadruple depth refinement.
[0024] Furthermore, the dual 3D CNN regularized cost volume described in step S3 is used to obtain a dual two-channel probability volume, and the specific process is as follows:
[0025] The cost volume is fed into the cost volume filtering network composed of dual 3D CNN. A convolution unit with an output channel of 2 is added at the end of each 3D CNN. The output dual two-channel volume is subjected to a softmax operation to obtain a dual two-channel probability volume.
[0026] Furthermore, the specific process of calculating the quadruple depth map and the quadruple reset confidence map from the dual two-channel probability volume described in step S3 is as follows:
[0027] First, we use deep regression to obtain the dual two-channel probability The dual two-channel depth map is calculated in That is, each two-channel depth map is calculated separately for all depth hypotheses d j The corresponding two-channel probability and:
[0028]
[0029] Where, d j (j=0,1,…,D-1, D is the total number of depth hypothesis planes) represents the depth hypothesis; then, by extracting the maximum and minimum values along the channel dimension, a quadruple depth map can be obtained from the dual two-channel depth map: and
[0030] Confidence is used to measure the quality of depth estimation and can be calculated as the sum of the probabilities of the four closest depth hypotheses. Therefore, the dual two-channel confidence map can be calculated from the dual two-channel probability body in a similar manner as above, and then the four reset confidence maps can be obtained from the dual two-channel confidence map.
[0031] Furthermore, the multiple loss functions described in step S3 are used to constrain the distribution of the quadruple depth map near the ground truth value, and the specific process is as follows:
[0032] First, the L1 loss function is used to constrain the four depth maps separately, making them smooth and close to the ground truth depth; then, a loss function is constructed to constrain each pair of depths: and ——Symmetrically distributed above and below the basic truth depth plane, that is, the basic truth depth D g lie in and Finally, the final wave depth map D 1 The sub-pixel accuracy is also constrained using the L1 loss function.
[0033] Furthermore, the wavy selection strategy described in step S3 is used to calculate a rough depth map with wavy depth geometry from the quadruple depth map, and a wavy confidence map corresponding to the wavy depth map is calculated from the quadruple reset confidence map. The specific process is as follows:
[0034] First, a coordinate mask image is calculated to select depth values in different depth maps at different coordinate positions of the coarse wavy depth map; then, the four intermediate value mean images and the overall mean image of the quadruple depth map are calculated. These five mean images are very close to the basic truth depth map; finally, the depth values are selected in turn from the quadruple depth map and the five mean images in a specific order, so that the depth in each wavy depth unit fluctuates slightly up and down along the basic truth depth surface, rather than all on one side of the basic truth depth surface, thus forming a coarse wavy depth map D^1; compared with the unilateral unit, the wavy depth unit has a smaller depth interpolation deviation in the depth fusion stage when the depth prediction deviation is the same, thereby generating a higher quality point cloud.
[0035] Furthermore, the specific process of determining the unified initial depth range of the refinement module based on the rough wave-shaped depth map in step S4 is as follows:
[0036] Calculate the maximum mean map and minimum mean map of the quadruple depth map respectively, calculate the difference between the maximum mean map and the minimum mean map, and then use the difference as the standard to extend a difference interval to both sides along the depth direction as the initial depth range. Since the quadruple depth map can be regarded as a repeated measurement of the rough depth prediction, the larger the difference between the maximum mean map and the minimum mean map, the less reliable the rough depth prediction is, so the corresponding initial depth range during refinement should be wider to ensure that the ground truth depth is still included in the initial depth range. In general, the initial depth range of the refinement module is smaller than the predefined depth range used for the rough estimate.
[0037] At this point, by filtering and fusing the final depth maps of all reference images, a dense point cloud of the scene corresponding to the set of multi-view images can be obtained, thus completing the multi-view Figure 3 Reconstruction.
[0038] The beneficial effects of the present invention are as follows: Based on various problems faced by the prior art, the present invention proposes a multi-view MVSNet network based on quadruple depth and wave-shaped depth geometric enhancement. Figure 3dimensional reconstruction method. The network model of the method has a progressive refinement architecture, which effectively avoids the error propagation problem inherent in the cascade architecture. The basic process is four-fold depth rough estimation, four-fold depth refinement, and depth optimization based on an efficient Gauss-Newton module. First, the present invention constructs an efficient multi-scale information-rich feature extractor, and the information-rich features extracted with a shorter running time help the network predict more accurate depth maps for challenging scenes. Secondly, the present invention constructs a visibility-aware view aggregation module based on an epipolar transformer, which uses the cross-attention idea and the robust visibility-aware weights calculated by the filtering network to help improve the reconstruction quality of invisible areas. Then, the present invention proposes to predict a four-fold depth map and calculate a coarse depth map with a wavy geometric structure from it, and this geometric structure reduces the interpolation bias in the depth fusion stage. Further, the present invention constructs a four-fold depth refinement module, determines a unified initial depth range based on the coarse wavy depth map, and realizes the refinement of the coarse wavy depth map. The above complete four-fold depth and wavy depth geometry correlation design significantly improves the accuracy and integrity of the reconstructed point cloud. Finally, the present invention constructs an efficient Gauss-Newton module to optimize the depth map and improve the output resolution, which effectively improves the reconstruction quality. The multi-view MVSNet network based on quadruple depth and wave-shaped deep geometric enhancement proposed in the present invention is trained and tested on the DTU dataset. Figure 3 The 3D reconstruction method can significantly improve the quality of the reconstructed point cloud. Despite the time-consuming dual 3D CNN structure, it can still complete the depth map prediction at a very fast running speed overall. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A multi-view MVSNet network based on quadruple depth and wave-shaped deep geometric enhancement Figure 3 Flowchart of the reconstruction method;
[0040] Figure 2 A multi-view MVSNet network based on quadruple depth and wave-shaped deep geometric enhancement Figure 3 The point cloud visualization result diagram of all scenes in the DTU dataset evaluation set reconstructed by the 3D reconstruction method;
[0041] Figure 3 Comparison of point cloud visualization results of different methods for reconstructing challenging scenes in the DTU dataset;
[0042] Figure 4 Comparison of GPU memory consumption, running time and overall reconstruction quality of different methods;
[0043] Figure 5 Ablation experiment diagram for quadruple depth coarse estimation, quadruple depth refinement, and depth optimization based on efficient Gauss-Newton module. DETAILED DESCRIPTION
[0044] In order to more clearly illustrate the intended objectives, technical means and effects of the present invention, the implementation methods of the present invention will be described in detail below in conjunction with the accompanying drawings and specific examples.
[0045] It should be noted that, unless otherwise specifically noted, all technical and scientific terms used in the present invention have the same common meanings as those in the technical field to which the present invention belongs. Experimental conditions not specifically mentioned in the following examples are all based on conventional technical means in the field. It should be understood that the described embodiments are only partial examples of the present invention, not all examples. Based on these embodiments, all other embodiments obtained by ordinary technicians in the field without creative work belong to the protection scope of the present invention.
[0046] Example 1
[0047] In order to solve the problems existing in the prior art, the reconstruction quality of the deep learning MVS method based on the progressive refinement architecture is improved from the perspective of deep fusion, the efficiency of the network running time is ensured through careful design, and a feasible technical solution is provided to achieve high-precision, high-integrity and high-efficiency 3D reconstruction. The embodiment of the present invention proposes a multi-view MVSNet network based on quadruple depth and wave-type deep geometric enhancement. Figure 3 The network model of this method has a progressive refinement architecture, and its basic process is four-fold depth rough estimation, four-fold depth refinement, and deep optimization based on efficient Gauss-Newton modules, such as Figure 1 The multi-view MVSNet network based on quadruple depth and wave-shaped deep geometric enhancement Figure 3 The specific steps of the reconstruction method are as follows:
[0048] Step S1: Obtain a set of multi-view images of a scene, select one of them as a reference image each time and select (N-1) source images with the smallest image angle for it, and use these N images as input; construct an efficient multi-scale information-rich feature extractor, use the full-resolution reference image and its corresponding source image as input, and obtain two sets of different quarter-resolution scale feature maps and And a set of half-resolution scale feature maps They are respectively used for quadruple depth coarse estimation, quadruple depth refinement, and depth optimization based on the efficient Gauss-Newton module.
[0049] Step S1.1: Obtain a set of multi-view images of a certain scene, select one of them as a reference image each time, and select (N-1) source images with the smallest image angle for it, and use the N images as input.
[0050] For example, when using the DTU public dataset, a set of multi-view images of a scene is selected from its training set or evaluation set for network training or evaluation. According to the image angle score predefined in its matching record file, the (N-1) source images with the best matching degree can be selected for each image in turn. In this embodiment, the value is: N=3.
[0051] Step S1.2: Construct an efficient multi-scale information-rich feature extractor, taking the full-resolution reference image and its corresponding source image as input, and obtain two sets of different quarter-resolution scale feature maps and And a set of half-resolution scale feature maps They are respectively used for quadruple depth coarse estimation, quadruple depth refinement, and depth optimization based on the efficient Gauss-Newton module.
[0052] First, multiple groups of convolutional layers are used to gradually collect semantic information from the input image, where each group of convolutional layers contains a convolution operation with a stride of 2 to reduce the resolution level; then, bilinear interpolation is used to upsample the low-resolution intermediate features containing semantic information, and these upsampled features are added to the original high-resolution features containing spatial information; finally, several convolutional layers are applied again to filter the features while changing the feature scale; through this iterative "top-down-bottom-up" processing, both high-level semantic information and low-level spatial information are gathered into the information-rich features output by the feature extraction network.
[0053] The multi-scale filter convolution block extracts features at different resolution levels for different subsequent stages, so the additional feature network originally required in the subsequent stages can be removed, which improves the computational efficiency and reduces the memory burden from the overall network reasoning level. The efficient multi-scale information-rich feature extractor itself also uses an efficient image processing method. Most previous MVS methods used a feature pyramid network to process images of different views in sequence. The efficient multi-scale information-rich feature extractor uses packaging processing to merge different views of the same batch of data into the batch size, thereby extracting information-rich features of all input images at one time. This efficient image processing method reduces network running time at the cost of a certain amount of memory consumption.
[0054] This efficient multi-scale information-rich feature extractor extracts information from the full-resolution input image. In this paper, we extract features of different scales. and They are used for four-fold depth rough estimation, four-fold depth refinement, and depth optimization based on efficient Gauss-Newton modules. In this embodiment, the values are: for the DTU training set, W×H=640×512; for the DTU evaluation set, W×H=1600×1152; the number of feature channels C=32.
[0055] Step S2: Through the differentiable homography transformation, a reference feature map is included and N-1 source feature maps N feature maps of Mapped onto the forward parallel plane divided based on several depth assumptions in the reference camera frustum, thereby generating N feature volumes accordingly; calculating the inter-group correlation similarity between the reference feature volume and the N-1 source feature volumes, and constructing N-1 similarity volumes; feeding the similarity volumes into the visibility-aware view aggregation module based on the epipolar transformer to generate a unified matching cost volume.
[0056] Step S2.1: Through the differentiable homography transformation, a reference feature map is included and N-1 source feature maps N feature maps of Mapped onto the forward parallel plane in the reference camera cone divided based on several depth assumptions, thereby generating N feature volumes accordingly.
[0057] The reference feature volume W0 can be obtained by simply replicating the reference feature along the depth dimension Get, and the source feature body W i Source features need to be Perform a differentiable homography transformation, referred to as a differentiable twist. Specifically, given the corresponding camera parameters It can be the pixel p in the reference view r Calculate its corresponding source view pixel p s :
[0058]
[0059] Where, d j (j=0,1,…,D-1, D is the total number of depth hypothesis planes) represents the depth hypothesis, that is, the forward parallel planes divided based on several depth hypotheses. min ,d max ) is obtained by uniformly sampling the inverse depth range. j :
[0060] 1 / d j =1 / d min -j / (D-1)(1 / d min -1 / d max )
[0061] In this embodiment, the relevant values are: For the DTU data set, (d min ,d max )=(425mm,935mm); for training, D=48; for evaluation, D=96.
[0062] Use the calculated sampling grid coordinates to map the source feature map Perform differentiable bilinear sampling to twist the source feature map to several depth hypothesis planes to obtain the source feature volume W i .
[0063] Step S2.2: Calculate the inter-group correlation similarity between the reference feature body and the N-1 source feature bodies, and construct N-1 similarity volumes.
[0064] In the field of MVS based on deep learning, feature maps are first transformed into feature volumes or similarity volumes, and then the cost volume is calculated.
[0065] By evenly dividing the feature channels C into G groups, the inter-group correlation similarity S between the feature bodies can be calculated i for:
[0066] S i (d j )=G / C <W0,W i (d j )>
[0067] In the formula, <·,·> is an inner product operation. In this embodiment, the value is: G=4.
[0068] Step S2.3: Feed the similarity volume into the epipolar transformer-based visibility-aware view aggregation module to generate a unified matching cost volume.
[0069] In the transformer, the original cross-attention mechanism between the query vector Q, key vector K, and value vector V is as follows:
[0070]
[0071] Where C represents the number of channels of Q and K. Similarly, the reference feature Treat it as a query vector and match it along the epipolar line to the key vector W i , the similarity volume S used to enhance the value vector can be calculated through the cross attention mechanism i Here, 2D CNN is used to further filter the intermediate results to enhance robustness. Finally, the visibility-aware weight w i The calculation is as follows:
[0072]
[0073] Where i=1,…,N-1 represents the corresponding source view number, Φ represents the 2D CNN filter network, and C represents the number of feature channels. This view weight is also directly used for quadruple depth refinement.
[0074] The unified matching cost volume C is calculated by weighted sum:
[0075]
[0076] Where S i is the similarity volume, and role is a vector of values.
[0077] Step S3: Using a dual 3D CNN regularized cost body to obtain a dual two-channel probability body; calculating a quadruple depth map and a quadruple reset confidence map from the dual two-channel probability body; using multiple loss functions to constrain the distribution of the quadruple depth map near the ground truth; using a wavy selection strategy, a rough wavy depth map with wavy depth geometry is calculated from the quadruple depth map, and a wavy confidence map corresponding to the rough wavy depth map is calculated from the quadruple reset confidence map, thereby completing a rough estimate of the quadruple depth.
[0078] Step S3.1: Use the dual 3D CNN regularized cost volume to obtain a dual two-channel probability volume.
[0079] The cost volume is fed into the cost volume filtering network composed of dual U-Net-based multi-scale 3D CNN, and the cost volume is regularized using the powerful noise information filtering ability of CNN. A convolution unit with an output channel of 2 is added to the end of each 3D CNN, and the output dual two-channel volume is subjected to a softmax operation to obtain the dual two-channel probability volume P.
[0080] Step S3.2: Compute a quadruple depth map and a quadruple reset confidence map from the dual two-channel probability volume.
[0081] In the field of MVS based on deep learning, the calculation of the depth map from the probability volume is performed by deep regression:
[0082]
[0083] On this basis, the present invention realizes its dual-channel input calculation, and adds three symbols Q, a and b to represent the dual channels.
[0084] First, we use deep regression to obtain the dual two-channel probability The dual two-channel depth map is calculated in That is, each two-channel depth map is calculated separately for all depth hypotheses d j The corresponding two-channel probability and:
[0085]
[0086] Where, d j(j=0,1,…,D-1, D is the total number of depth hypothesis planes) represents the depth hypothesis; then, by extracting the maximum and minimum values along the channel dimension, a quadruple depth map can be obtained from the dual two-channel depth map: and
[0087] Confidence is used to measure the quality of depth estimation and can be calculated as the sum of the probabilities of the four closest depth hypotheses. Therefore, the dual two-channel confidence map can be calculated from the dual two-channel probability body in a similar manner as above, and then the four reset confidence maps can be obtained from the dual two-channel confidence map.
[0088] Step S3.3: Use multiple loss functions to constrain the distribution of the quadruple depth map around the ground truth.
[0089] First, the L1 loss function is used to constrain the four depth maps separately, making them smooth and close to the ground truth depth; then, a loss function is constructed to constrain each pair of depths: and ——Symmetrically distributed above and below the basic truth depth plane, that is, the basic truth depth D g lie in and between:
[0090]
[0091] When the base truth depth D g Not located in and In between, Must be greater than Therefore, the L1 loss function makes Increase until the basic truth depth D g lie in and Between, at this time Less than After that, the L1 loss function promotes Reduce but keep the ground truth depth D g lie in and In between, even if and Approaching the ground truth depth D from both sides g .
[0092] Finally, the final wave depth map D 1 The sub-pixel accuracy is also constrained using the l1 loss function.
[0093] Step S3.4: adopt the wave selection strategy, calculate the rough wavy depth map with wavy depth geometry from the quadruple depth map, and calculate the wavy confidence map corresponding to the rough wavy depth map from the quadruple reset confidence map, thereby completing the rough estimation of the quadruple depth.
[0094] First, a coordinate mask map is calculated to select depth values in different depth maps at different coordinate positions of the rough wavy depth map; then, four intermediate value mean maps and the overall mean map of the quadruple depth map are calculated. These five mean maps are very close to the basic truth depth map; finally, depth values are selected from the quadruple depth map and the five mean maps in turn in a specific order, so that the depth in each wavy depth unit fluctuates slightly up and down along the basic truth depth surface, rather than all on one side of the basic truth depth surface, thereby forming a rough wavy depth map D 1 Compared with the unilateral unit, the wavy depth unit has a smaller depth interpolation deviation in the depth fusion stage when the depth prediction deviation is the same, thus generating a higher quality point cloud.
[0095] Step S4: Construct a quadruple depth refinement module, determine the uniform initial depth range of the refinement module based on the rough wave-shaped depth map, and use N feature maps Repeat steps S2 and S3 to obtain a refined wavy depth map and a confidence map.
[0096] First, the maximum mean map and minimum mean map of the quadruple depth map are calculated respectively, and the difference between the maximum mean map and the minimum mean map is calculated. Then, using the difference as the standard, a difference interval is extended to both sides along the depth direction as the initial depth range. Since the quadruple depth map can be regarded as a repeated measurement of the rough depth prediction, the larger the difference between the maximum mean map and the minimum mean map, the less reliable the rough depth prediction is. Therefore, the corresponding initial depth range during refinement should be wider to ensure that the ground truth depth is still included in the initial depth range. In general, the initial depth range of the refinement module is wider than the predefined depth range (d min ,d max ) should be smaller.
[0097] After that, given the initial depth range and N feature maps Repeat steps S2 and S3 to obtain the refined wavy depth map D 2 and confidence graph.
[0098] Step S5: Using the efficient multi-scale information-rich feature extractor described in step S1 to extract the Direct calculation removes the extra feature extraction network used in the traditional Gauss-Newton module, thereby constructing an efficient Gauss-Newton module, which is used to upsample and optimize the quarter-resolution refined wavy depth map to obtain a half-resolution final depth map.
[0099] The traditional Gauss-Newton module uses the Gauss-Newton iterative optimization algorithm, the goal of which is to minimize the error function at pixel p:
[0100]
[0101] In the formula, p i′ represents the reprojected pixel of p. and They represent the half-resolution reference features and source features extracted by the feature pyramid, respectively. In the error minimization process, the optimal depth value of pixel p is determined, thereby achieving depth map optimization. Obviously, the use of information-rich features can enhance the difference between source features and reference features, making the final depth value obtained by the iterative optimization process more accurate.
[0102] Therefore, by using the efficient multi-scale information enrichment feature extractor in step S1 to extract On the one hand, these information-rich features can be used to promote the depth optimization process. On the other hand, the feature pyramid used in the traditional Gauss-Newton module can be omitted, thereby reducing the memory burden and running time. An efficient Gauss-Newton module is constructed. This efficient Gauss-Newton module refines the wavy depth map D at a quarter resolution. 2 Upsample and optimize to obtain the final half-resolution depth map D 3 .
[0103] Step S6: Replace the reference image and repeat the above steps until the depth maps corresponding to all images of the target scene are obtained.
[0104] The supervised learning strategy is used to train the network model using the loss described in step S3 and the L1 loss between the effective value of the final depth map and the effective value of the ground truth depth map. Next, the trained network model is used to evaluate each scene and predict the corresponding depth map of all its reference images. Through depth map filtering and fusion, a dense point cloud is generated to achieve multi-view Figure 3 Multi-dimensional reconstruction. Multi-view MVSNet network based on quadruple depth and wavy deep geometric enhancement Figure 3 The point cloud visualization results of the whole scene reconstruction of the DTU dataset evaluation set are shown in the figure. Figure 2 Shown
[0105] As described above, the embodiment of the present invention adopts a multi-view MVSNet network based on quadruple depth and wave-shaped depth geometric enhancement. Figure 3 A 3D reconstruction method is proposed to reconstruct a dense 3D point cloud of a target scene from multiple 2D RGB images with known camera parameters.
[0106] Example 2
[0107] Experimental setup details
[0108] The DTU dataset is a large-scale indoor multi-view stereo vision dataset containing 124 different scenes. Each scene is scanned from 49 or 64 viewpoints under seven different lighting conditions. The dataset was collected in a controlled laboratory environment and the camera trajectory remained fixed. To ensure a fair comparison, we follow the base truth depth map and dataset partitioning used by previous methods. The official MATLAB evaluation code is used to calculate the distance indicators of the point cloud, namely the average accuracy (accuracy Acc. for short) and the average completeness (Comp. for short), in millimeters. The quality of the reconstructed point cloud is measured by the overall Overall, also known as the overall quality or overall reconstruction quality, and the calculation formula is:
[0109] Overall=(Acc.+Comp.) / 2
[0110] The experimental software environment used in the embodiments of the present invention is configured as follows: the operating system is Ubuntu 20.04; the GPU is NVIDIA GeForce RTX 3090; the deep learning framework uses PyTorch 1.8.1, CUDA 11.7 and cuDNN 8.6.0.
[0111] The specific experimental settings of the embodiment of the present invention are as follows: the total number of training cycles (epochs) is 16; the batch size (batch_size) is 4; the optimizer is Adam; the initial learning rate is 0.001, and at the 8th, 10th and 12th epochs, the learning rate decays to half of the original.
[0112] Comparative experiments with other network models
[0113] First, the proposed method is quantitatively compared with the traditional MVS method, the learning-based progressive refinement MVS method, and the learning-based cascade MVS method (see Table 1). The results show that the proposed method outperforms other learning-based methods in terms of completeness and overall quality, especially the progressive refinement methods with similar architectures. Specifically, compared with the recent state-of-the-art learning-based progressive refinement MVS method, ACINR-MVSNet, the proposed method performs 24.18% relatively better in terms of completeness and 8.96% relatively better in terms of overall reconstruction quality.
[0114] Table 1 Reconstruction quality comparison experimental results (the lower the distance index means the better the reconstruction quality)
[0115]
[0116]
[0117] In order to verify the robustness of this algorithm, three sets of "challenging scenes" in the DTU dataset, Scan24, Scan75 and Scan77, were selected and reconstructed using R-MVSNet, ACINR-MVSNet and the method proposed in this invention, respectively. A qualitative comparison was performed. The results are shown in the figure. Figure 3 As shown in Figure 2. The reconstruction results of R-MVSNet are obtained by optimizing the complete point cloud post-processing steps designed by it, while most other methods (including the method proposed in this paper) do not include these steps; ACINR-MVSNet is one of the most advanced learning-based progressive refinement MVS methods. Figure 3 It can be seen that since the proposed method uses quadruple depth prediction and wavy depth map, R-MVSNet and ACINR-MVSNet reconstruct more complete 3D point clouds in the above challenging scenarios. This finding is also supported by the comparison results of the completeness indicators in Table 1.
[0118] As shown in Table 2, the GPU memory and running time required by different learning-based MVS methods to predict a depth map are compared. The table also lists the quantitative evaluation indicators of the reconstruction quality of each method, as well as the network configuration or experimental settings that have a direct impact on the efficiency, including the resolution of the predicted depth map and the number of depth hypothesis planes.
[0119] Table 2 Efficiency comparison experimental results
[0120]
[0121] As shown in Table 2 and Figure 4 As shown in the figure, although the proposed method uses a time-consuming dual 3D CNN network to predict the quadruple depth, the application of efficient multi-scale information-rich feature extractors and efficient Gauss-Newton modules can still complete the depth map prediction at a faster running speed overall, and is even competitive compared to the classic cascade method known for its high efficiency. Specifically, compared with MVSNet, the depth map reconstructed by the proposed method is almost 9 times larger, only 2GB of memory is used more, the running time is relatively reduced by 63.81%, and the reconstruction quality performance is relatively improved by 33.98%.
[0122] The above experimental results verify that the MVSNet network based on quadruple depth and wave-shaped depth geometric enhancement proposed in this paper is effective in multi-view Figure 3This method not only significantly improves the quality of the reconstructed point cloud, but also takes up less memory and has faster inference speed than other progressive refinement methods.
[0123] Ablation experiment
[0124] In the ablation experiment, a network model with 3 training views was used for unified evaluation, and the number of evaluation views was 5. In order to verify the effectiveness of the quadruple depth and wavy geometry, quadruple depth refinement module and efficient Gauss-Newton module proposed in the present invention in the method, the efficient Gauss-Newton module, the quadruple depth refinement module, and the quadruple depth enhancement MVS module were gradually removed from the network of the method proposed in the present invention, and finally the ordinary MVS module was used to replace the quadruple depth enhancement MVS module for ablation experiments. In order to make a fair comparison, the depth maps predicted by different ablated network models were upsampled to the same resolution as the input image before depth fusion. The quantitative comparison results are shown in Table 3. With the addition of the quadruple depth enhancement MVS module, the quadruple depth refinement module and the efficient Gauss-Newton module to the method, the reconstructed point cloud has obvious performance improvements in terms of completeness and overallness.
[0125] Table 3 Ablation experiment results
[0126]
[0127] The qualitative comparison results are as follows Figure 5 As shown in Figure 2. With the addition of the quadruple depth enhanced MVS module, quadruple depth refinement module, and efficient Gauss-Newton module to our method, the depth map becomes more accurate, the point cloud becomes denser, and captures finer details. These improvements further demonstrate the effectiveness of our quadruple depth design, quadruple depth optimization module, and efficient Gauss-Newton optimization module.
[0128] in conclusion
[0129] In summary, this paper proposes a multi-view MVSNet network based on quadruple depth and wave-shaped depth geometric enhancement. Figure 3dimensional reconstruction method. The network model of the method has a progressive refinement architecture, which effectively avoids the error propagation problem inherent in the cascade architecture. Its basic process includes four-fold depth rough estimation, four-fold depth refinement, and depth optimization based on an efficient Gauss-Newton module. First, the present invention designs an efficient multi-scale information-rich feature extractor that can extract information-rich features in a shorter running time, thereby helping the network to more accurately predict the depth map in challenging scenes. Secondly, the present invention introduces a visibility-aware view aggregation module based on an epipolar transformer, which significantly improves the reconstruction quality of invisible areas by using a cross-attention mechanism and robust visibility-aware weights calculated by a filtering network. Then, the present invention proposes to predict a four-fold depth map and calculate a coarse depth map with a wavy geometric structure from it, which reduces the interpolation bias in the depth fusion stage. Further, the present invention constructs a four-fold depth refinement module to refine the coarse wavy depth map by determining a unified initial depth range, thereby significantly improving the accuracy and completeness of the reconstructed point cloud. The above complete four-fold depth and wavy depth geometry-related design significantly improves the accuracy and completeness of the reconstructed point cloud. Finally, the present invention develops an efficient Gauss-Newton module to optimize the depth map and improve the output resolution, thereby effectively improving the reconstruction quality. The multi-view MVSNet network based on quadruple depth and wave-shaped deep geometric enhancement proposed in the present invention is trained and tested on the DTU dataset. Figure 3 The 3D reconstruction method can significantly improve the quality of the reconstructed point cloud. Although this method uses a time-consuming dual 3D CNN structure, it can still complete the depth map prediction at a faster speed overall.
[0130] It is easy for those skilled in the art to understand that, under the premise of no conflict, the above-mentioned preferred embodiments can be freely combined and superimposed. The above content is a further detailed description of the present invention, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the protection scope of the present invention.
Claims
1. A multi-view 3D reconstruction method based on MVSNet network with quadruple depth and wave-shaped depth geometric enhancement, characterized in that: The following steps are involved: Step S1: Obtain a set of multi-view images of a scene, select one of them as a reference image each time and select (N-1) source images with the smallest image angle for it, and use these N images as input; construct an efficient multi-scale information-rich feature extractor, use the full-resolution reference image and its corresponding source image as input, and obtain two sets of different quarter-resolution scale feature maps and And a set of half-resolution scale feature maps They are used for quadruple depth coarse estimation, quadruple depth refinement, and depth optimization based on efficient Gauss-Newton modules; Step S2: Through the differentiable homography transformation, a reference feature map is included and N-1 source feature maps N feature maps of Mapping onto the forward parallel planes in the reference camera frustum divided based on several depth assumptions, thereby generating N feature volumes accordingly; calculating the inter-group correlation similarity between the reference feature volume and the N-1 source feature volumes, and constructing N-1 similarity volumes; feeding the similarity volumes into the visibility-aware view aggregation module based on the epipolar transformer to generate a unified matching cost volume; Step S3: using a double 3D CNN regularized cost volume to obtain a double two-channel probability volume; calculating a quadruple depth map and a quadruple reset confidence map from the double two-channel probability volume; using multiple loss functions to constrain the distribution of the quadruple depth map near the ground truth; using a wavy selection strategy, calculating a rough wavy depth map with wavy depth geometry from the quadruple depth map, and calculating a wavy confidence map corresponding to the rough wavy depth map from the quadruple reset confidence map, thereby completing a rough estimation of the quadruple depth; Step S4: Construct a quadruple depth refinement module, determine the uniform initial depth range of the refinement module based on the rough wave-shaped depth map, and use N feature maps Repeat steps S2 and S3 to obtain a refined wave-shaped depth map and a confidence map; Step S5: Using the efficient multi-scale information-rich feature extractor described in step S1 to extract the Direct calculation removes the extra feature extraction network used in the traditional Gauss-Newton module, thereby building an efficient Gauss-Newton module, which is used to upsample and optimize the quarter-resolution refined wave-shaped depth map to obtain a half-resolution final depth map; Step S6: Replace the reference image and repeat the above steps until the depth maps corresponding to all images of the target scene are obtained.
2. According to claim 1, a multi-view 3D reconstruction method based on the MVSNet network with quadruple depth and wave-shaped depth geometric enhancement is characterized in that: The efficient multi-scale information-rich feature extractor described in step S1 can extract information-rich features, and the specific process is as follows: First, multiple groups of convolutional layers are used to gradually collect semantic information from the input image, where each group of convolutional layers contains a convolution operation with a stride of 2 to reduce the resolution level; then, bilinear interpolation is used to upsample the low-resolution intermediate features containing semantic information, and these upsampled features are added to the original high-resolution features containing spatial information; finally, several convolutional layers are applied again to filter features while changing the feature scale; through this iterative "top-down-bottom-up" process, both high-level semantic information and low-level spatial information are gathered into the information-rich features output by the feature extraction network.
3. The multi-view 3D reconstruction method based on the MVSNet network with quadruple depth and wave-shaped depth geometric enhancement according to claim 1, characterized in that: The visibility-aware view aggregation module based on the epipolar transformer described in step S2 calculates the visibility-aware weight by combining the cross-attention idea of the transformer with the filtering network. The specific process is as follows: Reference feature As the query vector, the source feature The source feature volume W obtained by differentiable homography transformation i is regarded as a key vector, then assigned to the similarity volume S as a value vector i The visibility-aware weight of can be calculated by the cross-attention mechanism; the intermediate results are further filtered using 2D CNN to enhance robustness; the visibility-aware weight w i The calculation is as follows: Where i=1,…,N-1 represents the corresponding source view number, Φ represents the 2D CNN filter network, and C represents the number of feature channels; the view weight is also directly used for quadruple depth refinement.
4. The multi-view 3D reconstruction method based on the MVSNet network with quadruple depth and wave-shaped depth geometric enhancement according to claim 1, characterized in that: The dual 3D CNN regularized cost volume described in step S3 is used to obtain a dual two-channel probability volume, and the specific process is as follows: The cost volume is fed into the cost volume filtering network composed of dual 3D CNN. A convolution unit with an output channel of 2 is added at the end of each 3D CNN. The output dual two-channel volume is subjected to a softmax operation to obtain a dual two-channel probability volume.
5. The multi-view 3D reconstruction method based on the MVSNet network with quadruple depth and wave-shaped depth geometric enhancement according to claim 1, characterized in that: The specific process of calculating the quadruple depth map and the quadruple reset confidence map from the dual two-channel probability volume described in step S3 is as follows: First, we use deep regression to obtain the dual two-channel probability The dual two-channel depth map is calculated in That is, each two-channel depth map is calculated separately for all depth hypotheses d j The corresponding two-channel probability and: Where, d j (j=0,1,…,D-1, D is the total number of depth hypothesis planes) represents the depth hypothesis; then, by extracting the maximum and minimum values along the channel dimension, a quadruple depth map can be obtained from the dual two-channel depth map: and Confidence is used to measure the quality of depth estimation and can be calculated as the sum of the probabilities of the four closest depth hypotheses. Therefore, the dual two-channel confidence map can be calculated from the dual two-channel probability body in a similar manner as above, and then the four reset confidence maps can be obtained from the dual two-channel confidence map.
6. The multi-view 3D reconstruction method based on the MVSNet network with quadruple depth and wave-type depth geometric enhancement according to claim 1, characterized in that: The multiple loss functions described in step S3 are used to constrain the distribution of the quadruple depth map near the ground truth value. The specific process is as follows: First, the L1 loss function is used to constrain the four depth maps separately, making them smooth and close to the ground truth depth; then, a loss function is constructed to constrain each pair of depths: and ——Symmetrically distributed above and below the basic truth depth plane, that is, the basic truth depth D g lie in and Finally, for the rough wavy depth map D 1 The sub-pixel accuracy is also constrained using the L1 loss function.
7. The multi-view 3D reconstruction method based on the MVSNet network with quadruple depth and wave-shaped depth geometric enhancement according to claim 1, characterized in that: The wavy selection strategy described in step S3 is used to calculate a rough wavy depth map with wavy depth geometry from the quadruple depth map, and a wavy confidence map corresponding to the rough wavy depth map is calculated from the quadruple reset confidence map, thereby completing the rough estimation of the quadruple depth. The specific process is as follows: First, a coordinate mask map is calculated to select depth values in different depth maps at different coordinate positions of the rough wavy depth map; then, four intermediate value mean maps and the overall mean map of the quadruple depth map are calculated. These five mean maps are very close to the basic truth depth map; finally, depth values are selected from the quadruple depth map and the five mean maps in turn in a specific order, so that the depth in each wavy depth unit fluctuates slightly up and down along the basic truth depth surface, rather than all on one side of the basic truth depth surface, thereby forming a rough wavy depth map D 1 ; Compared with the unilateral unit, the wavy depth unit has a smaller depth interpolation deviation in the depth fusion stage when the depth prediction deviation is the same, thus generating a higher quality point cloud.
8. The multi-view 3D reconstruction method based on the MVSNet network with quadruple depth and wave-shaped depth geometric enhancement according to claim 1, characterized in that: The specific process of determining the unified initial depth range of the refinement module based on the rough wave-shaped depth map in step S4 is as follows: The maximum mean map and the minimum mean map of the quadruple depth map are calculated respectively, and the difference between the maximum mean map and the minimum mean map is calculated. Then, using the difference as a standard, a difference interval is extended to both sides along the depth direction as the initial depth range.
Citation Information
Patent Citations
A Deep Learning-Based 3D Reconstruction Method for UAV Aerial Imagery
CN111462329B
Image processing method of multi-view stereoscopic reconstruction network model MA-MVSNet based on multi-resolution self-adaptability
CN114937073A
Three-dimensional reconstruction method and system based on improved MVSNet
CN116912405A