Multi-view three-dimensional reconstruction method of MVSNet network based on epipolar line converter and implicit neural optimization enhancement

By introducing external pole line transformers and implicit neural optimization modules into the MVSNet network, combined with the enhanced Gaussian-Newtonian module, the problem of difficult to balance reconstruction efficiency and quality in the existing technology is solved, and a three-dimensional reconstruction with high precision, high integrity and high efficiency is achieved.

CN119991966AActive Publication Date: 2025-05-13BEIHANG UNIV

Patent Information

Application Number
CN202510171881.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-13
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The existing multi-view stereoscopic vision three-dimensional reconstruction method based on deep learning is difficult to balance between reconstruction efficiency and quality, and there are limitations in reconstruction of non-ideal scenarios and invisible areas, and complex network structures lead to long run time and heavy memory burden.

Method used

The MVSNet network based on external polar line transformer and implicit neural optimization enhancement is adopted to enhance the coarse depth estimation of the MVS module, the depth super-resolution of the implicit neural optimization module and the deep refinement of the Gaussian-Newtonian module through external polar line transformer to achieve a progressive refinement architecture and avoid error propagation in the cascading architecture.

Benefits of technology

It significantly improves the quality of predicted depth maps and reconstructed point clouds, reduces memory usage, and makes inference faster than other progressive refinement methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991966A_ABST
    Figure CN119991966A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view three-dimensional reconstruction method of an MVSNet network based on an epipolar line converter and implicit neural optimization enhancement, and belongs to the technical field of multi-view three-dimensional reconstruction in computer vision. The basic process of the network model is depth rough estimation based on an epipolar converter enhanced MVS module, depth super-resolution based on an implicit neural optimization module, and depth refinement based on an enhanced Gaussian-Newton module. According to the method provided by the invention, various problems caused by balance loss between reconstruction efficiency and reconstruction quality in the prior art are solved, the quality of a predicted depth map and the quality of a reconstructed point cloud can be remarkably improved, meanwhile, less memory is occupied, and compared with other learning-based progressive refinement MVS methods, the reasoning speed is higher, and the reasoning efficiency is improved. Therefore, an effective technical scheme is provided for high-precision, high-integrity and high-efficiency three-dimensional reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to multi-view in computer vision Figure 3 The technical field of multi-dimensional reconstruction, more specifically, relates to a multi-view MVSNet network based on an epipolar transformer and implicit neural optimization enhancement. Figure 3 Dimensional reconstruction method. Background Art

[0002] As a basic task in the field of computer vision, Multi-View Stereo (MVS) has been widely studied for decades. Its core idea is to establish pixel-level correspondences between different views. The MVS method aims to reconstruct the three-dimensional dense geometric structure of the target scene from multiple two-dimensional RGB images with known camera parameters. Since humans are naturally more inclined to perceive the three-dimensional world, the MVS method plays a vital role in many fields such as national defense security, disaster relief, agriculture and forestry, urban construction, autonomous navigation, archaeological research, art design, virtual / augmented reality, etc. According to different scene representations, MVS methods can be divided into voxel-based, point diffusion-based, patch-based and depth map-based methods. The depth map-based method is the most portable and accurate, and is applicable to a wide range of scene scales. Among them, outstanding methods such as COLMAP, Gipuma, and ACMM have emerged.

[0003] However, the traditional MVS method uses a hand-designed similarity metric, which is difficult to properly solve challenging problems such as illumination changes, weak texture areas, reflective surfaces, and repeated patterns in the target scene, which can easily cause point cloud holes and blurs. On the other hand, the engineered regularization method used by the traditional MVS cannot eliminate the negative impact of invisible pixels, including pixels located in the occluded area between views and outside the field of view, resulting in a decrease in the point cloud quality in the invisible area.

[0004] After entering the era of deep learning, Yao et al. took advantage of the powerful semantic information enrichment and noise information filtering capabilities of convolutional neural networks (CNNs), and first proposed an end-to-end trainable network MVSNet for MVS tasks in 2018. MVSNet uses 2D CNN to extract image features as a similarity metric and feeds the matching cost volume into 3D CNN for regularization, which greatly reduces the adverse effects of the above-mentioned challenging scenes and invisible areas. By utilizing the efficient computing power of the graphics processing unit (GPU) in the computer's graphics card hardware, MVSNet achieves significantly better reconstruction quality than traditional state-of-the-art MVS methods with higher reconstruction efficiency. Therefore, the MVS method based on deep learning has quickly become a multi-viewing Figure 3The research focus in the field of 3D reconstruction is overwhelming, and most of them use MVSNet as the actual network baseline, and are committed to improving various problems existing in the MVSNet basic pipeline. For example, in response to the heavy memory and computational burden of 3D CNN, on the one hand, recursive methods represented by R-MVSNet, D2HC-RMVSNet, AA-RMVSNet, etc. propose to use recursive convolutional units to regularize the 2D slices of the 3D cost volume in order, breaking the memory bottleneck in high-resolution reconstruction at the expense of running time; on the other hand, multi-stage cascade methods represented by CasMVSNet, CVP-MVSNet, UCS-MVSNet, etc. are based on the idea of ​​residual estimation to infer depth maps in a coarse-to-fine manner. Although the various cascade frameworks differ in the principles of determining the overall interval and sampling interval of the depth hypothesis in the refinement stage, they all greatly reduce the memory and time required for learning-based MVS methods.

[0005] However, for high-precision, high-completeness, and high-efficiency 3D reconstruction, current learning-based MVS methods still have some problems: (1) There are still significant limitations when reconstructing non-ideal scenes and invisible areas; (2) Some complex network structures introduced to improve prediction accuracy lead to long running times and heavy memory burdens; (3) Most current methods adopt a multi-stage cascade framework from an efficiency perspective, and the errors in the first stage can be propagated to the last stage through residual estimation, thereby deteriorating the final prediction results. Summary of the invention

[0006] In view of this, the present invention proposes a multi-view MVSNet network based on epipolar transformer and implicit neural optimization enhancement. Figure 3 The basic network process is a rough depth estimation based on the epipolar transformer-enhanced MVS module, a depth super-resolution based on the implicit neural optimization module, and a depth refinement based on the enhanced Gauss-Newton module. This method solves various problems caused by the imbalance between reconstruction efficiency and reconstruction quality in the existing technology, thereby providing an effective technical solution for high-precision, high-integrity and high-efficiency 3D reconstruction.

[0007] A multi-view MVSNet network based on epipolar transformer and implicit neural optimization enhancement Figure 3 The reconstruction method comprises the following steps:

[0008] Step S1: Obtain a set of multi-view images of the target scene, select one image from them each time as a reference image and select (N-1) source images with the smallest image angle for them, and use the N images as input; input the full-resolution reference image and the source image into a scene-aware feature extractor embedded with an epipolar transformer to obtain a quarter-resolution feature map of the reference image and the source image, i.e., a reference feature map and a source feature map;

[0009] Step S2: Map N-1 source feature maps to forward parallel planes divided by several depth hypotheses in the reference camera cone through differentiable homography transformation; calculate the inter-group correlation similarity between the reference feature map and the twisted source feature map to form N-1 similarity volumes; aggregate the similarity volumes into a unified matching cost volume through a view aggregation module embedded with an inter-view epipolar transformer; use 3D CNN to regularize the cost volume to obtain a probability volume along the depth direction; and then calculate a quarter-resolution depth map from the probability volume through a depth regression operation;

[0010] Step S3: construct an implicit neural optimization module to implement the depth map super-resolution process, taking quarter-resolution pixel coordinates, half-resolution pixel coordinates, quarter-resolution depth map, and full-resolution reference image as input, and outputting an optimized half-resolution depth map;

[0011] Step S4: Using the context-aware feature extractor described in step S1 to replace the feature pyramid used in the traditional Gauss-Newton module, thereby constructing an enhanced Gauss-Newton module; using this module to further refine the half-resolution depth map to obtain a final depth map, the size of which is still half resolution;

[0012] Step S5: Replace the reference image and repeat the above steps until the depth maps corresponding to all images of the target scene are obtained.

[0013] Furthermore, the context-aware feature extractor described in step S1 is constructed by embedding a view extrema transformer in a feature pyramid-based cyclic feature extractor: the feature pyramid-based cyclic feature extractor has a "top-down-bottom-up" cyclic reciprocating structure, which can extract context-aware features that contain both low-level spatial information and high-level semantic information; the view extrema transformer has a multi-head self-attention layer and a multi-head cross-attention layer, which can aggregate non-local features along the extrema, thereby further enhancing the distinguishability and describability of the features.

[0014] Furthermore, the similarity volumes are aggregated into a unified matching cost volume through the view aggregation module embedded with the inter-view epipolar line transformer described in step S2, and the process is as follows:

[0015] First, through the inter-view epipolar transformer, based on the cross-attention idea of ​​the transformer, a long-distance 3D correlation relationship along the depth direction is constructed between the matching pixels on the epipolar lines in different views, and the weight w of each similarity volume is calculated. i :

[0016]

[0017] Where F0 and W i They represent the reference feature and the source feature volume twisted according to the differentiable homography transformation, i=1,…,N-1 represents the corresponding source view number, Ch represents the number of feature channels; in the cross-attention mechanism, W0 plays the role of the query vector, W i The role of is the key vector, and the two are calculated as above. i , which can be used to weight the value vector; the unified matching cost volume C is calculated as follows:

[0018]

[0019] Where S i is the similarity volume, and its role is the value vector.

[0020] Furthermore, the implicit neural optimization process described in step S3 is as follows:

[0021] The implicit neural function f implemented by a multilayer perceptron θ , we can jointly estimate the interpolated depth value v q,p and the interpolation weight w q,p :

[0022] w q,p ,v q,p =f θ (d p ,g p ,g q -g p ,x q -x p )

[0023] In the formula, x p and x q denote quarter and half resolution pixel coordinates, respectively, g p and g q denote the quarter-resolution and half-resolution reference feature maps extracted from the full-resolution reference image using the feature pyramid, d p Represents the use of feature pyramids from a quarter-resolution coarse depth map D 1 The extracted quarter-resolution depth feature map; finally, the optimized half-resolution depth map D 2 The calculation is done by interpolation as follows:

[0024]

[0025] Where D 2 (x q ) means D 2 The mid-coordinate x q The depth value, represents the set of neighboring pixels of pixel q in the half-resolution image domain in the quarter-resolution image domain.

[0026] At this point, the final depth map of all reference images of the target scene or object is filtered and fused to obtain a dense point cloud, completing the multi-view Figure 3 Reconstruction.

[0027] The beneficial effects of the present invention are as follows: Based on the various problems faced by the existing multi-view stereoscopic vision 3D reconstruction technology based on deep learning, the present invention proposes a multi-view stereoscopic vision 3D reconstruction technology based on the epipolar transformer and the MVSNet network enhanced by implicit neural optimization. Figure 3 dimensional reconstruction method. The network model of the method has a progressive refinement architecture, and the basic process is a rough depth estimation based on the epipolar transformer-enhanced MVS module, a depth super-resolution based on the implicit neural optimization module, and a depth refinement based on the enhanced Gauss-Newton module, thereby avoiding the error propagation inherent in the cascade architecture. First, the present invention constructs a context-aware feature extractor with an embedded epipolar transformer between views, and the more expressive features extracted enable the network to more accurately reason about depth values ​​in challenging areas. Secondly, the present invention constructs a view aggregation module with an embedded inter-view epipolar transformer, which uses the view weights inferred by the cross-attention idea to improve the point cloud reconstruction effect of invisible areas. The MVS module enhanced by the above two epipolar transformers can effectively improve the quality of depth map and point cloud reconstruction. Then, the present invention constructs an implicit neural optimization module, which uses implicit neural functions to achieve coarse depth map super-resolution, and effectively improves the accuracy of the edge area of ​​foreground objects in the depth map and point cloud. Finally, the present invention constructs an enhanced Gauss-Newton module to further refine the depth map as a whole, thereby comprehensively improving the reconstruction quality. The multi-view MVSNet network based on epipolar transformer and implicit neural optimization enhancement is experimentally verified by training and testing on the DTU dataset. Figure 3 The dimensional reconstruction method can significantly improve the quality of predicted depth maps and reconstructed point clouds, while taking up less memory and having a faster reasoning speed than other progressive refinement methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 Multi-view based MVSNet network enhanced by epipolar transformer and implicit neural optimization Figure 3 Flowchart of the reconstruction method;

[0029] Figure 2 It is a schematic diagram of the principle of the intra-view epipolar line converter and the inter-view epipolar line converter;

[0030] Figure 3 Multi-view based MVSNet network enhanced by epipolar transformer and implicit neural optimization Figure 3 The point cloud visualization result diagram of all scenes in the DTU dataset evaluation set reconstructed by the 3D reconstruction method;

[0031] Figure 4 Comparison of point cloud visualization results of different methods for reconstructing challenging scenes in the DTU dataset;

[0032] Figure 5 Comparison of GPU memory consumption, running time and overall reconstruction quality of different methods;

[0033] Figure 6 Ablation experiment diagrams of two epipolar transformers, implicit neural optimization module and enhanced Gauss-Newton module. DETAILED DESCRIPTION

[0034] In order to make the intended purpose, technical means and effects of the present invention more clearly understood, the specific implementation methods of the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments.

[0035] It should be noted that, unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as in the technical field to which the present invention belongs. In the following examples, the methods for which the experimental conditions are not specifically indicated are all conventional technical means in the technical field. It should be understood that the described embodiments are only some examples of the present invention, not all examples. Based on these embodiments in the present invention, all other embodiments obtained by ordinary technicians in the field without paying creative work belong to the protection scope of the present invention.

[0036] Example 1

[0037] In order to solve various problems caused by the imbalance between reconstruction efficiency and reconstruction quality in the prior art and to provide an effective technical solution for high-precision, high-integrity and high-efficiency 3D reconstruction, an embodiment of the present invention proposes a multi-view MVSNet network based on an epipolar transformer and implicit neural optimization enhancement. Figure 3 The basic process of the network is the depth rough estimation based on the epipolar transformer enhanced MVS module, the depth super-resolution based on the implicit neural optimization module, and the depth refinement based on the enhanced Gauss-Newton module, such as Figure 1 The multi-view MVSNet network based on epipolar transformer and implicit neural optimization is shown in Figure 2. Figure 3 The reconstruction method specifically comprises the following steps:

[0038] Step S1: Obtain a set of multi-view images of the target scene, select one image each time as a reference image and select (N-1) source images with the smallest image angle for it, and use the N images as input; input the full-resolution reference image and the source image into a scene-aware feature extractor embedded with an intra-view epipolar transformer to obtain a quarter-resolution feature map of the reference image and the source image, i.e., a reference feature map and a source feature map.

[0039] Step S1.1: Obtain a set of multi-view images of the target scene, select one of them as a reference image each time, and select (N-1) source images with the smallest image angle for it, and use the N images as input.

[0040] For example, a DTU public dataset is selected. For different processes of network training and evaluation, multi-view images of a scene can be selected from the corresponding training set and evaluation set. According to the image angle scores defined in pair.txt, the (N-1) source images with the best matching degree can be selected for each reference image. In this embodiment, the value is: N = 5.

[0041] Step S1.2: Input the full-resolution reference image and the source image into a context-aware feature extractor embedded with an intra-view epipolar transformer to obtain a quarter-resolution feature map of the reference image and the source image, i.e., the reference feature map and the source feature map.

[0042] The context-aware feature extractor is constructed by embedding the view epipolar transformer in the feature pyramid-based cyclic feature extractor: the feature pyramid-based cyclic feature extractor has a "top-down-bottom-up" cyclic reciprocating structure, which can extract context-aware features that contain both low-level spatial information and high-level semantic information; the view epipolar transformer has a multi-head self-attention layer and a multi-head cross-attention layer, which can aggregate non-local features along the epipolar lines, such as Figure 2 As shown in Figure 2, the distinguishability and descriptiveness of the features are further enhanced. The context-aware feature extractor extracts the reference image I0 and the source image I0 of size W×H. In this embodiment, the reference features and source features of size (W / 4)×(H / 4) are extracted. In this embodiment, the values ​​are: for the DTU training set, W×H=640×512; for the DTU evaluation set, W×H=1600×1152.

[0043] Step S2: Map N-1 source feature maps to forward parallel planes in the reference camera cone divided by several depth hypotheses through differentiable homography transformation; calculate the inter-group correlation similarity between the reference feature map and the twisted source feature map to form N-1 similarity volumes; aggregate the similarity volumes into a unified matching cost volume through a view aggregation module embedded with an inter-view epipolar transformer; use 3D CNN to regularize the cost volume to obtain a probability volume along the depth direction; and then calculate a quarter-resolution depth map from the probability volume through a depth regression operation.

[0044] Step S2.1: Map the N feature maps to the forward parallel planes in the reference camera frustum divided by several depth hypotheses through differentiable homography transformation.

[0045] The forward parallel plane divided by several depth hypotheses, namely the depth hypothesis plane, is obtained by min ,d max ) uniform sampling depth assumption d j get:

[0046] d j =d min +j / D(d max -d min )

[0047] Where j = 0, 1, ..., D-1, D is the total number of depth hypothesis planes. In this embodiment, the value is: for the DTU dataset, (d min ,d max )=(425mm,935mm); for training, D=48; for evaluation, D=96.

[0048] The differentiable homography actually describes the potential pixel-level correspondence between the reference view and the source view. Specifically, given the corresponding camera parameters It can be pixel p in the reference view, that is, the depth hypothesis plane d in the reference camera frustum after twisting j Calculate the corresponding source view pixel p i (d j ):

[0049]

[0050] Then, by using the calculated sampling grid coordinates, differentiable bilinear sampling is performed on the source feature map, for example, using the F.grid_sample function in the Python language, the deformation operation of the source feature map can be completed.

[0051] Step S2.2: Calculate the inter-group correlation similarity between the reference feature map and the twisted source feature map to form N-1 similarity volumes.

[0052] In the field of MVS based on deep learning, feature maps are first transformed into feature volumes or similarity volumes, and then the cost volume is calculated.

[0053] By combining the reference feature map F0 and the twisted source feature map W i (d j )’s feature channels Ch are evenly divided into G groups, and the similarity between the g-th groups can be calculated as:

[0054]

[0055] In the formula, g = 0, 1, ..., G-1, <·,·> is the inner product operation. By calculating the similarity of G groups, N-1 similarity graphs with G channels are obtained; by calculating the similarity graphs at all depth hypothesis planes, N-1 similarity volumes S are obtained. i In this embodiment, the value is: G=4.

[0056] Step S2.3: Aggregate the similarity volumes into a unified matching cost volume through a view aggregation module embedded with an inter-view epipolar transformer.

[0057] The similarity volumes are aggregated into a unified matching cost volume through a view aggregation module embedded with an inter-view epipolar transformer. The process is as follows:

[0058] First, if Figure 2 As shown in the figure, through the inter-view epipolar transformer, based on the cross-attention idea of ​​the transformer, a long-distance 3D correlation relationship along the depth direction is constructed between the matching pixels on the epipolar lines in different views, and the weight w of each similarity volume is calculated. i :

[0059]

[0060] Where F0 and W i They represent the reference feature and the source feature volume twisted according to the differentiable homography transformation, i=1,…,N-1 represents the corresponding source view number, Ch represents the number of feature channels; in the cross-attention mechanism, W0 plays the role of the query vector, W i The role of is the key vector, and the two are calculated as above. i , which can be used to weight the value vector; the unified matching cost volume C is calculated as follows:

[0061]

[0062] Where S i is the similarity volume, and its role is the value vector.

[0063] Step S2.4: Use 3D CNN to regularize the cost volume to obtain a probability volume along the depth direction.

[0064] The cost volume is fed into a multi-scale 3D CNN based on U-Net, and the cost volume is regularized using the powerful noise information filtering capability of CNN. A convolution unit with an output channel of 1 is set at the end of the 3D CNN, and the output volume is subjected to a softmax operation, which is the probability volume P along the depth direction.

[0065] Step S2.5: A quarter-resolution depth map is calculated from the probability volume through a depth regression operation.

[0066] Corresponding to the depth hypothesis sampling, the probability weighted sum of all depth hypotheses is calculated, and a coarse depth map with a quarter resolution can be calculated from the probability volume. The formula is as follows:

[0067]

[0068] Where p is the corresponding pixel and the rough depth map D 1 Resolution and Features Figure 1 It is one quarter of the full-resolution input image, that is, its size is (W / 4)×(H / 4).

[0069] Step S3: construct an implicit neural optimization module to implement the depth map super-resolution process, taking quarter-resolution pixel coordinates, half-resolution pixel coordinates, quarter-resolution depth map, and full-resolution reference image as input, and outputting an optimized half-resolution depth map.

[0070] First, a quarter-resolution coarse depth map D is extracted using a feature pyramid 1 The quarter-resolution feature d p , extract the quarter-resolution features g of the full-resolution reference image p and its half-resolution feature g q The purpose is to p and g q Under the guidance, from d p Start by getting the depth value of the half-resolution depth map.

[0071] The depth map is considered as a two-dimensional image with a single channel. In the implicit neural representation, all images can be represented by the same implicit neural function f with parameters θ. θ express:

[0072] s=f θ (c,x)

[0073] Where s is the estimated signal value (depth value or RGB value), c is the potential code (usually a feature vector), and x is the two-dimensional coordinate in the image domain. Based on this, we can extend the definition of the depth value v corresponding to the quarter-resolution pixel coordinate in the half-resolution depth map: q,p for:

[0074] v q,p =f θ (d p ,g p ,x q -x p )

[0075] In the formula, x p and x q Represents quarter and half resolution pixel coordinates respectively. According to the edge weight calculation method in the graph attention mechanism, if the depth value v is used q,p The depth value at the half-resolution pixel coordinate is interpolated with the weight:

[0076] w q,p =f η (g p ,g q -g p )

[0077] Since the interpolated depth value and the interpolated weight have the same representation, the implicit neural function f can be defined as θ To jointly estimate the interpolated depth value v q,p and the interpolation weight w q,p :

[0078] w q,p ,v q,p =f θ (d p ,g p ,g q -g p ,x q -x p )

[0079] The implicit neural function can be implemented by a multilayer perceptron (MLP). In this embodiment, in order to improve efficiency, the number of hidden layers of the MLP is set to [256, 128, 64, 32].

[0080] Finally, the optimized half-resolution depth map D 2 The calculation is done by interpolation as follows:

[0081]

[0082] Where D 2 (x q ) means D 2 The mid-coordinate x q The depth value, represents the set of neighboring pixels of pixel q in the half-resolution image domain in the quarter-resolution image domain.

[0083] Step S4: Use the context-aware feature extractor described in step S1 to replace the feature pyramid used in the traditional Gauss-Newton module to construct an enhanced Gauss-Newton module; use this module to further refine the half-resolution depth map to obtain a final depth map, which is still half-resolution in size.

[0084] The traditional Gauss-Newton module can minimize the error function at pixel p through an iterative optimization algorithm:

[0085]

[0086] In the formula, p i′ represents the reprojected pixel of p. and Denote the half-resolution reference features and source features extracted from the full-resolution reference image using the feature pyramid, respectively. In this error minimization process, the optimal depth value of pixel p is determined. Obviously, using more expressive features can enhance the recognition of the difference between the source features and the reference features, thereby improving the accuracy of the final depth value obtained through the Gauss-Newton iterative optimization process.

[0087] Therefore, by using the context-aware feature extractor described in step S1 to replace the feature pyramid used in the traditional Gauss-Newton module, that is, using enhanced context-aware features to replace the original ordinary semantic features, an enhanced Gauss-Newton module can be constructed to extract the optimized half-resolution depth map D 2 Further refinement is performed to obtain the final half-resolution depth map D 3 .

[0088] Step S5: Replace the reference image and repeat the above steps until the depth maps corresponding to all images of the target scene are obtained.

[0089] The network model is trained using a supervised learning strategy using the L1 loss between the effective value of the predicted depth map and the effective value of the ground truth depth map. The trained network model is then used to evaluate each scene and predict the corresponding depth map of all its reference images. By filtering and fusing the final depth map of all reference images of the target scene or object, a dense point cloud can be generated to complete the multi-view Figure 3 Multi-view reconstruction using MVSNet network based on epipolar transformer and implicit neural optimization enhancement Figure 3The point cloud visualization results of the DTU dataset evaluation set reconstructed by the 3D reconstruction method are shown in the figure below. Figure 3 Shown

[0090] As described above, the embodiment of the present invention uses a multi-view MVSNet network based on an epipolar transformer and implicit neural optimization enhancement. Figure 3 A 3D reconstruction method is proposed to reconstruct the 3D dense geometric structure of the target scene from multiple 2D RGB images with known camera parameters.

[0091] Example 2

[0092] Experimental setup details

[0093] The DTU dataset is a large-scale indoor multi-view stereo vision dataset collected under controlled laboratory conditions with a fixed camera trajectory. The dataset contains 124 different scenes, each scanned from 49 or 64 viewpoints under 7 different lighting conditions. To ensure a fair comparison, the same scheme as MVSNet is used to obtain the ground truth depth map and the corresponding mask image, and the DTU dataset is divided into training, validation, and evaluation sets following the previous method. The official MATLAB evaluation code is provided to calculate the quantitative evaluation indicators of the point cloud, namely the average accuracy (referred to as accuracy) and the average completeness (referred to as completeness), in mm. Accuracy Acc. measures the distance from the reconstructed point cloud to the ground truth point cloud; completeness Comp. measures the distance from the ground truth point cloud to the reconstructed point cloud. The quality of the reconstructed point cloud is measured by the overall Overall, also known as the overall quality or overall reconstruction quality, calculated as the average of accuracy and completeness:

[0094] Overall=(Acc.+Comp.) / 2

[0095] The experimental software environment configuration used in the embodiment of the present invention is: Ubuntu 20.04 is selected as the operating system; NVIDIA GeForce RTX 3090 is selected as the GPU; in the deep learning framework, PyTorch uses version 1.8.1, CUDA uses version 11.7, and cuDNN uses version 8.6.0.

[0096] The specific experimental settings of the embodiment of the present invention are as follows: the total number of training cycles (epochs) is 16; the batch size (batch_size) is 4; the RMSProp (root mean square propagation) optimizer is used; the initial learning rate is 0.00025, and the learning rate is decayed by multiplying it by 0.9 every two training cycles.

[0097] Comparative experiments with other network models

[0098] First, the method proposed in the present invention is quantitatively compared with the traditional MVS method, the learning-based progressive optimization MVS method and the learning-based cascade MVS method. As shown in Table 1, the method proposed in the present invention is superior to other learning-based methods in terms of completeness and overallness, especially in the progressive optimization methods with the same architecture.

[0099] Table 1 Reconstruction quality comparison experimental results (the lower the distance index means the better the reconstruction quality)

[0100]

[0101]

[0102] In order to verify the robustness of this algorithm, three sets of challenging scenes in the DTU dataset, Scan11, Scan24, and Scan77, were selected and reconstructed qualitatively using different methods, R-MVSNet, ACINR-MVSNet, and the method proposed in this invention. The results are shown in Figure 2. Figure 4 As shown in Figure 2. The reconstruction results of R-MVSNet are obtained by optimizing the complete point cloud post-processing steps designed by it. Most other methods (including the method proposed in this invention) do not contain these steps. ACINR-MVSNet is the best recent learning-based progressive refinement MVS method. Figure 4 It can be seen from the figure that the proposed method embeds two epipolar transformers and uses an implicit neural optimization module and an enhanced Gauss-Newton module to reason about the depth map in a progressively refined architecture, so it reconstructs a more complete 3D point cloud for the above scene than R-MVSNet and ACINR-MVSNet. This finding is also supported by the comparison results of the corresponding integrity indicators in Table 1.

[0103] In practical applications, the efficiency of the learning-based MVS method that people are concerned about mainly refers to its efficiency in the test and evaluation stage. As shown in Table 2, the GPU memory and running time required to predict a depth map by different progressive refinement methods are compared. At the same time, the table lists the quantitative evaluation indicators of the reconstruction quality of each method, as well as the network configuration or experimental settings that have a direct impact on the efficiency of each method, including the input image resolution and the number of depth hypothesis planes.

[0104] Table 2 Efficiency comparison experimental results

[0105]

[0106] Given that MVSNet and Transformer are computationally expensive in terms of time and memory, the proposed method is carefully designed at each stage of the network to reduce the computational requirements. Specifically, by utilizing the most basic epipolar geometry principle in the MVS algorithm, the two epipolar transformers restrict the non-local feature aggregation and long-range 3D association relationship construction to the corresponding points on the epipolar lines of the reference view and the source view, rather than between all pixels in the view, such as Figure 2 In the implicit neural optimization module, fewer hidden layers are configured to build a lightweight implicit neural function decoder. As shown in Table 2 and Figure 5 As shown in the figure, the method proposed in the present invention is superior to the three representative progressive refinement methods in terms of efficiency, occupies less memory and is significantly faster in inferring a single depth map.

[0107] Compared with MVSNet and VA-Point-MVSNet, the proposed method achieved a relative reduction of 51.4% and 84.8% in running time on the DTU evaluation set, and achieved a relative performance improvement of 33.3% and 14.2% in the overall reconstruction of point clouds. The above experiments verify the multi-view MVSNet network based on the epipolar transformer and implicit neural optimization enhancement proposed in the present invention. Figure 3 The dimensional reconstruction method can significantly improve the quality of the reconstructed point cloud while taking up less memory and having a faster reasoning speed than other progressive refinement methods.

[0108] Ablation experiment

[0109] In the ablation experiment, a network model with 3 training views was used for unified evaluation, and the number of evaluation views was 5. In order to verify the effectiveness of the epipolar transformer enhanced MVS module, implicit neural optimization module and enhanced Gauss-Newton module constructed by the present invention in the method, the enhanced Gauss-Newton module and implicit neural optimization module were gradually removed from the network of the method proposed by the present invention, and finally the epipolar transformer enhanced MVS module was replaced with an ordinary MVS module to conduct an ablation experiment. In order to make a fair comparison, the depth maps predicted by different ablated network models were upsampled to the same resolution as the input image before depth fusion. The quantitative comparison results are shown in Table 3. With the addition of the epipolar transformer enhanced MVS module, the implicit neural optimization module and the enhanced Gauss-Newton module to the method, the reconstructed point cloud has obvious performance improvements in terms of completeness and overallness.

[0110] Table 3 Ablation experiment results

[0111]

[0112] The qualitative comparison results are as follows Figure 6As shown. A scenario-aware feature extractor with an embedded inter-view epipolar transformer is constructed, and the more expressive features extracted by it enable the network to reason about more accurate depth values, which helps in more complete reconstruction, especially in challenging areas such as reflective surfaces; a view aggregation module with an embedded inter-view epipolar transformer is constructed, which uses the view weights inferred by the cross-attention idea to achieve more accurate depth estimation in invisible areas and improve the point cloud reconstruction effect in the corresponding area. The constructed implicit neural optimization module generates clearer boundaries in the depth map and restores finer details in the point cloud. The constructed enhanced Gauss-Newton module further refines the depth map as a whole, thereby comprehensively improving the reconstruction quality. These components together improve the overall performance of the method proposed in the present invention.

[0113] in conclusion

[0114] In summary, the present invention proposes a multi-view MVSNet network based on epipolar transformer and implicit neural optimization enhancement. Figure 3 dimensional reconstruction method. The network model of the method has a progressive refinement architecture, and the basic process is a rough depth estimation based on the epipolar transformer-enhanced MVS module, a depth super-resolution based on the implicit neural optimization module, and a depth refinement based on the enhanced Gauss-Newton module, thereby avoiding the inherent error propagation of the cascade architecture. First, the present invention constructs a context-aware feature extractor with an embedded epipolar transformer between views, and the more expressive features extracted enable the network to more accurately reason about depth values ​​in challenging areas. Secondly, the present invention constructs a view aggregation module with an embedded inter-view epipolar transformer, which uses the view weights inferred by the cross-attention idea to improve the point cloud reconstruction effect of invisible areas. The MVS module enhanced by the above two epipolar transformers can effectively improve the quality of depth map and point cloud reconstruction. Then, the present invention constructs an implicit neural optimization module, which uses implicit neural functions to achieve coarse depth map super-resolution, and effectively improves the accuracy of the edge area of ​​foreground objects in the depth map and point cloud. Finally, the present invention constructs an enhanced Gauss-Newton module to further refine the depth map as a whole, thereby comprehensively improving the reconstruction quality. The multi-view MVSNet network based on epipolar transformer and implicit neural optimization enhancement is experimentally verified by training and testing on the DTU dataset. Figure 3 The dimensional reconstruction method can significantly improve the quality of predicted depth maps and reconstructed point clouds, while taking up less memory and having a faster reasoning speed than other progressive refinement methods.

[0115] It is easy for those skilled in the art to understand that, under the premise of no conflict, the above-mentioned preferred solutions can be freely combined and superimposed.

[0116] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A multi-view 3D reconstruction method based on an epipolar transformer and an MVSNet network enhanced by implicit neural optimization, characterized in that: The following steps are involved: Step S1: Obtain a set of multi-view images of the target scene, select one image from them each time as a reference image and select (N-1) source images with the smallest image angle for them, and use the N images as input; input the full-resolution reference image and the source image into a scene-aware feature extractor embedded with an epipolar transformer to obtain a quarter-resolution feature map of the reference image and the source image, i.e., a reference feature map and a source feature map; Step S2: Map N-1 source feature maps to forward parallel planes divided by several depth hypotheses in the reference camera cone through differentiable homography transformation; calculate the inter-group correlation similarity between the reference feature map and the twisted source feature map to form N-1 similarity volumes; aggregate the similarity volumes into a unified matching cost volume through a view aggregation module embedded with an inter-view epipolar transformer; use 3D CNN to regularize the cost volume to obtain a probability volume along the depth direction; and then calculate a quarter-resolution depth map from the probability volume through a depth regression operation; Step S3: construct an implicit neural optimization module to implement the depth map super-resolution process, taking quarter-resolution pixel coordinates, half-resolution pixel coordinates, quarter-resolution depth map, and full-resolution reference image as input, and outputting an optimized half-resolution depth map; Step S4: Using the context-aware feature extractor described in step S1 to replace the feature pyramid used in the traditional Gauss-Newton module, thereby constructing an enhanced Gauss-Newton module; using this module to further refine the half-resolution depth map to obtain a final depth map, the size of which is still half resolution; Step S5: Replace the reference image and repeat the above steps until the depth maps corresponding to all images of the target scene are obtained.

2. The multi-view 3D reconstruction method based on the MVSNet network enhanced by epipolar transformer and implicit neural optimization according to claim 1, characterized in that: The context-aware feature extractor described in step S1 is constructed by embedding a view extrema transformer in a feature pyramid-based cyclic feature extractor: the feature pyramid-based cyclic feature extractor has a "top-down-bottom-up" cyclic reciprocating structure, which can extract context-aware features that contain both low-level spatial information and high-level semantic information; the view extrema transformer has a multi-head self-attention layer and a multi-head cross-attention layer, which can aggregate non-local features along the extrema, thereby further enhancing the distinguishability and describability of the features.

3. The multi-view 3D reconstruction method based on the MVSNet network enhanced by epipolar transformer and implicit neural optimization according to claim 1, characterized in that: The similarity volumes are aggregated into a unified matching cost volume through the view aggregation module embedded with the inter-view epipolar line transformer described in step S2, and the process is as follows: First, through the inter-view epipolar transformer, based on the cross-attention idea of ​​the transformer, a long-distance 3D correlation relationship along the depth direction is constructed between the matching pixels on the epipolar lines in different views, and the weight w of each similarity volume is calculated. i : Where F0 and W i They represent the reference feature and the source feature volume twisted according to the differentiable homography transformation, i=1,…,N-1 represents the corresponding source view number, Ch represents the number of feature channels; in the cross-attention mechanism, W0 plays the role of the query vector, W i The role of is the key vector, and the two are calculated as above. i , which can be used to weight the value vector; the unified matching cost volume C is calculated as follows: Where S i is the similarity volume, and its role is the value vector.

4. The multi-view 3D reconstruction method based on the epipolar transformer and the MVSNet network enhanced by implicit neural optimization according to claim 1, characterized in that: The implicit neural optimization process described in step S3 is as follows: The implicit neural function f implemented by a multilayer perceptron θ , we can jointly estimate the interpolated depth value v q,p and the interpolation weight w q,p : w q,p ,v q,p =f θ (d p ,g p ,g q -g p ,x q -x p ) In the formula, x p and x q denote quarter and half resolution pixel coordinates, respectively, g p and g q denote the quarter-resolution and half-resolution reference feature maps extracted from the full-resolution reference image using the feature pyramid, d p Represents the use of feature pyramids from a quarter-resolution coarse depth map D 1 The extracted quarter-resolution depth feature map; finally, the optimized half-resolution depth map D 2 The calculation is done by interpolation as follows: Where D 2 (x q ) means D 2 The mid-coordinate x q The depth value, represents the set of neighboring pixels of pixel q in the half-resolution image domain in the quarter-resolution image domain.

Citation Information

Patent Citations

  • Real-time three-dimensional reconstruction method applied to binocular endoscope medical image

    CN110033465A

  • System and method for depth estimation by learning triangulation and densification of multi-view stereoscopic sparse points

    CN115210532A

  • Multi-view three-dimensional network three-dimensional reconstruction method based on attention cost body pyramid

    CN115239870A

  • Image Processing Apparatus, System, Method and Computer Program Product for 3D Reconstruction

    US20160210776A1

  • Method and system that uses an anisotropy parameter to generate high-resolution time-migrated image gathers for reservoir characterization, and interpretation

    US20220066059A1

Cited By

  • Scene and sensor data generation method based on remote sensing image and electronic equipment

    CN121962800A