Depth-priori-guided multi-view three-dimensional reconstruction method
By introducing depth priors and Transformer to perform cost-body aggregation in the multi-view three-dimensional reconstruction method, the problems of inaccurate depth estimation and insufficient adaptability of cost-aggregation in the prior art are solved, and higher quality depth maps and three-dimensional point cloud reconstruction are achieved.
Patent Information
- Application Number
- CN202510109915.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The existing multi-view three-dimensional reconstruction method based on deep learning is inaccurate in the depth estimation under the conditions of lighting, occlusion and large-scale changes, and the cost aggregation method fails to fully consider the differences between views, resulting in poor noise suppression effect.
The multi-view three-dimensional reconstruction method with depth prior guidance is adopted to obtain multi-view features through feature extraction networks, build feature bodies, and use Transformer to perform cost body aggregation in the depth map generation module. The depth prior guidance aggregation process is used to calculate the contribution weight of different views to cost body.
The quality of depth maps and three-dimensional point clouds is improved, the adaptability of cost body aggregation is enhanced, noise interference is reduced, error transmission is avoided by external models, and the efficiency and stability of the overall model are improved.
Smart Images

Figure CN120125635A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image technology, and particularly to a multi-view stereo three-dimensional reconstruction method guided by depth prior. Background Art
[0002] Multi-view stereo matching (MVS) is a technology for recovering the three-dimensional structure of a scene from multiple two-dimensional images, and is widely applied in the fields of computer vision, remote sensing, robot navigation, and three-dimensional reconstruction. Traditional MVS methods mainly rely on geometric algorithms to estimate the depth and structure of a three-dimensional scene by matching feature points or regions in multiple images with large perspective differences. However, traditional methods often face problems such as unstable matching and error propagation when dealing with situations such as illumination, occlusion, and large-scale changes, resulting in a decline in the quality of the reconstructed dense point cloud.
[0003] Although the MVS method based on deep learning has made significant progress in many fields, it still faces the problem of inaccurate depth estimation in some difficult conditions, such as scenes with geometric, specular reflection, and high illumination changes. In these scenarios, cost aggregation is one of the key steps affecting the accuracy of depth estimation. However, existing cost aggregation methods (such as simple addition aggregation or variance aggregation) have some inherent defects: they do not fully consider the differences between views and lack adaptive processing of factors such as occlusion and illumination changes. This makes it impossible for cost aggregation to effectively suppress noise in scenes with drastic geometric and illumination changes, resulting in depth estimation bias, and thus affecting the accuracy and stability of three-dimensional reconstruction.
[0004] Therefore, how to improve the adaptability in the cost aggregation process, especially in scenes with drastic geometric structures and illumination changes, has become one of the main technical problems faced by the current MVS method based on deep learning.
[0005] Currently, multi-view stereo 3D reconstruction methods, such as MVSNet and CASMVSNet, usually aggregate all cost volumes through variance; for example, DPSNet aggregates all cost volumes using simple addition. Whether it is variance aggregation or addition aggregation, the core principle is to consider the contributions of all views equally. To reduce invalid matches in cost volume aggregation, PVA-MVSNet uses gated convolution to adaptively aggregate the cost volume. This method adjusts the weights of occluded regions in the matching through a weight map, making the weights of occluded regions relatively small. The weight map is generated based on the cost volume information and follows a self-attention mechanism. This adaptive method depends on training parameters rather than simply dynamically updating based on the data itself. Vis-MVSNet explicitly introduces a measurement method for cost volumes by examining the uncertainty or confidence of the probability distribution, regarding it as a visibility measurement. Some researchers have also proposed an adaptive aggregation module to enhance the reliability of cost volume aggregation, and at the same time, by adding an edge detection branch, constrain the consistency of edge features in the epipolar direction.
[0006] However, existing algorithms have great limitations. Some simply aggregate the cost volume, which often leads to invalid matching relationships. Although some methods complete the adaptive aggregation of the cost volume, the aggregation operation itself often requires additional training of adaptive parameters. There are also methods that introduce some external models, such as monocular depth estimation networks or edge detection networks, to enhance the aggregation effect, but this may bring error transmission of the external models. Summary of the Invention
[0007] To solve the above problems existing in the prior art, the present invention provides a multi-view stereo 3D reconstruction method guided by depth prior, specifically including:
[0008] In a first aspect, the present invention provides a multi-view stereo 3D reconstruction method guided by depth prior, including:
[0009] Obtain N original images of the target to be reconstructed taken from different perspectives, where N is a positive integer greater than or equal to 2;
[0010] Traverse the obtained N original images, and one by one use each original image as a reference image and the remaining N - 1 original images as source images, and input them into a trained depth map generation network to obtain N depth maps;
[0011] Fuse the N depth maps to obtain the 3D point cloud of the target to be reconstructed;
[0012] Among them, the depth map generation network includes a feature extraction network, a rough-resolution depth map generation module, a first refined-resolution depth map generation module, and a second refined-resolution depth map generation module;
[0013] During the generation process of any first depth map among N depth maps:
[0014] A feature extraction network, configured to obtain three different-scale reference feature maps corresponding to the input reference image and three different-scale source feature maps corresponding to the input source image according to the input reference image and the input source image;
[0015] A rough-resolution depth map generation module, configured to obtain a first intermediate depth map based on the internal Transformer according to the reference feature map with the smallest scale and the source feature maps with the smallest scale at each scale;
[0016] A first refined-resolution depth map generation module, configured to obtain a first depth prior according to the first intermediate depth map and the input reference image, and complete cost volume aggregation based on the internal Transformer according to the first depth prior, the reference feature map with the middle scale, and the source feature maps with the middle scale at each scale, and obtain a second intermediate depth map based on the aggregation result;
[0017] A second refined-resolution depth map generation module, configured to obtain a second depth prior according to the second intermediate depth map and the input reference image, and complete cost volume aggregation based on the internal Transformer according to the second depth prior, the reference feature map with the largest scale, and the source feature maps with the largest scale at each scale, and obtain a first depth map based on the aggregation result.
[0018] In a second aspect, the present invention further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0019] The memory is used to store a computer program;
[0020] The processor is configured to implement any method provided in the first aspect when executing the program stored on the memory.
[0021] Advantages of the present invention:
[0022] The multi-view stereo three-dimensional reconstruction method guided by depth prior provided by the present invention obtains N original images of the target to be reconstructed taken from different perspectives; traverses the obtained N original images, takes each original image as a reference image one by one, and the remaining N-1 original images as source images, inputs them into a trained depth map generation network to obtain N depth maps; fuses the N depth maps to obtain the three-dimensional point cloud of the target to be reconstructed. Based on its network architecture, the depth map generation module at the back uses the depth map output by the depth map generation module at the front to generate a depth prior, and guides the Transformer in the depth map generation module at the back to complete cost volume aggregation through the depth prior to generate the depth map at the back, thereby effectively calculating the contribution weights of different views to the cost volume, improving the final quality of the depth map and the point cloud. At the same time, the generation of the depth prior completely depends on the internal structure design of the network, without relying on external models, avoiding the error transmission that may be brought by external models, and improving the efficiency and stability of the overall model.
[0023] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Description of the Drawings
[0024] Figure 1 It is a schematic flowchart of a multi-view stereo three-dimensional reconstruction method guided by depth prior provided by the present invention;
[0025] Figure 2 It is a schematic diagram of a simulation result provided by the present invention;
[0026] Figure 3 It is another schematic diagram of a simulation result provided by the present invention;
[0027] Figure 4 It is yet another schematic diagram of a simulation result provided by the present invention. Detailed Embodiments
[0028] The following further describes the present invention in detail in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0029] To solve the problems existing in the prior art, the present invention provides a multi-view stereo three-dimensional reconstruction method guided by depth prior. This method first uses a feature extraction network to obtain multi-view features and constructs a feature volume. At the initial stage of rough resolution, the feature volume is directly aggregated by the Transformer to generate a cost volume; at the subsequent stage of refined resolution, depth prior features are generated using the depth information estimated in the previous stage and the RGB image to further enhance the Transformer to achieve more accurate cost volume aggregation. At each stage, through 3D regularization, the depth map is gradually estimated. Finally, the final depth maps are fused to generate a dense three-dimensional point cloud.
[0030] Figure 1 The flowchart of a multi-view stereo three-dimensional reconstruction method guided by depth prior provided by the present invention is shown as Figure 1 follows. This method includes:
[0031] S101. Obtain N original images of the target to be reconstructed taken from different perspectives.
[0032] Where N is a positive integer greater than or equal to 2.
[0033] S102. Traverse the obtained N original images. One by one, take each original image as the reference image, and the remaining N - 1 original images as the source images, and input them into the trained depth map generation network to obtain N depth maps.
[0034] Exemplarily, assume that the total number of obtained original images is 3, namely image A, image B, and image C. Take image A as the reference image, and image B and image C as the source images and input them into the trained depth map generation network to obtain 1 depth map. Take image B as the reference image, and image A and image C as the source images and input them into the trained depth map generation network to obtain 1 depth map. Take image C as the reference image, and image A and image B as the source images and input them into the trained depth map generation network to obtain 1 depth map, thus obtaining 3 depth maps.
[0035] S103. After fusing the N depth maps, obtain the three-dimensional point cloud of the target to be reconstructed.
[0036] Where the depth map generation network includes a feature extraction network, a rough resolution depth map generation module, a first refined resolution depth map generation module, a second refined resolution depth map generation module, and a depth map generation module.
[0037] During the generation process of any first depth map among the N depth maps:
[0038] The feature extraction network is used to obtain three different-scale reference feature maps corresponding to the input reference image and three different-scale source feature maps corresponding to the input source image according to the input reference image and the input source image.
[0039] The rough resolution depth map generation module is used to obtain a first intermediate depth map based on the internal Transformer according to the reference feature map with the smallest scale and the source feature maps with the smallest scale of each scale.
[0040] The first refined resolution depth map generation module is used to obtain a first depth prior according to the first intermediate depth map and the input reference image, and based on the internal Transformer, complete cost volume aggregation according to the first depth prior, the reference feature map with the middle scale, and the source feature maps with the middle scale of each scale, and obtain a second intermediate depth map based on the aggregation result.
[0041] The second refined resolution depth map generation module is used to obtain a second depth prior based on the second intermediate depth map and the input reference image, and based on the internal Transformer, complete cost volume aggregation according to the second depth prior, the reference feature map with the largest scale, and the source feature maps with the largest scale at each scale, and obtain the first depth map based on the aggregation result.
[0042] The depth map generation module is used to obtain the depth map according to the aggregated cost volume.
[0043] Optionally, the feature extraction network is a multi-scale feature network, specifically a feature pyramid structure. That is, after inputting a single-scale image, the downsampling module of UNet is used to perform convolutional operations step by step to extract three different-scale feature maps. Then, through the convolutional operations and skip connections of the upsampling module, these feature maps are synthesized and output to obtain three different-scale feature maps.
[0044] Furthermore, in order to make the predicted depth closer to the real depth and obtain a more accurate depth map, before step S102, it also includes respectively setting the depth sampling number, depth sampling range, and depth sampling interval of the rough resolution depth map generation module, the first refined resolution depth map generation module, and the second refined resolution depth map generation module.
[0045] Optionally, the rough resolution depth map generation module, the first refined resolution depth map generation module, and the second refined resolution depth map generation module are connected in sequence. Among them, if the depth sampling number corresponding to the previous module is D, the depth of the depth map is d, and the sampling interval is L, then the sampling interval in the module is The depth sampling range is
[0046] Furthermore, based on the multi-stage cascaded structure provided by the present invention, which consists of a rough resolution depth map generation module, a first refined resolution depth map generation module, and a second refined resolution depth map generation module, in the initial stage, the rough resolution depth map generation module obtains features with a small scale, and in the gradually refined stages, the first refined resolution depth map generation module and the second refined resolution depth map generation module sequentially obtain features with a larger scale. The resolution of each stage is twice that of the previous stage. Optionally, the resolutions of the rough resolution depth map generation module, the first refined resolution depth map generation module, and the second refined resolution depth map generation module are successively: and H×W, where H represents the image height and W represents the image width. Correspondingly, when extracting image features, the different-scale features extracted will also correspond to the above resolutions, which are and H×W×C, where C represents the number of channel features.
[0047] Further, during the generation process of any first depth map among the N depth maps, the rough-resolution depth map generation module is specifically configured to perform the following steps A1 - A4:
[0048] A1. Obtain the smallest-scale reference feature volume and each smallest-scale source feature volume according to the smallest-scale reference feature map and each smallest-scale source feature map.
[0049] A2. Based on the Transformer inside the rough-resolution depth map generation module, use the smallest-scale reference feature volume as the query vector, use each smallest-scale source feature volume as the key vector, and determine the correlation weight between each smallest-scale source feature volume and the smallest-scale reference feature volume. The corresponding expression is:
[0050]
[0051] where w i represents the correlation weight between the i-th smallest-scale source feature volume and the smallest-scale reference feature volume, and Softmax(·) represents a function that normalizes the correlation weight into a probability distribution. represents the key vector transformed from the i-th smallest-scale source feature volume, Q r represents the query vector transformed from the smallest-scale reference feature volume, c represents the number of channels corresponding to the key vector and the query vector, T represents transpose, and i represents the index of the smallest-scale source feature volume.
[0052] A3. Determine the single-view cost volume corresponding to each smallest-scale source feature volume according to the smallest-scale reference feature volume and each smallest-scale source feature volume. The corresponding expression is:
[0053] C i = <F i · F r >,
[0054] where C i represents the single-view cost volume corresponding to the i-th smallest-scale source feature volume, F i represents the i-th smallest-scale source feature volume, F r represents the smallest-scale reference feature volume, and <·> represents the inner product.
[0055] A4. Aggregate the single-view cost volumes corresponding to each smallest-scale source feature volume, and obtain the first intermediate depth map according to the aggregation result.
[0056] Further, during the generation process of any first depth map among the N depth maps, the first refined-resolution depth map generation module is specifically configured to perform the following steps B1 - B6:
[0057] B1. Sample the input reference image to obtain a first reference image, where the scale of the first reference image is the same as that of the reference feature map centered at the scale.
[0058] B2. Obtain a first depth prior based on the first reference image and the first intermediate depth map.
[0059] B3. Obtain a reference feature volume centered at the scale and source feature volumes centered at each scale based on the reference feature map centered at the scale and the source feature maps centered at each scale.
[0060] B4. Based on the Transformer inside the first refined resolution depth map generation module, use the first depth prior as the query vector and the source feature volumes centered at each scale as the key vectors to determine the correlation weights between the source feature volumes centered at each scale and the reference feature volume centered at the scale. The corresponding expression is:
[0061]
[0062] Among them, represents the correlation weight between the i 1 -th source feature volume centered at the scale and the reference feature volume centered at the scale, represents the key vector transformed from the i 1 -th source feature volume centered at the scale, represents the query vector transformed from the reference feature volume centered at the scale, i 1 represents the index of the source feature volume centered at the scale, represents the query vector transformed from the first depth prior.
[0063] B5. Determine the single-view cost volume corresponding to each source feature volume centered at the scale based on the reference feature volume centered at the scale and the source feature volumes centered at each scale. The corresponding expression is:
[0064]
[0065] Among them, represents the single-view cost volume corresponding to the i 1 -th source feature volume centered at the scale, represents the i 1 -th source feature volume centered at the scale, represents the reference feature volume centered at the scale.
[0066] B6. Aggregate the single-view cost volumes corresponding to the source feature volumes centered at each scale, and obtain a second intermediate depth map based on the aggregation result.
[0067] This method uses the depth map generated by the previous depth map generation module and the corresponding reference image to generate a depth prior, and then introduces the depth prior as the query vector of the Transformer, so as to obtain more accurate adaptive weights and improve the cost volume aggregation effect.
[0068] During the generation process of any first depth map among N depth maps, the processing process of the second refined resolution depth map generation module is similar to that of the first refined resolution depth map generation module, specifically including the following steps C1 - C6:
[0069] C1. Sample the input reference image to obtain a second reference image, and the scale of the second reference image is the same as that of the reference feature map with the largest scale.
[0070] C2. Obtain a second depth prior according to the second reference image and the second intermediate depth map.
[0071] C3. Obtain the reference feature volume with the largest scale and each source feature volume with the largest scale according to the reference feature map with the largest scale and each source feature map with the largest scale.
[0072] C4. Based on the Transformer inside the second refined resolution depth map generation module, use the second depth prior as the query vector and each source feature volume with the largest scale as the key vector to determine the correlation weight between each source feature volume with the largest scale and the reference feature volume with the largest scale. The corresponding expression is:
[0073]
[0074] Among them, represents the correlation weight between the i 2 -th source feature volume with the largest scale and the reference feature volume with the largest scale, represents the key vector converted from the i 2 -th source feature volume with the largest scale, represents the query vector converted from the reference feature volume with the largest scale, i 2 represents the index of the source feature volume with the largest scale, represents the query vector converted from the second depth prior.
[0075] After adding the depth prior feature, the depth prior can guide cost aggregation, reduce noise interference, and more accurately capture the geometric distribution of the scene.
[0076] C5. Determine the single-view cost volume corresponding to each source feature volume with the largest scale according to the reference feature volume with the largest scale and each source feature volume with the largest scale. The corresponding expression is:
[0077]
[0078] Among them, represents the single-view cost volume corresponding to the source feature volume with the largest scale at the $i$-th 2 scale, represents the source feature volume with the largest scale at the $i$-th 2 scale, represents the reference feature volume with the largest scale.
[0079] C6. Aggregate the single-view cost volumes corresponding to the source feature volumes with the largest scale at each scale, and obtain the first depth map according to the aggregation result.
[0080] Specifically, the first refined-resolution depth map generation module includes several convolutional layers (encoder) and transposed convolutional layers (decoder). The first depth prior is generated through the encoder and decoder parts of the first refined-resolution depth map generation module. Among them, the encoder is connected to the encoder part through skip connections to enhance the transmission of detailed information. In addition, the decoder part of the first refined-resolution depth map generation module also performs feature dimension concatenation with the encoder part of the second refined-resolution depth map generation module to further optimize the expression of the depth prior. The second refined-resolution depth map generation module is also composed of several convolutional layers (decoder) and transposed convolutional layers (decoder). The input of this module includes the first depth prior passed to the first refined-resolution depth map generation module. It should be noted that the encoder part of the second refined-resolution depth map generation module is connected to the decoder part of the first refined-resolution depth map generation module through feature dimension concatenation, which enables the encoder of the second refined-resolution depth map generation module to receive more information from the decoder. Therefore, in the second refined-resolution depth map generation module, the depth prior undergoes more sufficient information fusion, further optimizing the expression of the depth prior. Finally, the second depth prior is generated through the encoder and decoder parts of the second refined-resolution depth map generation module. This module also adopts skip connections between the encoder and decoder to enhance the feature expression ability.
[0081] Furthermore, the coarse-resolution depth map generation module, the first refined-resolution depth map generation module, and the second refined-resolution depth map generation module are all involved in how to obtain the feature volume, how to aggregate the single-view cost volume, and how to obtain the depth map according to the aggregation result of the single-view cost volume. The implementation processes of each module are similar, as follows:
[0082] Optionally, obtaining the feature volume from the feature map includes: constructing a homography matrix according to the camera pose and depth hypothesis corresponding to the feature map; converting the feature map into a feature volume through the homography matrix.
[0083] The corresponding expression is:
[0084]
[0085] Among them, H represents the homography matrix, and K 2 represents the camera extrinsic parameter matrix corresponding to the feature Figure 2 viewpoint, R 1 represents the rotation matrix corresponding to the feature Figure 1 R 2 represents the rotation matrix corresponding to the feature Figure 2 I represents the identity matrix, and t 1 represents the translation matrix corresponding to the feature Figure 1 t 2 represents the translation matrix corresponding to the feature Figure 2 t, n represents the normal vector corresponding to the feature Figure 1 d 1 represents the depth hypothesis corresponding to the feature Figure 1 K 1 represents the camera extrinsic parameter matrix corresponding to the feature Figure 2 viewpoint. The superscript T represents the transpose. The feature Figure 2 is the reference feature map; the feature Figure 1 is the same as the feature Figure 2 , or is any source feature map with the same scale as the feature Figure 2
[0086] The rough-resolution depth map generation module, the first refined-resolution depth map generation module, and the second refined-resolution depth map generation module all perform homography transformation on the corresponding reference feature map and source feature map using the corresponding homography matrix, so as to project the pixel features on the pixel plane corresponding to the source feature map onto the pixel plane corresponding to the reference feature map, and then obtain a plurality of feature volumes.
[0087] Exemplarily, the dimensions of the feature volume at each stage are D 3 ×H×W×C, where D 1 , D 2 , D 3 is the number of depth hypotheses corresponding to the rough-resolution depth map generation module, the first refined-resolution depth map generation module, and the second refined-resolution depth map generation module. The feature volumes corresponding to each depth map generation module are divided into 1 reference feature volume F r and N - 1 source feature volumes F i represents the feature volume corresponding to the i-th source feature map. The reference feature volume is a copy of the original feature map along the depth direction, and the reference feature volume is equivalent to the true value. The closer the depth hypothesis is to the true depth, the higher the similarity between the source feature volume and the reference feature volume.
[0088] Optionally, aggregate each single-view cost volume, which is expressed as:
[0089]
[0090] Among them, C agg represents the aggregation result of each single-view cost volume, and w j represents the correlation weight between the source feature volume corresponding to the j-th single-view cost volume and the corresponding reference feature volume, and C j represents the j-th single-view cost volume, and J represents the total number of single-view cost volumes used for aggregation.
[0091] In the cost aggregation process of the present invention, no additional parameter learning is involved. Instead, the feature volume and the depth prior feature volume are dynamically calculated, and the cost volume is directly updated, effectively avoiding the learning burden brought by the introduction of additional parameters, and at the same time ensuring the flexibility and adaptability of the cost aggregation process.
[0092] Optionally, obtaining a depth map according to the aggregation result of the single-view cost volume includes: using a multi-scale 3D convolutional network to perform noise reduction processing on the aggregation result of the single-view cost volume, and converting the denoised single-view cost volume into a probability volume; obtaining depth hypotheses based on a preset depth sampling range, depth sampling interval, and number of depth samplings; calculating depth probabilities according to the probability volume; predicting the image depth based on the calculated depth probabilities and depth hypotheses to obtain a depth map, and the corresponding expression is:
[0093]
[0094] Among them, D yc represents the predicted image depth, [l min , l max represents the preset depth sampling range, l represents the depth hypothesis of the current image, and P(l) represents the probability under the depth hypothesis of l.
[0095] It should be noted that the depth hypothesis is for each pixel, and each pixel corresponds to a set of depth hypotheses. By selecting a suitable depth hypothesis for the corresponding pixel in a set of depth hypotheses corresponding to each pixel, a depth map closer to the real situation can be obtained.
[0096] Further optionally, the loss function of the depth map generation network is expressed as:
[0097]
[0098] Among them, L total represents the loss function of the depth map generation network, represents the loss function of the rough-resolution depth map generation module, represents the loss function of the first refined-resolution depth map generation module, represents the loss function of the second refined-resolution depth map generation module.
[0099] There are varying degrees of errors between the depth maps output by each depth map generation module and the true values. Cross-entropy can be used as the loss function for each depth map generation module, which is expressed as:
[0100]
[0101] Among them, represents the loss function of the g-th stage, ρ represents valid pixels, and P GT (z) represents the true value of the probability distribution, and P(z) represents the predicted probability distribution.
[0102] The multi-view stereo three-dimensional reconstruction method guided by depth prior provided by the present invention obtains N original images of the target to be reconstructed taken from different perspectives; traverses the obtained N original images, takes each original image as a reference image one by one, and the remaining N-1 original images as source images, inputs them into a trained depth map generation network to obtain N depth maps; fuses the N depth maps to obtain the three-dimensional point cloud of the target to be reconstructed. Based on its network architecture, the depth map generation module at the back uses the depth map output by the depth map generation module at the front to generate a depth prior, and completes cost volume aggregation through the Transformer in the depth map generation module at the back to generate the depth map at the back, thereby effectively calculating the contribution weights of different views to the cost volume, improving the final quality of the depth map and the point cloud. At the same time, the generation of the depth prior completely depends on the internal structure design of the network, without relying on external models, avoiding the error transmission that may be brought by external models, and improving the efficiency and stability of the overall model.
[0103] The present invention also provides a structure of an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus.
[0104] The memory is used to store computer programs.
[0105] When the processor is used to execute the programs stored on the memory, it implements the steps provided in the above method embodiments.
[0106] The communication interface is used for communication between the above electronic device and other devices.
[0107] To further prove the beneficial effects of the present invention, the present invention also provides a set of simulation experiments, which are specifically as follows:
[0108] Quantitative experiments were conducted on the DTU dataset. The DTU dataset is an indoor dataset with multi-view images and camera poses, containing 124 scenes, each scene covering 49 or 64 views and 7 lighting conditions. According to MVSNet, the dataset was divided into 79 training sets, 18 validation sets, and 22 test sets, with a total of 27,097 training samples.
[0109] The number of input images N = 5, the image size is 1152×864, depth filtering is performed through geometric constraints and photometric constraints, and depth fusion of Gipuma is used to obtain the final 3D point cloud. The accuracy (Acc.), completeness (Comp.), and overall performance (overall) of the reconstructed point cloud are calculated by using the official MATLAB code provided by DTU. The overall performance is the average of accuracy and completeness (the lower the better), and the calculation formula is as follows.
[0110] The reconstruction method of the present invention was compared with traditional methods and some learning-based MVS methods, and the quantitative results are shown in Table 1. The reconstruction method of the present invention is superior to most traditional methods and learning-based methods in terms of completeness, achieving better overall performance. As Figure 2 are the results corresponding to the CasMVSNET method, Figure 3 are the results corresponding to the MVSNET method, Figure 4 are the results corresponding to the method provided by the present invention. As Figure 2 - 4 shown, the details are highlighted in the rectangle, and it can be seen that the reconstruction results of the present invention are more complete. And the geometric structure is clearer.
[0111] Table 1 Quantitative results of the method of the present invention and other methods on the DTU test set
[0112] Method Acc.(mm) Comp.(mm) Overall(mm) Furu 0.613 0.941 0.777 Gipuma 0.283 0.873 0.578 COLMAP 0.400 0.664 0.532 MVSNet 0.396 0.527 0.462 CasMVSNet 0.325 0.385 0.355 PVA - MVSNet 0.379 0.336 0.357 Vis - MVSNet 0.369 0.361 0.365 Ours 0.382 0.291 0.336
[0113] For the electronic device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the specific content, beneficial effects, etc., please refer to the partial description of the method embodiment.
[0114] The terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0115] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A depth prior guided multi-view stereo 3D reconstruction method, characterized in that: include: Obtain N original images of the target to be reconstructed taken from different perspectives, where N is a positive integer greater than or equal to 2; Traversing the N original images obtained, taking each of the original images as a reference image one by one, and the remaining N-1 original images as source images, inputting them into the trained depth map generation network to obtain N depth maps; Fusing the N depth maps to obtain a three-dimensional point cloud of the target to be reconstructed; Wherein, the depth map generation network includes a feature extraction network, a coarse resolution depth map generation module, a first refined resolution depth map generation module and a second refined resolution depth map generation module; In the process of generating any first depth map among the N depth maps: The feature extraction network is used to obtain three reference feature maps of different scales corresponding to the input reference image and three source feature maps of different scales corresponding to the input source image according to the input reference image and the input source image; The coarse resolution depth map generation module is used to obtain a first intermediate depth map based on the internal Transformer according to the reference feature map with the smallest scale and the source feature maps with the smallest scales; The first refined resolution depth map generation module is used to obtain a first depth prior according to the first intermediate depth map and the input reference image, and based on the internal Transformer, complete cost volume aggregation according to the first depth prior, the scale-centered reference feature map and the scale-centered source feature maps, and obtain a second intermediate depth map based on the aggregation result; The second refined resolution depth map generation module is used to obtain a second depth prior based on the second intermediate depth map and the input reference image, and based on the internal Transformer, completes cost volume aggregation according to the second depth prior, the reference feature map with the largest scale, and the source feature maps with the largest scales, and obtains the first depth map based on the aggregation result.
2. The method according to claim 1, characterized in that: In the process of generating any first depth map among the N depth maps, the coarse resolution depth map generating module is specifically used to: According to the reference feature map with the smallest scale and each source feature map with the smallest scale, a reference feature body with the smallest scale and each source feature body with the smallest scale are obtained; Based on the Transformer inside the coarse resolution depth map generation module, the reference feature body with the smallest scale is used as a query vector, and each source feature body with the smallest scale is used as a key vector to determine the correlation weight between each source feature body with the smallest scale and the reference feature body with the smallest scale. The corresponding expression is: Among them, w i represents the correlation weight between the i-th smallest source feature body and the smallest reference feature body. Softmax(·) represents a function that normalizes the correlation weight to a probability distribution. The key vector representing the transformation of the i-th source feature volume with the smallest scale, Q r represents the query vector converted from the smallest-scale reference feature body, c represents the number of channels corresponding to the key vector and the query vector, T represents transpose, and i represents the index of the smallest-scale source feature body; According to the reference feature volume with the smallest scale and each source feature volume with the smallest scale, the single view cost volume corresponding to each source feature volume with the smallest scale is determined, and the corresponding expression is: C i =<F i ·F r >, Among them, C i represents the single view cost volume corresponding to the source feature volume with the smallest scale of i, F i represents the source feature body with the smallest scale of i, F r represents the reference feature volume with the smallest scale, and <·> represents the inner product; The single view cost volumes corresponding to the source feature volumes with the smallest scales are aggregated, and a first intermediate depth map is obtained according to the aggregation result.
3. The method according to claim 2, characterized in that In the process of generating any first depth map among the N depth maps, the first refined resolution depth map generating module is specifically configured to: Sampling the input reference image to obtain a first reference image, wherein the scale of the first reference image is the same as the scale of the reference feature map centered at the scale; Obtaining a first depth prior according to the first reference image and the first intermediate depth map; According to the scale-centered reference feature map and each scale-centered source feature map, a scale-centered reference feature volume and each scale-centered source feature volume are obtained; Based on the Transformer inside the first refined resolution depth map generation module, the first depth prior is used as a query vector, and each of the scale-centered source feature bodies is used as a key vector to determine the correlation weight between each of the scale-centered source feature bodies and the scale-centered reference feature body. The corresponding expression is: in, represents the correlation weight between the i1th scale-centered source feature and the scale-centered reference feature, represents the key vector for the transformation of the i1th scale-centered source feature volume, represents the query vector transformed from the scale-centered reference feature, i1 represents the index of the scale-centered source feature, represents the query vector transformed by the first depth prior; According to the scale-centered reference feature volume and each scale-centered source feature volume, a single view cost volume corresponding to each scale-centered source feature volume is determined, and the corresponding expression is: in, represents the single view cost volume corresponding to the source feature volume centered at the i1th scale, represents the source feature body centered at the i1th scale, represents a reference feature body centered on the scale; The single view cost volumes corresponding to the source feature volumes at the middle of the scales are aggregated, and a second intermediate depth map is obtained according to the aggregation result.
4. The method according to claim 3, characterized in that Aggregate the cost volumes of each single view, expressed as: Among them, C agg represents the aggregation result of each single view cost volume, w j represents the correlation weight between the source feature volume corresponding to the j-th single view cost volume and the corresponding reference feature volume, C j represents the jth single view cost volume, and J represents the total number of single view cost volumes used for aggregation.
5. The method according to claim 4, characterized in that The depth map is obtained based on the aggregation result of the single view cost volume, including: Use a multi-scale 3D convolutional network to denoise the aggregation result of the single-view cost volume and convert the denoised single-view cost volume into a probability volume; A depth hypothesis is obtained based on a preset depth sampling range, a depth sampling interval, and a depth sampling number; Calculating depth probability according to the probability body; Based on the calculated depth probability and the depth hypothesis, the image depth is predicted to obtain a depth map. The corresponding expression is: Among them, D yc represents the predicted image depth, [l min ,l max ] represents the preset depth sampling range, l represents the depth hypothesis of the current image, and P(l) represents the probability under the depth hypothesis l.
6. The method according to claim 5, characterized in that The coarse resolution depth map generation module, the first refined resolution depth map generation module and the second refined resolution depth map generation module are connected in sequence, wherein, if the number of depth samples corresponding to the previous module is D, the depth of the depth map is d, and the sampling interval is L, then the sampling interval in the module is The depth sampling range is 7. The method according to claim 6, characterized in that The loss function of the depth map generation network is expressed as: L total =L stage1 +L stage2 +L stage3 , Among them, L total represents the loss function of the deep graph generation network, L stage1 represents the loss function of the coarse resolution depth map generation module, L stage2 represents the loss function of the first refined resolution depth map generation module, L stage3 Represents the loss function of the second refined resolution depth map generation module.
8. The method according to claim 7, characterized in that The feature body is obtained according to the feature graph, including: According to the camera pose and depth hypothesis corresponding to the feature map, the homography matrix is constructed, and the corresponding expression is: Among them, H represents the homography matrix, K2 represents the camera external parameter matrix corresponding to the view of feature map 2, R1 represents the rotation matrix corresponding to feature map 1, R2 represents the rotation matrix corresponding to feature map 2, I represents the identity matrix, t1 represents the translation matrix corresponding to feature map 1, t2 represents the translation matrix corresponding to feature map 2, n represents the normal vector corresponding to feature map 1, d1 represents the depth hypothesis corresponding to feature map 1, K1 represents the camera external parameter matrix corresponding to the view of feature map 2, the superscript T represents the transpose, feature map 2 is the reference feature map; feature map 1 is the same as feature map 2, or is any source feature map with the same scale as feature map 2; The feature map is converted into a feature volume by using the homography matrix.
9. The method according to claim 8, characterized in that The resolutions of the coarse resolution depth map generation module, the first refined resolution depth map generation module and the second refined resolution depth map generation module are respectively: H×W, where H represents the image height and W represents the image width.
10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing any of the methods described in claims 1-9 when executing a program stored in a memory.
Citation Information
Patent Citations
Visible light multi-view image three-dimensional reconstruction method based on deep learning
CN115564888A
Multi-view stereo reconstruction method based on adaptive learning and aggregation
CN115631223A
Deep learning multi-view stereoscopic three-dimensional reconstruction algorithm based on path aggregation
CN116778091A
Multi-view three-dimensional reconstruction method, measurement method and system
CN116958434A
Multi-view three-dimensional reconstruction system based on feature aggregation Transform
CN117765175A
Cited By
Point cloud and image dynamic depth map alignment method and device
CN120823137A