Dynamic Adaptive Multi-View Stereo Reconstruction Method and Model Based on Dual-Domain Information Fusion
By adopting dual-domain information fusion and dynamic adaptive propagation technology in the multi-view stereo reconstruction method, the existing methods solve the problems of high-frequency details loss and insufficient long-range information correlation capabilities when dealing with complex modes and reflective surfaces, and achieve higher quality three-dimensional reconstruction.
Patent Information
- Application Number
- CN202510238779.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Existing multi-view stereo reconstruction methods tend to lose high-frequency details when dealing with complex modes and reflective surfaces, and lack adaptive long-range information correlation capabilities, resulting in reduced reconstruction performance.
The dynamic adaptive multi-view stereo reconstruction method based on dual-domain information fusion is adopted, and the multi-scale airspace, frequency domain information fusion network and multi-scale depth prediction network are used to achieve airspace details enhancement and frequency domain dynamic aggregation. Combined with the dynamic range multi-scale adaptive propagation mechanism, robust depth assumptions are generated.
It significantly improves adaptability to complex geometric scenarios, enhances reconstruction quality, and can more effectively capture high-frequency details and complex structures.
Smart Images

Figure CN119741434B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a three-dimensional reconstruction method and model, and particularly to a dynamic adaptive multi-view stereo reconstruction method and model based on dual-domain information fusion. Background Art
[0002] Multi-view stereo reconstruction is a core task in the field of computer vision, aiming to recover and reconstruct the dense geometric representation of a scene from multiple overlapping perspective images. Its application scope includes autonomous driving, virtual reality, augmented reality, cultural relic protection, etc., and has received extensive attention and research from relevant scholars in the past few decades.
[0003] Multi-view stereo (MVS) methods use manually crafted feature representations for matching, thus facing challenges in areas such as repetitive patterns, weak textures, drastic illumination changes, and reflections. To make up for the weak representativeness of MVS methods, learning-based methods have emerged. These methods replace the manually crafted similarity metric and cost volume aggregation modules with deep learning-based feature extraction and 3D regularization networks, and thus exhibit excellent performance. In particular, learning-based MVS methods can better introduce global semantic information, such as specular and reflection priors, to support subsequent robust feature matching and depth generation.
[0004] However, existing methods only mine the feature representation of input images in the spatial domain, have limited processing capabilities for complex patterns such as rotation and scaling, are prone to losing high-frequency details, and especially difficult to extract robust features from reflective surfaces, weak textures, and textureless regions. In addition, considering computational efficiency and memory limitations, existing methods usually use a fixed range of attention or only use attention along the channel dimension to explore the spatial relationship between pixels, lacking the ability of adaptive long-range information association, resulting in a reduction in image reconstruction performance. Therefore, how to effectively utilize the advantages of frequency domain processing in capturing high-frequency details and complex structures, and explore the long-range spatial dependencies in the depth propagation process is crucial for improving the reconstruction quality of the scene. Summary of the Invention
[0005] The object of the present invention is to solve the deficiencies in the prior art that use attention along the channel dimension to explore the spatial relationship between pixels, lack the ability of adaptive long-range information association, and result in a reduction in image reconstruction performance, and to provide a dynamic adaptive multi-view stereo reconstruction method and model based on dual-domain information fusion.
[0006] The inventive concept of the present invention is as follows: In the dynamic adaptive multi-view stereo reconstruction method based on dual-domain information fusion of the present invention, first, multi-view images of the target scene and the internal and external parameters of the camera are obtained, the multi-view images are input, and the multi-view images are divided into one reference image and multiple source images. Then, a multi-scale spatial and frequency domain information fusion network is constructed to achieve spatial detail enhancement and frequency domain dynamic aggregation and enhancement, so as to obtain multi-scale enhanced features of the input images. Then, a multi-scale depth prediction network is constructed using a coarse-to-fine framework. For the features at the lowest scale, depth hypotheses are uniformly sampled along the inverse depth range; for the features at the other scales, an uncertainty hypothesis mechanism based on dynamic range multi-scale adaptive propagation is constructed, and through the dynamic range multi-scale adaptive propagation mechanism, depth hypotheses at each scale are obtained. Using the internal and external parameters of the camera, the multi-scale features are warped to the front parallel plane of the reference camera frustum through differentiable homography transformation, and an adaptive group correlation strategy is used to adaptively aggregate the reference features and the source features to generate a cost volume. Finally, the cost volume is regularized using 2D regularization compensated by probability volume geometric priors to obtain the predicted depth, and finally the predicted depth is converted to obtain the three-dimensional point cloud file of the corresponding scene. In the process of establishing a model that can implement the above method, the network training is guided by the L1 loss at each stage. This method is accurate, flexible and robust, and provides effective technical support for the multi-view stereo matching and reconstruction task.
[0007] To achieve the above object, the technical solution provided by the present invention is as follows:
[0008] A dynamic adaptive multi-view stereo reconstruction method based on dual-domain information fusion, characterized in that it includes the following steps:
[0009] S1, input N multi-view images of the same scene to obtain multi-scale features , where, represents the scale, , represents the serial number of the multi-view image, taking values from 0 to N-1, and one of the N multi-view images is used as the reference image, represents the reference image, and the remaining N-1 images are used as source images;
[0010] S2, perform spatial detail enhancement on the multi-scale features to output the features corresponding to each multi-view image ; perform frequency domain dynamic aggregation and enhancement on the multi-scale features to output the features corresponding to each multi-view image ;
[0011] S3, fuse the corresponding features with the features to obtain the fused feature , and then The upsampling is respectively added to the corresponding multi-scale features to obtain multi-scale enhanced features ;
[0012] S4. Determine the probability volume of the reference image and its corresponding predicted depth at the scale of k = 0 ; ;
[0013] S5. Determine the predicted depths at the scales of k = 1, 2, and 3 ; The predicted depths at the corresponding scales are specifically obtained through the following method:
[0014] 1) Based on the multi-scale enhanced features , use the uncertainty hypothesis mechanism based on dynamic range multi-scale adaptive propagation to determine the corresponding depth hypotheses;
[0015] 2) Use the depth hypotheses to perform differentiable homography transformations on the enhanced features corresponding to the current scale respectively, to obtain the transformed feature volume , and apply the adaptive group correlation strategy to aggregate the transformed feature volume to obtain the cost volume at the current scale;
[0016] 3) Use 2D regularization to process the cost volume at the current scale to obtain the corresponding probability volume ; Weight-sum the probability volume along the depth direction to obtain the predicted depth at the current scale ;
[0017] S6. Convert the predicted depth at the scale of k = 3 into a three-dimensional point cloud file of the corresponding scene, and complete the dynamic adaptive multi-view stereo reconstruction based on dual-domain information fusion.
[0018] Furthermore, step S2 is specifically as follows:
[0019] S2.1. Perform depth convolution on the multi-scale features to obtain intermediate features ;
[0020] S2.2. Perform a proportional splitting operation on the intermediate features along the channels to obtain , , , and perform depth convolution on , respectively to obtain convolutional features and convolutional features ;
[0021] S2.3. Take , , After cascading along the channel direction, the channel information is integrated through pointwise convolution to obtain the features corresponding to each multi-view image ;
[0022] S2.4. For the multi-scale features The features within the corresponding frequency range are adaptively weighted according to the frequency, highlighting the required frequencies and suppressing the useless frequencies, to obtain the weighted spectral data corresponding to each multi-view image ;
[0023] S2.5. The weighted spectral data is transformed from the spatial domain to the frequency domain through Fourier transform, and the amplitude and phase are calculated, and they are respectively dynamically enhanced through the modulation network, and then the residual connections are added respectively to obtain the new amplitude and the new phase ;
[0024] S2.6. Use the new amplitude and the new phase to reconstruct the complex tensor, and then obtain the features corresponding to each multi-view image through inverse Fourier transform .
[0025] Furthermore, in step S2.1, the depth convolution uses depth convolution , where d is the dilation rate;
[0026] In step S2.2, the ratio is: ; For , the depth convolution is specifically: , respectively use depth convolution and depth convolution for processing;
[0027] In step S2.3, the pointwise convolution uses pointwise convolution;
[0028] Step S2.4 is specifically to obtain the weighted spectral data through the following formula:
[0029] ;
[0030] Among them, is a two-dimensional convolution layer with a Sigmoid activation function for calculating the weights of low-frequency features, It is a two-dimensional convolutional layer with a Sigmoid activation function for calculating the weights of high-frequency features. is the low-frequency feature. is the multi-scale feature a simplified representation;
[0031] In step S2.5, the modulation network consists of , LeakyReLU, and efficient channel attention ; respectively performing dynamic enhancement on them through the modulation network, and then adding residual connections respectively to obtain a new amplitude and a new phase , specifically:
[0032]
[0033] .
[0034] Furthermore, 1) in step S5 is specifically:
[0035] a. Respectively at the scale, reducing the enhanced feature of the reference image at the previous scale to obtain a reduced-dimensional feature , and then respectively performing sorting on the reduced-dimensional feature along the axis direction and the axis direction to obtain the corresponding features and , and recording the indices in the X-axis direction and the Y-axis direction, ; respectively performing layer normalization and one-dimensional convolution on the features and in the X-axis direction and the Y-axis direction, and then restoring the order according to the indices and to obtain the features and ;
[0036] Calculating the first cross-attention and the second cross-attention through the following formula:
[0037]
[0038]
[0039] where, represents adjusting the tensor shape, is the activation function, represents matrix dot product, T is the transpose, , , is the corresponding QKV matrix, , , is the corresponding QKV matrix;
[0040] Fuse the first cross-attention and the second cross-attention , then perform residual connection with the downsampled feature to obtain the fused feature of the reference image at the corresponding scale ;
[0041] b. At the scales of respectively, based on the fused feature , for each pixel in the reference image, use the spatial correlation between the pixels in the neighborhood and itself to guide depth adaptive propagation, and obtain 8 depth hypotheses at the current scale;
[0042] c. At the current scale, depth sample within the depth search range of the pixel to obtain depth hypotheses. Among them, is the number of depth hypotheses at the th scale. The depth sampling points are determined using the Gaussian quantile function. The depth hypotheses corresponding to depth sampling points are obtained through the following formula:
[0043]
[0044] where is the th depth hypothesis at the th scale, , is the optimized standard deviation, is the quantile function of the standard normal distribution, is the predicted depth at the th scale, represents the quantile function of the normal distribution with a mean of and a standard deviation of ; represents the total probability covered within the depth search range, and λ is a scalar parameter that controls the width of the sampling interval.
[0045] Furthermore, in step a, the convolutional kernel sizes of the one-dimensional convolution are and respectively;
[0046] Step b specifically is: For each pixel in the reference image , apply the dilated convolution to the fused feature to learn the offsets of the 8 pixels in the neighborhood centered on the pixel in the two directions of and . According to the offsets , apply bilinear interpolation to the predicted depth at the th scale to obtain 8 depth hypotheses at the current scale;
[0047] In step c, the depth search range of the pixel is obtained by the following method:
[0048] Based on the fused feature , obtain the normalized weights of the pixel and the 8 pixels in itsneighborhood, as well as the standard deviation of the predicted depth at the th scale, where , is the serial number of the pixel and the 8 pixels in its neighborhood;
[0049] Obtain the standard deviation of the offset pixel and its neighborhood pixels, and weighted to obtain the optimized standard deviation : ;
[0050] Calculate the depth search range R of the pixel through the optimized standard deviation: , where is a scalar parameter for controlling the width of the sampling interval, is the predicted depth at the th scale.
[0051] Furthermore, in step c, the normalized weight is specifically obtained by the following method: Process the fused feature through a 2D weight learning network with a shallow residual connection to learn the weight distribution of the pixel and the 8 pixels in its neighborhood, and then apply the Softmax operation to obtain the normalized weight ;
[0052] The standard deviation of the predicted depth at the Obtained by the following formula:
[0053]
[0054] Wherein, is the number of depth hypotheses at the th scale, is the th depth hypothesis at the th scale, is the probability that the depth is at the th scale, is the predicted depth at the th scale;
[0055] The standard deviation of the offset pixel and its neighboring pixels is interpolated based on the offset obtained in step b and the standard deviation .
[0056] Furthermore, 2) in step S5 is specifically:
[0057] The multi-scale enhanced features are respectively subjected to a differentiable homography transformation and distorted to the front parallel plane of the reference camera frustum corresponding to the reference image, and the transformed feature volume corresponding to the current scale is obtained. Define as the reference feature volume, as the source feature volume, and the transformed feature volume passes through a 3D weight learning network to generate the per-pixel similarity weight of the source feature volume relative to the reference feature volume at each scale; the transformed feature volume is evenly divided into groups along the channel direction, and the cost volume at the current scale is constructed by the following formula:
[0058]
[0059] Wherein, is the number of channels of the feature volume, is matrix dot product, is the simplified representation of , is the th group of features of the reference feature volume, is the rd group of features of the th source feature volume;
[0060] In step 3) of step S5, 2D regularization is used to process the cost volume at the current scale, specifically: in the depth prediction process corresponding to the current scale, the probability volume is upsampled and concatenated with the cost volume along the depth dimension, and then passed through a 2D U-Net with batch normalization and ReLU activation function; the Softmax function is used to normalize the processed feature volume to obtain the probability volume at the current scale.
[0061] Furthermore, step S4 is specifically as follows:
[0062] S4.1. For the lowest-scale feature at scale , sample depth hypotheses at equal intervals along the inverse depth interval:
[0063]
[0064] where is the preset maximum depth limit, and is the preset minimum depth limit;
[0065] S4.2. Use differentiable homography transformation to convert into the feature volume in the reference image view, and calculate the variance of the feature volume to construct the cost volume;
[0066] S4.3. Normalize the cost volume obtained in step along the depth direction to obtain the probability volume at scale , and weighted to obtain the predicted depth at the lowest scale:
[0067]
[0068] where represents the probability of the depth hypothesis .
[0069] At the same time, the present invention also provides a dynamic adaptive multi-view stereo reconstruction model based on dual-domain information fusion, which is used to implement the above-mentioned dynamic adaptive multi-view stereo reconstruction method based on dual-domain information fusion, and its special features are: it includes a multi-scale spatial domain, a frequency domain information fusion network and a multi-scale depth prediction network; the multi-scale spatial domain, frequency domain information fusion network includes a feature extraction module, a spatial domain detail enhancement module, a frequency domain dynamic aggregation module and a dual-domain information fusion module; the feature extraction module is used to perform multi-scale feature extraction on multi-view images to obtain multi-scale features ; The spatial detail enhancement module is used to perform deep convolution on multi-scale features to achieve spatial detail enhancement and output features ; The frequency domain dynamic aggregation module is used to analyze multi-scale features Perform dynamic aggregation and enhancement in the frequency domain and output features ; The dual-domain information fusion module is based on features ,feature and multi-scale features The multi-scale enhanced features are obtained by fusion of features; the multi-scale depth prediction network includes a depth initialization module, a dynamic range adaptive propagation module, a dynamic range uncertainty perception module, a feature body adaptive aggregation module and a cost body regularization module; the depth initialization module is used to determine Probability volume of reference image at different scales and predicted depth ; The dynamic range adaptive propagation module is used to construct a The dynamic range uncertainty perception module is used to construct a partial depth hypothesis for the reference image. Another part of the depth hypothesis at the scale of k=1,2,3; the feature body adaptive aggregation module is used to perform differentiable homography transformation on the multi-scale enhanced features at the scale of k=1,2,3 to obtain the transformed feature body, and then adaptively aggregate the transformed feature body to obtain the cost body at the current scale; the cost body regularization module is used to process the cost body at the scale of k=1,2,3 to obtain the corresponding probability body and predicted depth.
[0070] Furthermore, the feature extraction module includes 8 2D convolutional layers, each of which is followed by batch normalization and activation functions, where the convolution kernel size of the 3rd, 5th, and 7th convolutional layers is , the stride and padding are 2, and the convolution kernel size of the remaining convolutional layers is , stride and padding are 1.
[0071] Beneficial effects of the present invention:
[0072] 1. The dynamic adaptive multi-view stereo reconstruction method based on dual-domain information fusion in this invention adopts a coarse-to-fine framework to construct a multi-scale depth prediction network. The depth hypotheses of the lowest-scale features are sampled at equal intervals along the inverse depth interval. In the remaining stages, an uncertainty hypothesis mechanism based on dynamic range multi-scale adaptive propagation is used to learn the corresponding depth hypotheses, significantly improving the adaptability to complex geometric scenes, making the generated candidate depths more effective, and laying a foundation for the final high-quality predicted depth.
[0073] 2. In the multi-scale spatial and frequency domain information fusion network of the dynamic adaptive multi-view stereo reconstruction model based on dual-domain information fusion of the present invention, a spatial domain detail enhancement module and a frequency domain dynamic aggregation module are provided to respectively optimize the spatial domain details of the multi-scale features and perform frequency domain dynamic aggregation and enhancement, improving the ability to capture and express targets of different scales, and making full use of the characteristics of frequency domain information to help the multi-scale spatial and frequency domain information fusion network capture richer and more important features. Brief Description of the Drawings
[0074] Figure 1 is a schematic structural diagram of an embodiment of the dynamic adaptive multi-view stereo reconstruction model based on dual-domain information fusion of the present invention;
[0075] Figure 2 is a flowchart of an embodiment of the dynamic adaptive multi-view stereo reconstruction method based on dual-domain information fusion of the present invention;
[0076] Figure 3 In (a) to (h) in is a schematic diagram of the reconstructed point cloud image obtained by an embodiment of the dynamic adaptive multi-view stereo reconstruction method based on dual-domain information fusion of the present invention. Detailed Description of the Embodiment
[0077] The dynamic adaptive multi-view stereo reconstruction model based on dual-domain information fusion of the present invention includes a multi-scale spatial and frequency domain information fusion network and a multi-scale depth prediction network. As Figure 1 shown, the multi-scale spatial and frequency domain information fusion network includes a feature extraction module, a spatial domain detail enhancement module (SDDE), a frequency domain dynamic aggregation module (FDDA), and a dual-domain information fusion module (DDF). The multi-scale depth prediction network includes a depth initialization module, a dynamic range adaptive propagation module (DRAP), a dynamic range uncertainty perception module (UA), a feature volume adaptive aggregation module, and a cost volume regularization module.
[0078] Before using the above dynamic adaptive multi-view stereo reconstruction model based on dual-domain information fusion for reconstruction, the model needs to be trained. The present invention introduces a loss to constrain the difference between the predicted depth map and the ground truth depth map, thereby guiding the network training to obtain a trained dynamic adaptive multi-view stereo reconstruction model.
[0079] Loss: ;
[0080] Total loss:
[0081] Among them, represents the ground truth depth at scale , is the predicted depth at scale , represents the weight coefficient of each scale.
[0082] After training, the image to be reconstructed is input into the model, and the dynamic adaptive multi-view stereo reconstruction method based on dual-domain information fusion of the present invention is used for reconstruction, as shown in Figure 2 shown, which specifically includes the following steps:
[0083] Step 1, use the multi-scale spatial and frequency domain information fusion network for feature extraction.
[0084] 1.1. Obtain N multi-view images of the same scene , as well as the internal parameters and external parameters of the camera corresponding to the N multi-view images respectively, where 3 ≤ N ≤ 6, represents the serial number of the multi-view image, taking values from 0 to N - 1, and respectively represent the rotation component and translation component in the external camera parameters corresponding to the th multi-view image. Take one of the N multi-view images as the reference image, and the remaining N - 1 images as the source images, represents the reference image, respectively represent the N - 1 source images, and define that the camera corresponding to the reference image is the reference camera, and the cameras corresponding to the source images are the source cameras.
[0085] 1.2. Multi-scale feature extraction
[0086] Through the feature extraction module, multi-scale feature extraction is respectively performed on an input reference image and source images to obtain the corresponding multi-scale features , where represents the scale, . The feature extraction module contains 8 2D convolutional layers, and each convolutional layer is equipped with batch normalization and activation functions. Among them, the convolutional kernel sizes of the 3rd, 5th, and 7th convolutional layers are , the stride and padding are 2, and the convolutional kernel sizes of the remaining convolutional layers are , the stride and padding are 1.
[0087] 1.3. Use the spatial domain detail enhancement module for spatial domain detail enhancement;
[0088] 1.3.1. Use depth convolution to encode multi-order features;
[0089] Use depth convolution to process multi-scale features , obtaining intermediate features As shown in the following formula, where d is the dilation rate. For the sake of simplicity in subsequent steps, the subscripts and superscripts are omitted and abbreviated as . .
[0090]
[0091] Divide along the channels proportionally to perform a splitting operation to obtain 、 、 , where , are processed by depth convolution and respectively to obtain convolutional features , , remains unchanged;
[0092]
[0093]
[0094] 1.3.2. Concatenate , , along the channel direction and then pass them through a pointwise convolution to integrate the channel information, obtaining the features corresponding to each multi-view image output by the spatial domain detail enhancement module:
[0095]
[0096] In this embodiment, the intermediate feature is divided into three parts proportionally along the channel direction, and , are respectively applied with depth convolution using convolution kernels of different sizes and dilation rates, while is not processed. This way of dividing and processing channels can effectively avoid redundant calculations for all features and improve the calculation efficiency. Applying depth convolution with different dilation rates and convolution kernel sizes to , enhances the ability to capture and represent targets of different scales. In addition, It can be directly retained as the original feature, providing untransformed detailed information and avoiding the loss of detailed information or feature blurring caused by deep convolution operations.
[0097] 1.4. Use the frequency-domain dynamic aggregation module for dynamic aggregation and enhancement
[0098] 1.4.1. For multi-scale features Adaptive weighting is performed on the high-frequency and low-frequency features within the corresponding frequency range to obtain weighted spectral data corresponding to each multi-view image . Among them, the low frequency refers to the global structural contour, and the high frequency refers to the edge region and texture details. The low-frequency and high-frequency features are highlighted through adaptive weighting, and harmful and useless frequency information such as noise and background is suppressed. The low-frequency feature Is obtained through pooling operations with different window sizes on the multi-scale features . The weighted spectral data Is specifically shown as follows:
[0099]
[0100]
[0101] Among them, Is a two-dimensional convolutional layer with a Sigmoid activation function for calculating the weight of the low-frequency feature, Is a two-dimensional convolutional layer with a Sigmoid activation function for calculating the weight of the high-frequency feature, Represents the pooling operation. For the sake of simplicity, only the formula representation when pooling once is listed here.
[0102] The multi-scale pooling kernel in the present invention can capture information in different frequency ranges, fully consider the influence of different frequency features on the reconstruction quality, perform adaptive weighting on the low-frequency and high-frequency features respectively, highlight the frequency features related to the scene, and suppress the interference of useless frequencies such as noise and background on the reconstruction quality, which can improve the generalization ability of the multi-scale spatial and frequency-domain information fusion network for different scenes and enhance its robustness.
[0103] 1.4.2. Perform dynamic aggregation and enhancement on the weighted spectral data
[0104] The weighted spectral data Is transformed from the spatial domain to the frequency domain through Fourier transform, and the amplitude And phase Are calculated, and they are respectively dynamically enhanced through the same modulation network, and then residual connections are added respectively to obtain the new amplitude And the new phase , specifically shown as follows. Among them, the modulation network consists of , LeakyReLU and Efficient Channel Attention constitute
[0105]
[0106]
[0107] Use the new amplitude and the new phase to reconstruct the complex tensor, and then obtain the features corresponding to each multi-view image output by the frequency-domain dynamic aggregation module through the inverse Fourier transform .
[0108] The amplitude reflects the global energy distribution and is suitable for capturing overall features, while the phase contains local detail information, which helps to enhance the fine-grained feature expression. Dynamically enhancing and residual connecting them can make full use of the characteristics of the frequency-domain information to help the multi-scale spatial and frequency-domain information fusion network capture richer and more important features. In other embodiments of the present invention, the amplitude and the phase can also be dynamically enhanced through different modulation networks. The modulation network used in the present invention is a lightweight modulation network, which can dynamically adjust the weights of the spectral information, highlight the key frequency bands and suppress the useless frequencies, and enhance the robustness and discriminability of the features. Among them, setting the residual connection enables the frequency-domain dynamic aggregation module to learn the spectral feature changes more efficiently, and at the same time improves the gradient fluidity of the network, making the optimization process more stable.
[0109] 1.5. Dual-domain Information Interaction and Fusion
[0110] Through the dual-domain information fusion module, fuse the corresponding feature and feature through and operations to obtain the fused feature , thereby improving the model's ability to capture multi-scale frequency-domain details and global information:
[0111]
[0112] The dual-domain information fusion module upsamples the fused feature and adds it to the corresponding scale of to obtain the multi-scale enhanced feature of dual-domain information fusion .
[0113] Step 2, determine the probability volume of the multi-scale enhanced feature at the k = 0 scale .
[0114] 2.1. To allocate limited resources to more likely depth hypotheses, the depth sampling initialization module samples the lowest-scale features in the multi-scale enhanced features at equal intervals along the inverse depth interval to obtain depth hypotheses:
[0115]
[0116] Among them, is the corresponding serial number of the depth hypothesis, and are the preset maximum depth limit and minimum depth limit respectively. In this embodiment, and are set to 935mm and 425mm respectively, and takes the value of 48.
[0117] 2.2. Determine the cost volume at the lowest scale
[0118] Use differentiable homography transformation to transform into the feature volume in the reference image view;
[0119] For the feature volume calculate the variance to construct the cost volume :
[0120]
[0121] Among them, represents the feature volume mean, and is the simplified representation of .
[0122] 2.3. The depth sampling initialization module performs a normalization operation on the cost volume along the depth direction to obtain the probability volume at the scale, , and then weighted to obtain the predicted depth at the lowest scale:
[0123]
[0124] Among them, represents the probability of the depth hypothesis .
[0125] Steps 2.2 and 2.3 are as shown in the homography transformation and variance aggregation (HW&Variance) in Figure 1 .
[0126] Step three, determine the predicted depths at k = 1, 2, 3 scales The predicted depth corresponding to the scale It is obtained through the following method specifically:
[0127] 1) Determine the multi-scale enhanced features when k = 1, 2, 3 Depth hypothesis
[0128] The present invention proposes an uncertainty search method for dynamic range adaptive propagation, which uses the uncertainty hypothesis mechanism based on dynamic range multi-scale adaptive propagation. By aggregating cross-scale dynamic range information, it learns robust depth hypotheses for uncertainty regions.
[0129] a. Aggregation of dynamic range information
[0130] The dynamic range adaptive propagation module At the scale of, the enhanced feature of the previous scale of the reference image, that is, the enhanced feature of the scale of the reference image , through Dimensionality reduction to obtain the dimensionality-reduced feature , and then for the dimensionality-reduced feature Respectively along the axis direction (width direction of the image) and axis direction (height direction of the image) perform sorting according to the eigenvalues from large to small ( ), to obtain the corresponding features and , which are respectively expressed as: , and record the indexes in the X-axis direction and Y-axis direction , .
[0131] On the X-axis direction and Y-axis direction, respectively perform layer normalization and one-dimensional convolution on the above-obtained features and to achieve cross-scale dynamic range feature aggregation in two directions, and obtain the features and . The sizes of the convolution kernels used are respectively and , and use to represent layer normalization, then:
[0132]
[0133]
[0134] According to the indexes and restore the features and to obtain the features and , where Indicates the restoration operation by index:
[0135]
[0136]
[0137] Through the above sorting and index restoration operations, one-dimensional convolution can operate between similar pixels across regions rather than adjacent pixels, thereby effectively capturing and focusing on uncertain regions while maintaining geometric structure consistency.
[0138] Use the reshape operation on and to reshape and obtain the corresponding QKV matrices, which are respectively denoted as 、 、 and 、 、 , where Q, K, and V respectively represent the Query, Key, and Value feature vectors in the attention mechanism.
[0139] Finally, perform biaxial cross-attention, and calculate the first cross-attention through the following formula :
[0140] ;
[0141] where, represents adjusting the tensor shape, is the activation function, represents matrix dot product, and T is the transpose.
[0142] Similarly, calculate the second cross-attention through the following formula :
[0143] .
[0144] Fuse the two cross-attentions and add a residual connection to obtain the fused feature containing rich dynamic range global information of the reference image at the corresponding scale : .
[0145] The present invention sets biaxial cross-attention between the dynamic range enhancement features in two directions, which are respectively used to focus on the correlations in the horizontal and vertical directions. The biaxial cross-attention can better capture global information, enhance the modeling ability of the multi-scale depth prediction network for complex structures and textures, and significantly reduce the computational complexity compared with the conventional self-attention.
[0146] b. The dynamic range adaptive propagation module (DRAP) constructs depth hypotheses
[0147] The dynamic range adaptive propagation module, at the scales of respectively, for each pixel in the reference image , uses the spatial correlation between its neighboring pixels and itself to guide depth adaptive propagation. Specifically: Dilated convolution is introduced to act on the fused feature at the corresponding scale to learn the offsets of the 8 pixels in the neighborhood centered on the pixel in the and directions. According to the offset , bilinear interpolation is applied to the predicted depth at the previous scale to generate 8 depth hypotheses based on dynamic range adaptive propagation at the current scale.
[0148] In the present invention, dilated convolution is applied to the fused feature to calculate the adaptive offset, which can, while enhancing the receptive field, dynamically adjust the depth sampling position according to the scene, flexibly capture complex geometric structure and disparity change information, significantly improve the adaptability to complex geometric scenes, and lay a foundation for achieving high-quality depth prediction.
[0149] c. Constructing depth hypotheses through the dynamic range uncertainty awareness module (UA)
[0150] The dynamic range uncertainty awareness module, at the scales of respectively, processes the fused feature through an existing 2D weight learning network with shallow residual connections to learn the weight distribution of the pixel and the 8 pixels in its neighborhood, and then applies the Softmax operation to obtain the normalized weight , which is simply denoted as , where is the serial number of the pixel and the 8 pixels in its neighborhood.
[0151] According to the probability volume and the predicted depth at the previous scale, the standard deviation of the predicted depth of the pixel at the th scale is calculated. For simplicity of expression, the superscript representing the scale and the content in the brackets are omitted, and the standard deviation of the predicted depth of the pixel is denoted as , as shown in the following formula:
[0152]
[0153] Among them, is the number of depth hypotheses at the th scale, is the th depth hypothesis at the th scale, is the probability that the depth is at the th scale, is the predicted depth at the th scale.
[0154] Furthermore, according to the previously learned offset to adaptively propagate neighborhood coordinates, and interpolate to obtain the offset pixels and the standard deviation of its neighborhood pixels based on the standard deviation , and weighted to obtain the optimized standard deviation : .
[0155] Therefore, the present invention can, based on the fused features, adaptively assign different weights to the pixel and its neighborhood pixels. The normalized weights ensure that the contributions of each pixel are effectively balanced in the standard deviation calculation, thereby reducing the interference of noise or outliers on depth prediction. In addition, weak texture regions usually lack obvious features, and it is difficult to obtain accurate depth through traditional methods. This method can effectively supplement the regions with insufficient texture information by weighted combination of the features and standard deviation information of the pixel and its neighborhood pixels, and improve the reliability of depth prediction.
[0156] The dynamic range uncertainty perception module calculates the depth search range of the pixel by optimizing the standard deviation : , where is a scalar parameter that controls the width of the sampling interval.
[0157] At the current scale, depth sampling is performed within the depth search range of the pixel , and the number of samples is , where represents the number of depth hypotheses at the th scale. In this embodiment, = the number of depth hypotheses M 1 、M 2 、M 3They are set to 32, 16, and 16 respectively, and can be set and adjusted according to experience. The Gaussian quantile function (i.e., the inverse cumulative distribution function of the Gaussian distribution) is introduced to determine the depth sampling points. The sampling points generated in this way will be centered around and present a Gaussian distribution. The sampling is denser closer to the position of and sparser farther away from the position of . The depth hypotheses corresponding to depth sampling points are obtained through the following formula:
[0158] ;
[0159] where is the th depth hypothesis at the th scale, , is the quantile function of the standard normal distribution, represents the quantile function of a normal distribution with a mean of and a standard deviation of ; represents the total probability coverage within the depth search range .
[0160] The present invention introduces the Gaussian quantile function to generate depth sampling points, making the sampling points present a Gaussian distribution centered around the predicted depth, densely distributed in the region with higher possibility, and sparsely covering the low-probability region far from the center. The advantages of this method are as follows: effectively concentrating computing resources in the high-confidence range of depth estimation, improving computing efficiency; taking into account the coverage of the low-probability region, enhancing the robustness and global search ability of the present invention, thereby reducing the computing overhead while maintaining the accuracy.
[0161] 2) The feature volume adaptive aggregation module performs feature volume aggregation;
[0162] The feature volume adaptive aggregation module respectively warps the multi-scale enhanced features corresponding to the current scale to the front parallel plane of the reference camera frustum after differentiable homography transformation (HW), and obtains the transformed feature volume corresponding to the current scale. For the sake of simplicity, the transformed feature volume is denoted as , and it is defined that the transformed feature volume of is the reference feature volume, and
[0163]
[0164] where represents any pixel in the reference image, denote at the th multi-view image corresponding pixel; and represent the rotation and translation components in the extrinsic parameters of the reference camera, and respectively represent the rotation component and the translation component in the extrinsic parameters of the camera corresponding to the th multi-view image; is the intrinsic parameter of the reference camera, is the intrinsic parameter of the camera corresponding to the i-th multi-view image; is the rd scale and the th depth hypothesis.
[0165] To balance the computational efficiency and the matching quality, an adaptive group correlation (GC) strategy is applied to construct the cost volume. The present invention designs a lightweight 3D weight learning network for generating the per-pixel similarity weights of the source feature volume relative to the reference feature volume at the current scale , which helps to reduce the false matches caused by occlusion between views. Table 1 shows the specific parameters of the lightweight 3D weight learning network. The transformed feature volume is sequentially processed by three 3D convolutional layers with a stride of 1, and then the per-pixel similarity weights of each source feature volume to the reference feature volume are generated by the Softmax operation . The first two 3D convolutional layers have batch normalization and activation functions, and the third 3D convolutional layer has no batch normalization and activation function.
[0166] Table 1 Specific parameters of the lightweight 3D weight learning network
[0167]
[0168] Then, the transformed feature volume at the current scale is evenly divided into groups along the channel direction, and the cost volume aggregating the information of multiple feature volumes at the current scale is constructed through the following formula :
[0169] ;
[0170] wherein, is the number of channels of the feature volume, is the matrix dot product, is the simplified representation of , is the th group feature of the reference feature volume, takes values from 1 to G, is the th source feature volume's Group features.
[0171] 3) The cost volume regularization module with probability volume geometric prior compensation performs cost volume regularization;
[0172] During the depth prediction process of scales other than the lowest scale, due to the increase in the resolution of the cost volume, 3D convolution will bring extremely high video memory consumption. Therefore, a 2D regularization with probability volume geometric prior compensation is introduced. Through 2D regularization, the probability volume at the scale is embedded into the cost volume at the scale to enhance the internal representation ability of the cost volume. Specifically, during the depth prediction process corresponding to the current scale, the probability volume is upsampled and then concatenated with the cost volume along the depth dimension. After being processed by a 2D regularization network (U-Net network) with batch normalization and ReLU activation function, it not only realizes the lightweight of the model but also ensures good spatial information integration ability. Finally, the Softmax function is used to normalize the processed cost volume to obtain the probability volume at the current scale.
[0173] The predicted depth at the current scale of the reference image is obtained by weighted summation of the probability volume along the depth direction, as shown in the following formula:
[0174]
[0175] where represents the probability that the depth is .
[0176] Step four, use the publicly available Open3d code to convert the predicted depth at the scale into a three-dimensional point cloud file of the corresponding scene. Then use CloudCompare software to view the quality of the reconstructed point cloud image. As Figure 3 shown, eight reconstructed point cloud images obtained through this embodiment are presented.
Claims
1. A dynamic adaptive multi-view stereo reconstruction method based on dual-domain information fusion, characterized in that: The following steps are involved: S1, input N multi-view images of the same scene to obtain multi-scale features Wherein, k represents the scale, k = 0, 1, 2, 3, i represents the sequence number of the multi-view image, ranging from 0 to N-1, one of the N multi-view images is used as the reference image, i = 0 represents the reference image, and the remaining N-1 images are used as the source images; S2, for multi-scale features Perform spatial detail enhancement and output the features F corresponding to each multi-view image s ; For multi-scale features Perform dynamic aggregation and enhancement in the frequency domain and output the features F corresponding to each multi-view image f ; S3, the corresponding feature F s With feature F f Fusion obtains fusion feature F m , F m Upsampling and corresponding multi-scale features Add together to get multi-scale enhanced features S4, determine the probability volume P of the reference image at scale k = 0 0 and its corresponding predicted depth D 0 ; S5, determine the predicted depth D at scale k = 1, 2, 3 k ; The predicted depth D of the corresponding scale k Specifically obtained through the following methods: 1) Based on multi-scale enhanced features The corresponding depth hypothesis is determined using the uncertainty hypothesis mechanism based on dynamic range multi-scale adaptive propagation; specifically: a. At scales k=1, 2, and 3, the previous scale feature of the reference image is enhanced. Dimensionality reduction obtains dimensionality reduction feature S k-1 , and then reduce the dimension feature S k-1 Sorting is performed along the X-axis and Y-axis directions respectively to obtain the corresponding features and And record the index idex in the X-axis direction and the Y-axis direction X ,idex Y ; In the X-axis direction and Y-axis direction, the features and Perform layer normalization and one-dimensional convolution, and then according to the index idex X and idex Y Restore the order to get feature F X and F Y ; Calculate the first cross attention Attn1 and the second cross attention Attn2 and merge them, and then combine them with the dimension reduction feature S k-1 Residual connection to obtain the fusion features of the reference image at the corresponding scale b. At scales k=1, 2, and 3, based on fusion features For each pixel p in the reference image, the spatial correlation between the pixels in the 3×3 neighborhood and itself is used to guide the depth adaptive propagation, and 8 depth hypotheses of the current scale are obtained; c. At the current scale, the depth of the pixel p is sampled within the depth search range to obtain M k -8 depth sampling points, among which M k is the number of depth hypotheses at the kth scale. The depth sampling points are determined using the Gaussian quantile function and then calculated to obtain the value of M k - Depth hypothesis corresponding to 8 depth sampling points; 2) Multi-scale enhanced features After differentiable homography transformation, it is warped to the front parallel plane of the reference camera frustum corresponding to the reference image, and the transformed feature volume corresponding to the current scale is obtained Definition: the transformed feature body with i=0 is the reference feature body, the transformed feature body with i=1,2,…,N-1 is the source feature body, and the transformed feature body After the 3D weight learning network, the pixel-by-pixel similarity weights of the source feature body relative to the reference feature body are generated at each scale. The transformed feature bodies Divide into G groups along the channel direction and construct the cost volume C at the current scale k ; 3) In the depth prediction process corresponding to the current scale k, the probability volume P k-1 After upsampling and cost volume C k Cascade along the depth dimension, pass through 2D U-Net with batch normalization and ReLU activation function; use Softmax function to normalize the processed feature volume to obtain the corresponding probability volume P k ; For probability body P k The weighted sum along the depth direction is used to obtain the predicted depth D of the current scale. k ; S6, the predicted depth D at scale k = 3 3 Convert it into a three-dimensional point cloud file of the corresponding scene, and complete dynamic adaptive multi-view stereo reconstruction based on dual-domain information fusion.
2. The dynamic adaptive multi-view stereoscopic reconstruction method based on dual-domain information fusion according to claim 1, characterized in that: Step S2 is specifically as follows: S2.
1. Multi-scale features Perform deep convolution to obtain the intermediate feature X0; S2.
2. Split the intermediate feature X0 along the channel in proportion to obtain X1, X2, and X3, and perform deep convolution on X1 and X2 to obtain convolution features X 1′ and the convolutional features X 2′ ; S2.
3. X 1′ , X 2′ , X3 are cascaded along the channel direction and the channel information is integrated through point-by-point convolution to obtain the features F corresponding to each multi-view image. s ; S2.
4. Multi-scale features The features within the corresponding frequency range are adaptively weighted according to the frequency, highlighting the required frequencies and suppressing the useless frequencies, and the weighted spectrum data F corresponding to each multi-view image is obtained. a ; S2.
5. The weighted spectrum data F a The spatial domain is converted to the frequency domain through Fourier transform, the amplitude mag and phase pha are calculated, and they are dynamically enhanced through the modulation network respectively, and then residual connections are added to obtain the new amplitude mag′ and new phase pha′ respectively; S2.
6. Use the new amplitude mag′ and the new phase pha′ to reconstruct the complex tensor, and then use the inverse Fourier transform to obtain the features F corresponding to each multi-view image f .
3. The dynamic adaptive multi-view stereoscopic reconstruction method based on dual-domain information fusion according to claim 2, characterized in that: In step S2.1, the depth convolution adopts the depth convolution DWConv 3×3,d=1 , where d is the expansion rate; In step S2.2, the ratio is: ratio = {0.25, 0.25, 0.5}; the depth convolution of X1 and X2 is specifically: X1 and X2 are respectively subjected to the depth convolution DWConv 5×5,d=2 And deep convolution DWConv 7×7,d=3 to process; In step S2.3, the point-by-point convolution adopts 1×1 point-by-point convolution; Step S2.4 is specifically to obtain the weighted spectrum data F by the following formula: a : F a =Freq_Weight L (F)·Low_Freq(F)+Freq_Weight H (F)·(F-Low_Freq(F)); Among them, Freq_Weight L Freq_Weight is a two-dimensional convolutional layer with a Sigmoid activation function for calculating the weight of low-frequency features. H is a two-dimensional convolutional layer with a Sigmoid activation function for calculating the weight of high-frequency features, Low_Freq(F) is the low-frequency feature, and F is the multi-scale feature A simplified representation of ; In step S2.5, the modulation network consists of Conv 1×1 , LeakyReLU and efficient channel attention ECA; the dynamic enhancement is performed by the modulation network respectively, and then the residual connection is added to obtain the new amplitude mag′ and the new phase pha′, specifically: mag′=Conv 1×1 (ECA(LeakyReLU(Conv 1×1 (mag))))+mag pha′=Conv 1×1 (ECA(LeakyReLU(Conv 1×1 (pha))))+pha。 4. The dynamic adaptive multi-view stereoscopic reconstruction method based on dual-domain information fusion according to any one of claims 1 to 3, characterized in that: In step 1) of step S5: The first cross attention Atten1 and the second cross attention Atten2 are calculated by the following formula: Attn1=reshape(softmax(F YQ F XK T )⊙F XV ) Attn2=reshape(softmax(F XQ F YK T )⊙F YV ) Among them, reshape means adjusting the shape of the tensor, softmax is the activation function, ⊙ means matrix dot multiplication, T is transpose, F XQ 、F XK 、F XV F X The corresponding QKV matrix, F YQ 、F YK 、F YV F Y The corresponding QKV matrix; With M k The depth assumption corresponding to the 8 depth sampling points is calculated by the following formula: in, is the j-th depth hypothesis at the k-th scale, To optimize the standard deviation, Φ -1 () is the quantile function of the standard normal distribution, D k-1 is the predicted depth at the k-1th scale, F -1 () indicates the mean is D k-1 The standard deviation is Quantile function of the normal distribution; represents the total probability of coverage within the depth search range, and λ is a scalar parameter that controls the width of the sampling interval.
5. The method for dynamic adaptive multi-view stereoscopic reconstruction based on dual-domain information fusion according to claim 4, characterized in that: In step a, the convolution kernel sizes of the one-dimensional convolution are 1×3 and 3×1 respectively; Step b is as follows: for each pixel p in the reference image, use dilated convolution to act on the fusion feature We learn the offset of 8 pixels in the 3×3 neighborhood centered on pixel p in both the X and Y directions. According to the offset The predicted depth D at the k-1th scale k-1 Apply bilinear interpolation to obtain 8 depth hypotheses for the current scale; In step c, the depth search range of the pixel p is obtained by the following method: Based on fusion features Get the normalized weights of pixel p and its 8 pixels in its 3×3 neighborhood And the standard deviation σ of the predicted depth at the k-1th scale, where s is the sequence number of the pixel p and its 8 pixels in the 3×3 neighborhood; Get the standard deviation σ of the offset pixel p and its neighboring pixels s , weighted to obtain the optimal standard deviation The depth search range R of pixel p is calculated by optimizing the standard deviation: Among them, λ is a scalar parameter that controls the width of the sampling interval, D k-1 is the predicted depth at the k-1th scale.
6. The method for dynamic adaptive multi-view stereoscopic reconstruction based on dual-domain information fusion according to claim 5, characterized in that: In step c, the normalized weight The specific method is as follows: After being processed by a shallow 2D weight learning network with residual connections, the weight distribution of pixel p and its 8 pixels in its 3×3 neighborhood is learned, and then the Softmax operation is applied to obtain the normalized weight The standard deviation σ of the predicted depth at the k-1th scale is obtained by the following formula: Among them, M k-1 is the number of depth hypotheses at the k-1th scale, is the jth depth hypothesis at the k-1th scale, P j The depth at the k-1th scale is The probability of D k-1 is the predicted depth at the k-1th scale; The standard deviation σ of the pixel p after the shift and its neighboring pixels s According to the offset obtained in step b It is obtained by interpolation based on the standard deviation σ.
7. The dynamic adaptive multi-view stereoscopic reconstruction method based on dual-domain information fusion according to claim 6, characterized in that: In step S5, in step 2), the cost body C k It is obtained by the following formula: Among them, C1 is the number of feature volume channels, ⊙ is the matrix dot product, α i for A simplified representation of is the g-th group of features of the reference feature body, is the g-th set of features of the i-th source feature body.
8. The dynamic adaptive multi-view stereoscopic reconstruction method based on dual-domain information fusion according to claim 7, characterized in that: Step S4 is specifically as follows: S4.
1. is the lowest scale feature at scale k = 0 Sample M0 depth hypotheses at equal intervals along the inverse depth interval Among them, d max is the preset maximum depth limit, d min is the pre-set minimum depth limit; S4.
2. Using differentiable homography to transform Transformed into the feature volume under the reference image perspective For feature body Find the variance to construct the cost volume; S4.
3. Normalize the cost volume at the scale obtained in step S4.2 along the depth direction to obtain the probability volume P at the scale k = 0 0 , Weighted to obtain the predicted depth D at the lowest scale 0 : in, Depth Hypothesis probability.
9. A dynamic adaptive multi-view stereoscopic reconstruction model based on dual-domain information fusion, used to implement the dynamic adaptive multi-view stereoscopic reconstruction method based on dual-domain information fusion according to any one of claims 1 to 8, characterized in that: Including multi-scale spatial domain, frequency domain information fusion network and multi-scale depth prediction network; The multi-scale spatial domain and frequency domain information fusion network includes a feature extraction module, a spatial domain detail enhancement module, a frequency domain dynamic aggregation module and a dual-domain information fusion module; the feature extraction module is used to extract multi-scale features from multi-view images to obtain multi-scale features. The spatial detail enhancement module is used to perform deep convolution on multi-scale features to achieve spatial detail enhancement and output feature F s ; The frequency domain dynamic aggregation module is used to analyze multi-scale features Perform dynamic aggregation and enhancement in the frequency domain and output feature F f ; The dual-domain information fusion module is based on feature F s , Feature F f and multi-scale features The feature fusion of is used to obtain multi-scale enhanced features; The multi-scale depth prediction network includes a depth initialization module, a dynamic range adaptive propagation module, a dynamic range uncertainty perception module, a feature volume adaptive aggregation module, and a cost volume regularization module; The depth initialization module is used to determine the probability volume P of the reference image at scale k = 0 0 and predicted depth D 0 ; The dynamic range adaptive propagation module is used to construct partial depth hypotheses at scales k=1, 2, and 3 for the reference image; The dynamic range uncertainty perception module is used to construct another part of the depth hypothesis at the scale of k=1,2,3 for the reference image; The feature volume adaptive aggregation module is used to perform differentiable homography transformation on the multi-scale enhanced features at scales k = 1, 2, and 3 to obtain the transformed feature volume, and then adaptively aggregate the transformed feature volume to obtain the cost volume at the current scale; The cost body regularization module is used to process the cost body at the scale of k = 1, 2, 3 to obtain the corresponding probability body and predicted depth.
10. The dynamic adaptive multi-view stereo reconstruction model based on dual-domain information fusion according to claim 9, characterized in that: The feature extraction module contains 8 2D convolutional layers, each of which is followed by batch normalization and activation functions. The convolution kernel size of the 3rd, 5th, and 7th convolutional layers is 5×5, with a step size and padding of 2, and the convolution kernel size of the remaining convolutional layers is 3×3, with a step size and padding of 1.
Citation Information
Patent Citations
Remote sensing image panchromatic sharpening method based on multi-scale double-domain information fusion technology
CN116402700A
Low-light image enhancement method based on frequency domain and spatial domain perception
CN118674628A