Three-dimensional reconstruction method based on guided filtering and Mama geometric feature fusion
By combining adaptive guided filtering and global attention regularization with the Mamba geometric feature fusion network, the problems of missing global context information, error accumulation, and noise interference in 3D reconstruction are solved, achieving high-precision and high-robust 3D reconstruction.
Patent Information
- Application Number
- CN202511456207.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-13
AI Technical Summary
Existing 3D reconstruction methods suffer from problems such as lack of global context information in cost volume regularization, accumulation of errors in cross-stage feature fusion, noise interference, and structural degradation, resulting in insufficient reconstruction accuracy and robustness.
An adaptive guided filtering network and a global attention regularization module are combined with a Mamba geometric feature fusion network. The cost volume is processed by adaptive filtering and regularization to dynamically evaluate the confidence of depth information, gradually optimize the depth estimation, and selectively fuse high-confidence geometric features in conjunction with the Mamba network.
It significantly improves the accuracy and robustness of 3D reconstruction, effectively preventing error accumulation in complex scenes such as weak textures and occlusion, maintaining geometric edge structure, and improving point cloud quality and density.
Smart Images

Figure CN120912796A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional reconstruction, and particularly discloses a three-dimensional reconstruction method based on guided filtering and Mamba geometric feature fusion. BACKGROUND
[0002] Three-dimensional reconstruction (3D Reconstruction) refers to a technical process of recovering three-dimensional geometric structure and appearance information of an object or a scene from two-dimensional images. In a traditional three-dimensional reconstruction method, a multi-view stereo (MVS) network is widely used, in which a cost volume is used as a core data structure for encoding matching costs or uncertainties between pixels in different view images. However, the existing MVS-based scheme still has the following key problems: 1. Lack of global context information in cost volume regularization: In the initial stage of the cascaded MVS process, a cost volume needs to be constructed to represent the matching cost under different depth hypotheses, and a regularization network is used to optimize the cost volume. The existing method generally uses a three-dimensional convolutional neural network (3D CNN) to complete this regularization process. However, the receptive field of the three-dimensional convolutional neural network has a local limitation, and it can only perceive the information of the adjacent region when processing the cost volume, which makes it difficult to effectively capture the global geometric context and long-distance dependency relationship of the overall scene, thereby restricting the consistency and accuracy of depth estimation.
[0003] 2. Error accumulation and amplification in cross-stage feature fusion: The cascaded MVS architecture usually uses the depth prediction result of the previous coarse stage to guide the depth estimation of the subsequent finer stage. In this process, the existing method often fuses the geometric information (such as the up-sampled depth map) of the previous stage and the image features of the current stage through simple feature concatenation or addition, etc. This fusion strategy lacks an evaluation mechanism for the reliability of the depth information of the previous stage. If an estimation bias occurs in weak texture areas and the like, the error information will be forced to be transmitted to the next stage, resulting in the accumulation and even amplification of errors in the cascading process, which seriously reduces the final reconstruction accuracy.
[0004] 3. Noise interference and structure degradation of cost volume in cascading transmission: In the cascade processing of MVS, the cost volume generated by each stage often contains noise, especially in challenging areas such as low texture, occlusion or non-Lambertian surface. These noises can form false low cost values, interfere with the depth probability distribution, and thus cause noise and artifacts in the depth map. The prior art lacks a direct optimization mechanism for the cost volume across stages, and cannot effectively "purify" the cost volume itself. Noise and structural uncertainty are transmitted to subsequent stages through the depth estimation process, affecting the setting of the depth hypothesis interval and the construction of the new cost volume, and thus causing error propagation. Although a general filtering method can be introduced to suppress noise, it often indiscriminately smooths all information, resulting in the loss of key geometric details such as edges and thin-walled structures, causing structural degradation.
[0005] Therefore, how to provide a three-dimensional reconstruction method based on guided filtering and Mamba geometric feature fusion to improve the accuracy and robustness of three-dimensional reconstruction has become a technical problem to be solved. SUMMARY
[0006] The technical problem to be solved by the present application is to provide a three-dimensional reconstruction method based on guided filtering and Mamba geometric feature fusion to improve the accuracy and robustness of three-dimensional reconstruction.
[0007] The present application is implemented as follows: a three-dimensional reconstruction method based on guided filtering and Mamba geometric feature fusion, comprising the following steps: Step S1, acquiring multiple-view RGB images, calibrating the intrinsic and extrinsic parameters of the cameras for acquiring the RGB images to obtain a projection matrix and an intrinsic parameter matrix, and obtaining a reference map for three-dimensional reconstruction; Step S2, inputting each of the RGB images into a feature extraction network for feature extraction to obtain image features; Step S3, obtaining a plurality of depth hypothesis planes by inversely and uniformly sampling the reference map, performing homography transformation on each of the image features, and aggregating the depth hypothesis planes to obtain a cost volume; Step S4, performing a filtering operation on the cost volume through an adaptive guided filtering network; Step S5, performing a regularization operation on the cost volume after the filtering operation through a regularization network; Step S6, performing exponential normalization on the channel dimension of the cost volume after the regularization operation to obtain a probability volume, and constructing a depth map and a confidence map based on the probability volume; Step S7, after performing frequency domain filtering and upsampling on the depth map and the confidence map, inputting them into an adaptive geometric feature fusion network for feature fusion to obtain a view depth; Step S8, performing a three-dimensional reconstruction operation based on each of the view depths, the projection matrix and the intrinsic parameter matrix to obtain a dense point cloud.
[0008] Further, the filtering process of the adaptive guidance filtering network in step S4 is specifically: In the 0th stage, the reference map is used as the guidance map; in the subsequent stages, the depth map generated in the previous stage is used as the guidance map; The horizontal gradient and the vertical gradient of the guidance map are calculated: ; ; wherein, represents the horizontal gradient; represents the vertical gradient; represents the guidance map; represents the Sobel convolution kernel in the horizontal direction; represents the Sobel convolution kernel in the vertical direction; The edge-preserving weight is calculated based on the horizontal gradient and the vertical gradient: ; wherein, represents the edge-preserving weight; represents the natural exponential function; represents the edge-preserving strength control hyperparameter; The enhanced guidance representation is constructed based on the edge-preserving weight: ; wherein, represents the enhanced guidance representation; for each pixel of the cost volume slice of the cost volume , an adaptive window is constructed: ; wherein, represents the adaptive window, i.e. the window centered at ; represents the pixel coordinate of the RGB image; represents the gray value of the guidance map at ; represents the gray value of the guidance map at ; represents the continuity threshold; represents the maximum window radius; represents the maximum value; represents the maximum value of the horizontal distance between x and , and the maximum value of the vertical distance between y and ; The linear regression coefficient is calculated within the adaptive window: ; wherein, and denote linear regression coefficients; denotes a covariance function; denotes a variance function; denotes the k-th adaptive window; denotes the intercept of the guidance map within ; denotes the intercept of the cost volume slice p within ; denotes a regularization parameter to prevent the denominator from being zero; denotes the mean of the cost volume slice p within ; denotes the mean of the guidance map within ; performing slice-wise filtering along the depth dimension for each of the cost volumes based on the enhanced guidance representation: ; wherein, denotes the filtered cost volume slice; denotes an adaptive guidance filtering operation based on the linear regression coefficients; denotes the d-th depth hypothesis corresponding cost volume slice; performing weighted fusion of the filtered cost volume slices: ; wherein, denotes the weighted fused cost volume slice; denotes an adaptive easing factor.
[0009] Further, in the step S5, the regularization network is constructed based on a global attention regularization module and a three-dimensional U-shaped module; the global attention regularization module is used to perform a regularization operation on the cost volume of the 0th stage, and the three-dimensional U-shaped module is used to perform a regularization operation on the cost volumes of the remaining stages.
[0010] Further, the regularization process of the global attention regularization module is specifically: generating a three-dimensional coordinate grid matching the dimensions of the cost volume and normalizing, independently applying sinusoidal position encoding to each coordinate axis of the normalized three-dimensional coordinate grid to obtain encoding features, projecting the encoding features through a convolution layer to obtain projection features, adding the projection features to the cost volume to obtain a projected cost volume, inputting the projected cost volume into a stride 3D convolution layer for down-sampling to obtain a compact cost volume; The compact cost volume is flattened along the spatial and depth dimensions to obtain a one-dimensional sequence. The one-dimensional sequence is then input into several Transformer encoders employing a global attention mechanism to calculate the global attention value. The global attention value is reshaped into a three-dimensional feature volume, and then the three-dimensional feature volume is upsampled through a 3D transposed convolutional layer to restore the original depth and spatial resolution, resulting in a regularized cost volume.
[0011] Furthermore, the formula for the adaptive geometric feature fusion network is: ; in, The depth map is represented by the first... The viewing depth corresponding to each pixel; Indicates a Mamba network; This represents the visual features extracted from the reference image; Indicates a splicing operation; Indicates reliability weight; It represents the Hadamardi (or Hadama) stack; Represents a geometric branching network; Indicates reference image; This represents the depth map obtained from the previous upsampling stage.
[0012] The advantages of this invention are: 1. By collecting multi-view RGB images, the camera for collecting RGB images is calibrated to obtain the projection matrix and the intrinsic matrix, and the reference graph for three-dimensional reconstruction is obtained; then the RGB images are input into the feature extraction network to obtain the image features, and a plurality of depth hypothesis planes are obtained by inversely and uniformly sampling the reference graph, the homography transformation is performed on the image features, and the depth hypothesis planes are aggregated to obtain a cost volume; then the adaptive guided filtering network is used to filter the cost volume, the regularized network is used to regularize the filtered cost volume, the channel dimension of the regularized cost volume is exponentially normalized to obtain a probability volume, the depth map and the confidence map are constructed based on the probability volume, and after the frequency domain filtering and up-sampling of the depth map and the confidence map, the adaptive geometric feature fusion network is input to fuse the features to obtain the perspective depth, and finally the three-dimensional reconstruction is performed based on the perspective depth, the projection matrix and the intrinsic matrix to obtain the dense point cloud; that is, the adaptive guided filtering network is used to suppress the noise transmission of the cost volume and maintain the geometric edge structure, the global attention regular module is used to capture the long distance dependence to enhance the context consistency of the cost volume, the reliability weight mechanism in the adaptive geometric feature fusion network is used to dynamically evaluate the confidence of the cross-stage depth information, and the Mamba network is used to selectively fuse the high-confidence geometric features, so as to effectively block the error accumulation in the complex scene such as weak texture and occlusion, and greatly improve the accuracy and robustness of the three-dimensional reconstruction.
[0013] 2. By introducing the adaptive guided filtering network, the accuracy and stability of the filtering are effectively improved by dynamically adjusting the guide map (using the reference graph in the initial stage and the depth map of the previous stage in the subsequent stage) and combining the edge preservation weight and the adaptive window; the edge preservation weight is based on the gradient calculated by the Sobel operator, the edge intensity is controlled by the exponential function, the texture details are protected in the filtering process, the edge blur in the reconstruction is reduced, and the point cloud quality is improved; the window size is adaptively adjusted according to the pixel gray continuity and the spatial distance, compared with the fixed window filtering, the scene light change and the occlusion problem can be better handled, and the robustness to complex environment is enhanced; the depth map of the previous stage is used as the guide map in the subsequent stage, realizing adaptive iterative optimization, gradually refining the depth estimation, reducing the noise influence, improving the reconstruction consistency, finally solving the problems of losing details and over-smoothing in the traditional guided filtering in the three-dimensional reconstruction, and improving the generalization ability and practicality.
[0014] 3. The combination of global attention regularization module and 3D U-shaped module is used in the regularization network to optimize the feature representation of the cost volume; the global attention regularization module processes one-dimensional sequences through the encoding of three-dimensional coordinate grids and the Transformer encoder to capture global context dependencies (such as the correlation between depth dimension and spatial dimension), improve the consistency of features, and reduce the voids or distortions in reconstruction; the 3D U-shaped module is used in the subsequent stage to effectively handle local details and multi-scale features by combining down-sampling and up-sampling operations, improving the resolution recovery capability of the cost volume; the module design avoids redundant calculations (such as down-sampling compact cost volume), combines convolution and transposed convolution, reduces the computational overhead while maintaining accuracy, and is suitable for real-time 3D reconstruction applications; this dual-module architecture integrates global and local regularization, overcoming the limitations of single regularization methods (such as pure CNN) in balancing global consistency and local details, and improving the density and accuracy of reconstruction.
[0015] 4. The adaptive geometric feature fusion network uses Mamba network to combine visual features and geometric features, achieving efficient feature fusion; as a state space model, Mamba has linear computational complexity and long sequence processing capability compared to traditional RNN or Transformer, which can efficiently fuse multi-view features and improve the efficiency of depth map estimation; by combining reliability weights through concatenation operation (⊕) and Hadamard product (⊙), the geometric branch output is dynamically weighted, enhancing the robustness to unreliable areas (such as occlusion or low texture) and reducing reconstruction errors; frequency domain filtering and up-sampling are performed before fusion to preprocess the depth map, ensuring the quality of input features and further improving the accuracy of view depth; this fusion mechanism solves the problem of insufficient fusion of visual and geometric features in traditional methods, and significantly improves the accuracy and speed of depth estimation by combining the advanced architecture of Mamba.
[0016] 5. The complete end-to-end flow is formed from image acquisition to point cloud generation, and the key steps are designed iteratively; integrating feature extraction, cost volume construction, filtering, regularization, and fusion reduces manual intervention and has high automation, making it suitable for large-scale 3D reconstruction scenarios; the outputs of previous stages (such as depth maps as guide maps) are used in multiple stages (such as filtering and fusion) to achieve progressive optimization, gradually improving the quality of depth maps and enhancing the adaptability of the method to noise and view changes; the probability volume is obtained by exponentially normalizing the cost volume, combined with the confidence map, to provide reliable depth uncertainty estimation, improving the density and reliability of the point cloud; this flow design improves the robustness and scalability of the method, reduces the risk of error accumulation compared to independent algorithms in different stages, and is easy to integrate into existing systems.
[0017] 6, Based on standard RGB images and camera parameters (intrinsic / extrinsic), without special hardware (such as depth sensors), strong compatibility; inverse depth uniform sampling and homography transformation simplify the cost volume construction, easy to implement; through adaptive filtering and Mamba fusion, it is possible to surpass traditional multi-view stereo (MVS) methods in accuracy (edge preservation) and speed (Mamba efficiency), suitable for high dynamic scenes.
[0018] 7, By innovatively integrating adaptive guided filtering mechanism, efficient regularization network and Mamba geometric feature fusion framework, through dynamic adjustment of guided image (reference image→depth map iterative optimization) combined with edge preservation weight and adaptive window, the reconstruction edge detail preservation ability and scene adaptability are significantly improved; by using the dual-path design of global attention regularization module and three-dimensional U-shaped module, the global consistent features and local detail recovery are balanced, and the computational overhead is reduced; by means of the efficient long sequence processing capability of Mamba network, the visual features and weighted geometric information are fused, the depth estimation problem of occlusion and low texture area is effectively solved, and finally the end-to-end high-precision and high-robustness three-dimensional reconstruction from multi-view RGB images to dense point cloud is realized. BRIEF DESCRIPTION OF DRAWINGS
[0019] The application will be further described below with reference to the accompanying drawings and embodiments.
[0020] Figure 1 is a flowchart of a three-dimensional reconstruction method based on guided filtering and Mamba geometric feature fusion according to the application.
[0021] Figure 2 is a flowchart of the application.
[0022] Figure 3 is a schematic diagram of the adaptive guided filtering network according to the application.
[0023] Figure 4 is a schematic diagram of the global attention regularization module according to the application.
[0024] Figure 5 is a schematic diagram of the adaptive geometric feature fusion network according to the application.
[0025] Figure 6 is a comparison diagram before and after filtering according to the application.
[0026] Figure 7 is a comparison diagram of the application and traditional methods on the DTU dataset.
[0027] Figure 8 is a three-dimensional reconstruction diagram of the application on the Tanks&Temples dataset. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the present application have the following general idea: the adaptive guided filtering network is used to suppress the transmission of cost volume noise and maintain the geometric edge structure, the global attention regular module is used to capture long-distance dependence to enhance the context consistency of the cost volume, the reliability weight mechanism in the adaptive geometric feature fusion network is used to dynamically evaluate the cross-stage depth information confidence, the Mamba network is used to selectively fuse high-confidence geometric features, so as to effectively block error accumulation in complex scenes such as weak texture and occlusion, and thus the accuracy and robustness of three-dimensional reconstruction are improved.
[0029] Please refer to Figures 1 to 8 The preferred embodiment of the three-dimensional reconstruction method based on guided filtering and Mamba geometric feature fusion of the present application comprises the following steps: Step S1, acquiring multiple-view RGB images, calibrating the intrinsic and extrinsic parameters of the cameras for acquiring the RGB images to obtain a projection matrix and an intrinsic parameter matrix, and obtaining a reference image for three-dimensional reconstruction; Step S2, inputting each of the RGB images into a feature extraction network for feature extraction to obtain image features; in the feature extraction process, a feature pyramid network is used in the 0th stage to obtain coarse features, and an adaptive geometric feature fusion network based on confidence guidance is used in the subsequent stages to fuse geometric information with fine features of the visual features of the current stage; Step S3, obtaining a plurality of depth hypothesis planes by inversely and uniformly sampling the reference image, performing homography transformation on each of the image features, and aggregating the depth hypothesis planes to obtain a cost volume; Step S4, performing a filtering operation on the cost volume by an adaptive guided filtering network (AGF), and gently smoothing the cost volume along each depth hypothesis plane with a gentle coefficient to suppress noise and preserve edges; Step S5, performing a regularization operation on the cost volume after the filtering operation by a regularization network (AttentionCostReg); Step S6, performing exponential normalization on the channel dimension of the cost volume after the regularization operation to obtain a probability volume, constructing a depth map and a confidence map based on the probability volume; the calculation formula of the probability volume is: ; Among them, P represents the probability volume; P represents the cost volume after the regularization operation; Step S7, after performing frequency domain filtering and upsampling on the depth map and the confidence map, inputting the adaptive geometric feature fusion network (GeoMambaFusion) for feature fusion to obtain a perspective depth; and using the depth map of the previous stage to adaptively guide the filtering of the cost volume; Step S8, based on each of the depth of view, projection matrix and intrinsic matrix, performing a three-dimensional reconstruction operation to obtain a dense point cloud (fusion using open3d library).
[0030] The application adopts a multi-stage cascaded MVS architecture to gradually refine depth estimation at different resolutions; in the initial stage, a global attention regular module is used instead of a traditional 3D CNN to better capture global context information; in all stages, an adaptive guided filtering network is introduced to refine the cost volume.
[0031] The filtering process of the adaptive guided filtering network in step S4 is specifically: In the 0th stage, the reference image is used as the guide image; in the subsequent stage, the depth map generated in the previous stage is used as the guide image; in the initial stage, due to the lack of reliable depth prior, the reference image is used as the guide image; in the subsequent stage, the more reliable depth map generated in the previous stage is used as the guide image.
[0032] The horizontal gradient and the vertical gradient of the guide image are calculated: ; ; Wherein, represents the horizontal gradient; represents the vertical gradient; represents the guide image, R is a real number, B is the batch size, is the number of guide image channels, H is the height of the guide image, and W is the width of the guide image; represents the Sobel convolution kernel in the horizontal direction; represents the Sobel convolution kernel in the vertical direction; The edge preservation weight is calculated based on the horizontal gradient and the vertical gradient: ; Wherein, represents the edge preservation weight; represents the natural exponential function; represents the edge preservation strength control hyperparameter; The enhanced guide representation is constructed based on the edge preservation weight: ; Wherein, represents the enhanced guide representation; is each pixel of the cost volume slice (a two-dimensional plane divided along the depth dimension, each slice corresponds to the matching cost distribution of a specific depth hypothesis) of the cost volume , an adaptive window is constructed: ; wherein, denotes an adaptive window, i.e. a window centered at ; denotes the pixel coordinates of the RGB image; denotes the gray value of the guide image at ; denotes the gray value of the guide image at ; denotes a continuity threshold (e.g. depth or color difference), i.e. dynamically adjust the filter window according to the continuity of local depth (or color); denotes the maximum window radius; denotes taking the maximum value; denotes taking the maximum value of the horizontal distance between x and , and the maximum value of the vertical distance between y and ; compute the linear regression coefficients within the adaptive window: ; wherein, and both denote the linear regression coefficients; denotes the covariance function; denotes the variance function; denotes the k-th adaptive window; denotes the truncation of the guide image within ; denotes the truncation of the cost volume slice p within ; denotes a regularization parameter to prevent the denominator from being zero; denotes the mean value of the cost volume slice p within ; denotes the mean value of the guide image within ; based on the enhanced guide representation, filter each of the cost volumes along the depth dimension slice-by-slice: ; wherein, denotes the filtered cost volume slice; denotes an adaptive guide filtering operation based on the linear regression coefficients; denotes the d-th depth hypothesis corresponding cost volume slice; depth hypothesis means D candidate depths obtained by discretely sampling the scene depth at each stage, used to construct the cost volume and perform matching and regression, taking the 1-Dth candidate plane; weightedly fuse the filtered cost volume slices: ; wherein, denotes the cost volume slice after weighted fusion; denotes the adaptive easing factor, , denotes the base easing factor.
[0033] In the step S5, the regularization network is constructed based on a global attention regularization module and a three-dimensional U-shaped module (3D U-Net) based on probability volume geometry embedding; the global attention regularization module is used to perform a regularization operation on the cost volume of the 0th stage, and the three-dimensional U-shaped module is used to perform a regularization operation on the cost volumes of the remaining stages.
[0034] The adaptive guided filtering network directly refines the cost volume between each stage of the cascade, uses a reference map or a depth map of the previous stage as a guide, adaptively suppresses noise in smooth areas while maintaining the sharpness of geometric details at edges and thin-walled structures, realizes "purification" of the cost volume without sacrificing the integrity of key structures, thereby providing higher-quality input for subsequent stages and breaking the error propagation chain.
[0035] The adaptive guided filtering network ensures that more original cost volume details are retained at geometric edges, while stronger smoothing filtering is applied in flat areas; the goal of effectively filtering cost volume noise while maintaining its geometric structure is achieved. Compared with traditional fixed kernel filtering, the filtering strength can be adaptively adjusted according to the local geometric properties, thereby providing effective denoising effect in weak texture areas while maintaining the depth discontinuity at the object boundary, providing higher-quality cost volume for subsequent stages and effectively suppressing error propagation.
[0036] The regularization process of the global attention regularization module is specifically as follows: A three-dimensional coordinate grid matching the dimension of the cost volume is generated and normalized, and a sinusoidal position encoding is applied to each coordinate axis of the normalized three-dimensional coordinate grid to obtain an encoding feature, a convolution layer is used to project the encoding feature to obtain a projection feature, the projection feature is added to the cost volume to obtain a projected cost volume, and the projected cost volume is input into a 3D convolution layer with a stride to downsample to obtain a compact cost volume; ; wherein, denotes the compact cost volume; denotes the original cost volume; denotes feature projection; denotes sinusoidal position encoding; denotes three-dimensional coordinate grid; The compact cost volume is flattened along the spatial dimension and the depth dimension to obtain a one-dimensional sequence, and the one-dimensional sequence is input into several Transformer encoders using a global attention mechanism to calculate global attention values; ; wherein, represents the global attention value; Q represents a query matrix; K represents a key matrix; and V represents a value matrix; represents a key vector dimension; and T represents transposition; The design enables each position in the cost volume to directly interact with all other positions, thereby comprehensively capturing long-range dependencies and modeling the geometric consistency of the entire three-dimensional scene.
[0037] The global attention values are reshaped into a three-dimensional feature volume, which is then upsampled by a 3D transposed convolution layer to restore the original depth and spatial resolution, thereby obtaining a regularized cost volume.
[0038] To enable the Transformer encoder, which does not inherently have position awareness, to understand the three-dimensional spatial relationships in the cost volume, a three-dimensional position encoding mechanism is introduced, which encodes the normalized index coordinates of the cost volume to enhance the generalization ability of the model.
[0039] The global attention regularization module uses an attention mechanism with three-dimensional position encoding to process the cost volume in the initial stage of the cascaded network, thereby effectively capturing long-range dependencies in the entire three-dimensional scene, overcoming the limitations of the local receptive field of traditional 3D CNNs, and fundamentally improving the global consistency and accuracy of the initial cost volume estimation.
[0040] The formula of the adaptive geometric feature fusion network is: ; wherein, represents the perspective depth corresponding to the i th pixel of the depth map; represents the Mamba network; represents the visual features extracted from the reference map; represents a concatenation operation; represents a reliability weight; represents a Hadamard product; represents the geometric branch network; represents the reference map; represents the depth map obtained by upsampling in the previous stage.
[0041] The adaptive geometric feature fusion network can dynamically evaluate the reliability of geometric prior according to the confidence of the prediction result of the previous stage, adaptively adjust the weight of the prior information in feature fusion according to the reliability of the prior information, and ensure that only reliable geometric information is used to guide the next stage, thereby effectively blocking the transmission path of the error information and preventing the propagation of error or noise serious geometric data to the next stage, which not only suppresses the accumulation and amplification of errors in the cascade process, but also promotes the network to generate a feature representation with stronger geometric structure perception and higher robustness, thereby providing a solid foundation for the subsequent cost matching process.
[0042] The fusion step of the adaptive geometric feature fusion network is as follows: (1) Geometric prior reliability evaluation and weight generation: the core idea is that not all geometric priors from the previous rough stage are equally reliable, therefore, the probability volume generated in the previous stage (stage ℓ) is used to derive a confidence map, which is used as an indicator of prediction quality. Specifically, for any pixel z in the current refinement stage (stage ℓ+1), the upsampled confidence map is input into a small convolutional network g and a Sigmoid activation function, thereby generating a reliability weight in the range [0, 1], which is represented by the following formula: ; The reliability weight quantitatively represents the degree to which the geometric prior information at the corresponding position should be trusted, and a value close to 1 indicates high confidence and reliable prior, while a value close to 0 indicates low confidence and unreliable prior.
[0043] (2) Dual-branch feature extraction and parallel processing of visual appearance information and geometric structure information: the visual feature branch extracts basic visual features from the reference image through a feature pyramid network (FPN) backbone; the geometric feature branch (geometric branch network) inputs the reference image and the rough depth map obtained by upsampling the depth map of the previous stage, and extracts depth-related geometric structure features.
[0044] (3) Gated adaptive feature fusion and enhancement: the core is a gated fusion mechanism controlled by the reliability weight, which uses the reliability weight generated in the first step to dynamically modulate the output of the geometric feature branch through Hadamard product operation, then element-wise adds the modulated geometric features and visual features, and finally sends the complete feature representation after fusion to the depth-aware Mamba layer FM for processing to capture long-distance dependencies.
[0045] In summary, the advantages of the present application are: 1. By collecting multi-view RGB images, the camera for collecting RGB images is calibrated to obtain the projection matrix and the intrinsic matrix, and the reference graph for three-dimensional reconstruction is obtained; then the RGB images are input into the feature extraction network to obtain the image features, and a plurality of depth hypothesis planes are obtained by inversely and uniformly sampling the reference graph, the homography transformation is performed on the image features, and the depth hypothesis planes are aggregated to obtain a cost volume; then the adaptive guided filtering network is used to filter the cost volume, the regularized network is used to regularize the filtered cost volume, the channel dimension of the regularized cost volume is exponentially normalized to obtain a probability volume, the depth map and the confidence map are constructed based on the probability volume, and after the frequency domain filtering and up-sampling of the depth map and the confidence map, the adaptive geometric feature fusion network is input to fuse the features to obtain the perspective depth, and finally the three-dimensional reconstruction is performed based on the perspective depth, the projection matrix and the intrinsic matrix to obtain the dense point cloud; that is, the adaptive guided filtering network is used to suppress the noise transmission of the cost volume and maintain the geometric edge structure, the global attention regular module is used to capture the long distance dependence to enhance the context consistency of the cost volume, the reliability weight mechanism in the adaptive geometric feature fusion network is used to dynamically evaluate the confidence of the cross-stage depth information, and the Mamba network is used to selectively fuse the high-confidence geometric features, so as to effectively block the error accumulation in the complex scene such as weak texture and occlusion, and greatly improve the accuracy and robustness of the three-dimensional reconstruction.
[0046] 2. By introducing the adaptive guided filtering network, the accuracy and stability of the filtering are effectively improved by dynamically adjusting the guide map (using the reference graph in the initial stage and the depth map of the previous stage in the subsequent stage) and combining the edge preservation weight and the adaptive window; the edge preservation weight is based on the gradient calculated by the Sobel operator, the edge intensity is controlled by the exponential function, the texture details are protected in the filtering process, the edge blur in the reconstruction is reduced, and the point cloud quality is improved; the window size is adaptively adjusted according to the pixel gray continuity and the spatial distance, compared with the fixed window filtering, the scene light change and the occlusion problem can be better handled, and the robustness to complex environment is enhanced; the depth map of the previous stage is used as the guide map in the subsequent stage, realizing adaptive iterative optimization, gradually refining the depth estimation, reducing the noise influence, improving the reconstruction consistency, finally solving the problems of losing details and over-smoothing in the traditional guided filtering in the three-dimensional reconstruction, and improving the generalization ability and practicality.
[0047] 3. The combination of global attention regularization module and 3D U-shaped module is used in the regularization network to optimize the feature representation of the cost volume; the global attention regularization module processes one-dimensional sequences through the encoding of three-dimensional coordinate grids and the Transformer encoder to capture global context dependencies (such as the correlation between depth dimension and spatial dimension), improve the consistency of features, and reduce the voids or distortions in reconstruction; the 3D U-shaped module is used in the subsequent stage to effectively handle local details and multi-scale features by combining down-sampling and up-sampling operations, improving the resolution recovery capability of the cost volume; the module design avoids redundant calculations (such as down-sampling compact cost volume), combines convolution and transposed convolution, reduces the computational overhead while maintaining accuracy, and is suitable for real-time 3D reconstruction applications; this dual-module architecture integrates global and local regularization, overcoming the limitations of single regularization methods (such as pure CNN) in balancing global consistency and local details, and improving the density and accuracy of reconstruction.
[0048] 4. The adaptive geometric feature fusion network uses Mamba network to combine visual features and geometric features, achieving efficient feature fusion; as a state space model, Mamba has linear computational complexity and long sequence processing capability compared to traditional RNN or Transformer, which can efficiently fuse multi-view features and improve the efficiency of depth map estimation; by combining reliability weights through concatenation operation (⊕) and Hadamard product (⊙), the geometric branch output is dynamically weighted, enhancing the robustness to unreliable areas (such as occlusion or low texture) and reducing reconstruction errors; frequency domain filtering and up-sampling are performed before fusion to preprocess the depth map, ensuring the quality of input features and further improving the accuracy of view depth; this fusion mechanism solves the problem of insufficient fusion of visual and geometric features in traditional methods, and significantly improves the accuracy and speed of depth estimation by combining the advanced architecture of Mamba.
[0049] 5. The complete end-to-end flow is formed from image acquisition to point cloud generation, and the key steps are designed iteratively; integrating feature extraction, cost volume construction, filtering, regularization, and fusion reduces manual intervention and has high automation, making it suitable for large-scale 3D reconstruction scenarios; the outputs of previous stages (such as depth maps as guide maps) are used in multiple stages (such as filtering and fusion) to achieve progressive optimization, gradually improving the quality of depth maps and enhancing the adaptability of the method to noise and view changes; the probability volume is obtained by exponentially normalizing the cost volume, combined with the confidence map, to provide reliable depth uncertainty estimation, improving the density and reliability of the point cloud; this flow design improves the robustness and scalability of the method, reduces the risk of error accumulation compared to independent algorithms in different stages, and is easy to integrate into existing systems.
[0050] 6、Based on standard RGB images and camera parameters (intrinsic / extrinsic), no special hardware (such as depth sensors) is required, and compatibility is strong; inverse depth uniform sampling and homography transformation simplify cost volume construction, and are easy to implement; through adaptive filtering and Mamba fusion, it is possible to surpass traditional multi-view stereo (MVS) methods in accuracy (edge preservation) and speed (Mamba efficiency), and it is suitable for high dynamic scenes.
[0051] 7、By innovatively integrating adaptive guided filtering mechanism, efficient regularization network and Mamba geometric feature fusion framework, through dynamic adjustment of guided image (reference image→depth map iterative optimization) combined with edge preservation weight and adaptive window, the reconstruction edge detail preservation ability and scene adaptability are significantly improved; by using the dual-path design of global attention regularization module and three-dimensional U-shaped module, the global consistent features and local detail recovery are balanced, and the computational overhead is reduced; by means of the efficient long sequence processing capability of the Mamba network, the visual features and weighted geometric information are fused, the depth estimation problem of occlusion and low texture area is effectively solved, and finally the end-to-end high-precision and high-robustness three-dimensional reconstruction from multi-view RGB images to dense point cloud is realized.
[0052] Although the specific embodiments of the present application are described above, those skilled in the art should understand that the specific examples described are only illustrative, and are not intended to limit the scope of the present application, and equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present application should be covered within the scope of the claims of the present application.
Claims
1. A method for three-dimensional reconstruction based on guided filtering and Mamba geometric feature fusion, characterized in that: The method comprises the following steps: Step S1, acquiring multi-view RGB images, calibrating the intrinsic and extrinsic parameters of a camera for acquiring the RGB images to obtain a projection matrix and an intrinsic matrix, and obtaining a reference graph for three-dimensional reconstruction; Step S2, inputting each RGB image into a feature extraction network for feature extraction to obtain image features; Step S3, obtaining a plurality of depth hypothesis planes by inversely and uniformly sampling the reference graph, performing homography transformation on each image feature, and aggregating the depth hypothesis planes to obtain a cost volume; Step S4, performing a filtering operation on the cost volume through an adaptive guided filtering network; Step S5, performing a regularization operation on the cost volume after the filtering operation through a regularization network; Step S6, performing exponential normalization on the channel dimension of the cost volume after the regularization operation to obtain a probability volume, and constructing a depth map and a confidence map based on the probability volume; Step S7, after performing frequency domain filtering and upsampling on the depth map and the confidence map, inputting the depth map and the confidence map into an adaptive geometric feature fusion network for feature fusion to obtain a view depth; Step S8, performing a three-dimensional reconstruction operation based on each view depth, the projection matrix, and the intrinsic matrix to obtain a dense point cloud.
2. The method of claim 1, wherein the method is a method of 3D reconstruction based on fusion of guided filtering and Mamba geometric features. In the step S4, the filtering process of the adaptive guided filtering network is specifically as follows: In the 0th stage, the reference graph is used as a guide graph; in subsequent stages, the depth map generated in the previous stage is used as a guide graph; The horizontal gradient and the vertical gradient of the guide graph are calculated: ; ; wherein, represents a horizontal gradient; represents a vertical gradient; represents a guide map; represents a Sobel kernel in a horizontal direction; represents a Sobel kernel in a vertical direction; The edge preservation weight is calculated based on the horizontal gradient and the vertical gradient: ; wherein, denotes an edge-preserving weight; denotes a natural exponential function; denotes an edge-preserving strength control hyperparameter; An enhanced guide representation is constructed based on the edge preservation weight: ; wherein represents an enhanced guidance representation; for each pixel of a cost volume slice of the cost volume , a self-adaptive window is constructed: ; wherein, denotes an adaptive window, i.e. a window centered at ; denotes the pixel coordinates of the RGB image; denotes the gray value of the guide map at ; denotes the gray value of the guide map at ; denotes the continuity threshold; denotes the maximum window radius; denotes taking the maximum value; denotes taking the maximum value of the horizontal distance between x and , and the maximum value of the vertical distance between y and ; The linear regression coefficient is calculated within the adaptive window: ; wherein, and both represent linear regression coefficients; represents a covariance function; represents a variance function; represents the k-th adaptive window; represents the intercept of the guidance map within ; represents the intercept of the cost volume slice p within ; represents a regularization parameter that prevents the denominator from being zero; represents the mean of the cost volume slice p within ; represents the mean of the guidance map within ; Each cost volume is filtered along the depth dimension based on the enhanced guide representation: ; wherein, denotes a filtered cost volume slice; denotes an adaptive guided filtering operation based on linear regression coefficients; denotes a cost volume slice corresponding to the dth depth hypothesis; The filtered cost volume slices are fused by weighting: ; wherein, denotes the cost volume slice after weighted fusion; denotes the adaptive easing factor.
3. The method of claim 1, wherein the method is based on fusing guided filtering and Mamba geometric features for 3D reconstruction. In the step S5, the regularization network is constructed based on a global attention regularization module and a three-dimensional U-shaped module; the global attention regularization module is used to perform a regularization operation on the cost volume in the 0th stage, and the three-dimensional U-shaped module is used to perform a regularization operation on the cost volume in the remaining stages.
4. The method of claim 3, wherein the method is based on fusing guided filtering and Mamba geometric features for 3D reconstruction. The regularization process of the global attention regularization module is specifically as follows: A three-dimensional coordinate grid matching the dimension of the cost volume is generated and normalized, and the normalized three-dimensional coordinate grid is independently applied with sinusoidal position encoding to obtain an encoding feature; a projection feature is obtained by projecting the encoding feature through a convolution layer; the projection feature is added to the cost volume to obtain a projected cost volume; and the projected cost volume is input into a 3D convolution layer with a stride to downsample the projected cost volume to obtain a compact cost volume; The compact cost volume is flattened along the spatial dimension and the depth dimension to obtain a one-dimensional sequence; the one-dimensional sequence is input into a plurality of Transformer encoders adopting a global attention mechanism to calculate a global attention value; The global attention value is reshaped into a three-dimensional feature volume, and then upsampled through a 3D transpose convolution layer to restore the original depth and spatial resolution to obtain a regularized cost volume.
5. The method of claim 1, wherein: The formula of the adaptive geometry feature fusion network is: ; wherein, represents the depth of a pixel in the depth map; represents the depth of a pixel in the depth map; represents the Mamba network; represents the visual features extracted from the reference map; represents the concatenation operation; represents the reliability weight; represents the Hadamard product; represents the geometry branch network; represents the reference map; represents the depth map obtained by upsampling the previous stage.
Citation Information
Patent Citations
A cross-scale-based random walk stereo matching method
CN109887021A
Deep learning-based aneurysm detection and rupture risk assessment method and system
CN118864407A
Laser scanning three-dimensional imaging method, device, medium and equipment
CN120468879A
Multi-view three-dimensional reconstruction method based on multi-scale feature fusion
CN120580362A
Systems and methods for regularized reconstructions in MRI using side information
US20130278261A1
Cited By
Gold ore particle size distribution three-dimensional reconstruction method and device based on GPU and medium
CN122066872A
A GPU-based gold mine particle size distribution three-dimensional reconstruction method, device and medium
CN122066872B