Depth completion method based on geometric perception and channel attention mechanism

Through GAC-Net structure optimization, combined with PointNet++ and CSPN++ modules, the adaptive adjustment feature fusion is used to adaptively adjust the feature fusion of channel attention mechanism, which solves the problems of insufficient utilization of three-dimensional geometric information and high computational complexity in the existing depth completion method, and achieves an efficient depth completion effect.

CN120388061APending Publication Date: 2025-07-29ZHEJIANG UNIV OF SCI & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510460558.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing depth completion method ignores three-dimensional geometric information and lacks the utilization of the spatial structure of sparse depth maps, resulting in low completion accuracy of complex boundaries and large void areas, high computational complexity, and the feature fusion method is not efficient enough to dynamically adjust according to scene changes.

Method used

The GAC-Net structure based on geometric perception and channel attention mechanism is adopted, and global 3D geometric features are extracted through PointNet++, combined with U-Net and CSPN++ modules for multimodal feature fusion and deep refinement, the channel attention mechanism is used to adaptively adjust the feature contribution ratio, and 3D features are used in necessary stages to reduce calculation costs.

Benefits of technology

It improves scene geometry perception ability, improves adaptability to complex environments, enhances completion accuracy and boundary details recovery ability, reduces computing complexity, and improves network inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388061A_ABST
    Figure CN120388061A_ABST
Patent Text Reader

Abstract

The invention discloses a depth completion method based on geometric perception and a channel attention mechanism, and relates to the technical field of computer vision and deep learning. Comprising the following steps: generating an initial dense depth map by using a nonlinear propagation model; pointNet + + is adopted to extract global 3D geometric features; performing preliminary fusion on the initial dense depth map and the image features by using U-Net; performing weighted fusion on the global 3D geometric features and the preliminary fusion features through a multi-modal fusion module based on a channel attention mechanism to generate optimized fusion features; residual learning is carried out by using the fusion features to correct the initial dense depth map; and optimizing the complementation result by adopting CSPN + + in combination with the original sparse truth value in the sparse depth map. According to the method, the GAC-Net structure is integrally optimized, the scene geometric perception capability is improved, the 3D global features are fully utilized, and the adaptability to the complex environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and deep learning, and particularly relates to a depth completion method based on geometric perception and channel attention mechanism. Background Art

[0002] Depth completion is an important research direction in computer vision, aiming to recover high-quality dense depth maps from sparse depth measurement data. This technology has important application values in fields such as autonomous driving, robot navigation, 3D reconstruction, and object detection. However, due to the hardware limitations of lidar sensors, the obtained depth data is usually sparse and unevenly distributed, which affects the realization of high-precision depth perception tasks. Therefore, how to utilize the limited sparse depth information and combine it with auxiliary data such as RGB images to achieve higher-precision depth completion has become an important challenge in current research.

[0003] In recent years, significant progress has been made in depth completion models based on deep learning methods, such as methods like BPNet. However, there are the following limitations:

[0004] 1. Ignoring 3D geometric information: Most methods only rely on 2D image features for completion, lacking the utilization of the spatial structure information of sparse depth maps, resulting in low completion accuracy in complex boundary and large hole regions.

[0005] 2. Simply splicing multi-modal features: Some methods attempt to fuse RGB images and depth data, but usually use channel splicing or fixed-weight fusion, lacking adaptive adjustment of the contributions of different modal features, which affects the completion effect.

[0006] 3. Insufficient depth refinement ability: Existing methods often suffer from problems such as blurred boundaries, enhanced noise, and inconsistent local regions after completion, and more efficient depth propagation and optimization strategies are needed.

[0007] In recent years, some studies have attempted to introduce 3D geometric features into the depth completion task. Methods such as GraphCSPN and KBNet use point clouds to represent scene information. However:

[0008] 1. Directly using point clouds has a large computational amount, requires an additional point cloud processing network, and has a relatively high computational complexity.

[0009] 2. Some methods only adopt global feature splicing and fail to fully utilize local geometric information, which affects the completion accuracy.

[0010] 3. The feature fusion method is not efficient enough. The fusion strategy of RGB and 3D information is usually relatively fixed and cannot be dynamically adjusted according to scene changes.

[0011] Therefore, it is an urgent problem for those skilled in the art to propose a depth completion method based on geometric perception and channel attention mechanism to solve the difficulties existing in the prior art. Summary of the Invention

[0012] In view of this, the present invention provides a depth completion method based on geometric perception and channel attention mechanism. By optimizing the overall GAC-Net structure, the scene geometric perception ability is improved, the 3D global features are fully utilized, and the adaptability to complex environments is enhanced.

[0013] In order to achieve the above object, the present invention adopts the following technical solutions:

[0014] A depth completion method based on geometric perception and channel attention mechanism, comprising the following steps:

[0015] S1. Collect the sparse lidar point cloud data and RGB images obtained synchronously;

[0016] S2. Construct a GAC-Net architecture based on geometric perception and channel attention mechanism. The GAC-Net architecture includes a non-linear propagation model, a PointNet++ module, a U-Net structure, a channel attention feature fusion module, a CSPN++ module, and a six-scale pyramid structure;

[0017] S3. Use the non-linear propagation model to generate an initial dense depth map;

[0018] S4. Extract the global 3D geometric features of the initial sparse depth map by using the PointNet++ module;

[0019] S5. Based on the U-Net structure, preliminarily fuse the initial dense depth map and the RGB image to obtain an initial fusion feature;

[0020] S6. Use the channel attention feature fusion module for multi-modal feature fusion. Based on the SE mechanism, calculate the channel weights of the RGB image, the initial dense depth map, and the global 3D geometric features through global pooling, and use the fully connected layer for channel-level attention allocation to calculate the final fusion feature;

[0021] S7. Introduce residual learning to correct the initial dense depth map;

[0022] S8. Use the CSPN++ module and the six-scale pyramid structure to refine the corrected initial dense depth map, and combine the original sparse ground truth in the sparse lidar point cloud data to obtain the final dense depth map.

[0023] Optionally, in S1, the sparse lidar point cloud data and RGB images collected synchronously are used for depth completion.

[0024] Optionally, in S4, the PointNet++ module is used to extract the global 3D geometric features of the initial sparse depth map, that is, the encoder part is used to extract the scene structure information, generate the global 3D geometric features, and share them at all scales.

[0025] Optionally, in S6, a channel attention feature fusion module is used for multimodal feature fusion. Based on the SE mechanism, the channel weights of the RGB image, the initial dense depth map, and the global 3D geometric features are calculated through global pooling, and channel-level attention allocation is performed using a fully connected layer. The specific content of calculating the final fusion feature is as follows:

[0026] Based on the SE mechanism, the channel weights of the RGB image, the initial dense depth map, and the global 3D geometric features are calculated through global pooling, and the contribution ratio of the RGB image features and the global 3D geometric features is adaptively adjusted. The formula is as follows:

[0027] A U-Net = σ(FC(GlobalPool(F U-Net )))

[0028] A 3D = σ(FC(GlobalPool(F 3D )))

[0029] And channel-level attention allocation is performed using a fully connected layer to calculate the final fusion feature. The formula is as follows:

[0030] F fused = (A U-Net · F U-Net ) + (A 3D · F 3D )

[0031] Among them, F U-Net is the initial fusion feature output by the U-Net, F 3D is the global 3D geometric feature extracted by the PointNet++ module, σ is the Sigmoid activation function, FC is the fully connected layer, GlobalPool is the global pooling operation, A U-Net is the attention weight vector of the initial fusion feature, A 3D is the attention weight vector of the 3D global feature, and F fused is the final fusion feature.

[0032] Optionally, it also includes using the L2 supervised loss function for multi-scale depth supervision.

[0033] As can be seen from the above technical solutions, compared with the prior art, the present invention provides a depth completion method based on geometric perception and channel attention mechanism, which has the following beneficial effects:

[0034] (1) The present invention optimizes the overall GAC-Net structure, improves the scene geometry perception ability, makes full use of 3D global features, and enhances the adaptability to complex environments;

[0035] (2) Based on the self-adaptive fusion of the attention mechanism, it effectively optimizes the complementarity of RGB and depth information and improves the completion accuracy;

[0036] (3) The multi-scale depth optimization mechanism, combined with CSPN++, improves the boundary detail restoration ability and reduces noise and blurred areas;

[0037] (4) Efficient computational optimization, using 3D features only in necessary stages, reducing the computational cost and improving the network inference speed.

[0038] (5) Experimental results show that the present invention achieves an RMSE of 680.82 mm in the KITTI depth completion benchmark test, which is better than the existing state-of-the-art methods. The present invention is applicable to multiple computer vision tasks such as autonomous driving, 3D reconstruction, and object detection, and can effectively improve the 3D environment perception ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0040] Figure 1 It is a schematic diagram of the GAC-Net architecture based on geometric perception and channel attention mechanism provided by the present invention;

[0041] Figure 2 It is a schematic diagram of multi-modal feature fusion using the channel attention feature fusion module provided by the present invention;

[0042] Figure 3 It is a comparison diagram of the completion effects between the present invention and other state-of-the-art depth completion networks provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0044] Reference Figure 1 As shown, the present invention discloses a depth completion method based on geometric perception and channel attention mechanism, including the following steps:

[0045] S1. Collect the simultaneously acquired sparse lidar point cloud data and RGB images;

[0046] S2. Construct a GAC-Net architecture based on geometric perception and channel attention mechanism. The GAC-Net architecture includes a non-linear propagation model, a PointNet++ module, a U-Net structure, a channel attention feature fusion module, a CSPN++ module, and a six-scale pyramid structure;

[0047] S3. Use the non-linear propagation model to generate an initial dense depth map;

[0048] S4. Use the PointNet++ module to extract the global 3D geometric features of the initial sparse depth map;

[0049] S5. Based on the U-Net structure, preliminarily fuse the initial dense depth map and the RGB image to obtain an initial fusion feature;

[0050] S6. Use the channel attention feature fusion module for multi-modal feature fusion. Based on the SE mechanism, calculate the channel weights of the RGB image, the initial dense depth map, and the global 3D geometric features through global pooling, and use a fully connected layer for channel-level attention allocation to calculate the final fusion feature;

[0051] S7. Introduce residual learning to correct the initial dense depth map;

[0052] S8. Use the CSPN++ module and the six-scale pyramid structure to refine the corrected initial dense depth map, and combine the original sparse ground truth in the sparse lidar point cloud data to obtain the final dense depth map.

[0053] Furthermore, in S1, the simultaneously acquired sparse lidar point cloud data and RGB images are used for depth completion.

[0054] Furthermore, in S4, the PointNet++ module is used to extract the global 3D geometric features of the initial sparse depth map, that is, the encoder part is used to extract the scene structure information, generate the global 3D geometric features, and share them at all scales.

[0055] Furthermore, in S6, the channel attention feature fusion module is used for multi-modal feature fusion. The specific content of calculating the channel weights of the RGB image, the initial dense depth map, and the global 3D geometric features through global pooling based on the SE mechanism, and using a fully connected layer for channel-level attention allocation to calculate the final fusion feature is:

[0056] Based on the SE mechanism, calculate the channel weights of the RGB image, the initial dense depth map, and the global 3D geometric features through global pooling, and adaptively adjust the contribution ratio of the RGB image features and the global 3D geometric features. The formula is as follows:

[0057] A U-Net =σ(FC(GlobalPool(F U-Net )))

[0058] A 3D =σ(FC(GlobalPool(F 3D )))

[0059] And use the fully connected layer to perform channel-level attention allocation to calculate the final fused feature. The formula is as follows:

[0060] F fused =(A U-Net ·F U-Net )+(A 3D ·F 3D )

[0061] Where F U-Net is the initial fused feature output by U-Net, F 3D is the global 3D geometric feature extracted by the PointNet++ module, σ is the Sigmoid activation function, FC is the fully connected layer, GlobalPool is the global pooling operation, A U-Net is the attention weight vector of the initial fused feature, A 3D is the attention weight vector of the 3D global feature, and F fused is the final fused feature.

[0062] Specifically, the specific content of introducing residual learning in S7 to correct the initial dense depth map is as follows:

[0063] (1) Obtain the initial dense depth map D pre ;

[0064] (2) Based on the final multi-modal fused feature F fused , extract the residual information ΔD, and this residual is modeled by a convolutional neural network, that is:

[0065] ΔD = Conv(F fused )

[0066] (3) Add the residual term to the original depth map to obtain the corrected initial depth map D0:

[0067] D0 = D pre +ΔD

[0068] The above process enables the depth estimation network to retain fine-grained structural information through a residual-guided approach without directly regressing the absolute depth, thereby improving the local accuracy and overall consistency of depth estimation.

[0069] Specifically, the present invention uses a propagation network based on the CSPN++ module for depth refinement. By combining the local neighborhood affinity matrix to optimize the depth information propagation method, it improves the detail recovery ability of the completion, enhances the depth consistency in the edge region, reduces the blur effect during the completion process, makes the generated dense depth map smoother, while maintaining clear boundary features and more accurately fitting the true depth distribution.

[0070] The present invention adopts a six-scale network architecture to gradually refine the depth completion results at different scales. By combining the global information at low resolution and the local detail information at high resolution, this method can effectively improve the prediction accuracy and robustness of the completion network, ensuring high-quality depth completion results can still be provided in complex scenarios.

[0071] Furthermore, it also includes using the L2 supervision loss function for multi-scale depth supervision.

[0072] Specifically, a six-scale supervision loss function is adopted. This loss calculation method is based on the L2 loss framework and is optimized for the predicted depth maps at different scales. Specifically:

[0073] (1) Adopt a six-scale architecture to perform L2 supervision on the prediction results of each scale;

[0074] (2) Only calculate the valid pixels of the true depth map to reduce the influence of invalid data on training;

[0075] (3) Use bilinear interpolation for scale alignment to ensure the consistency of multi-scale loss calculation;

[0076] (4) Adopt a multi-scale weighted strategy to evenly distribute the loss influence across different scales.

[0077] In a specific embodiment, it includes the following content:

[0078] First, set the dataset:

[0079] The present invention uses the KITTI depth completion dataset for training and testing. This dataset consists of synchronously acquired RGB images and sparse depth maps (sparse lidar point cloud data), including complex environments such as outdoor urban roads, suburbs, villages, pedestrians, and vehicles in different scenarios.

[0080] Dataset division: Training set: 85,898 groups of data, Validation set: 1,000 groups of data, Test set: 1,000 groups of data.

[0081] Input data format: RGB image: size H×W×3; Sparse depth map: size H×W, only containing depth values of some pixels;

[0082] Output data format: Dense depth map, size H×W, used to generate a complete 3D environmental representation.

[0083] The present invention adopts the GAC-Net architecture based on geometric perception and channel attention mechanism, and its core modules include:

[0084] (1) Nonlinear propagation model:

[0085] Completes the input sparse depth map to generate an initial dense depth map.

[0086] (2) PointNet++ module extracts global 3D geometric features:

[0087] Adopts the encoder part to extract scene structure information;

[0088] Generates global 3D geometric features and shares them at all scales to improve inference efficiency.

[0089] (3) U-Net performs preliminary fusion:

[0090] Adopts multi-scale feature extraction to fuse RGB image and initial dense depth information.

[0091] (4) As Figure 2 shown, multi-modal fusion based on channel attention mechanism (CAFFM):

[0092] Calculates the channel weights of the RGB image, initial dense depth map and global 3D geometric features through the SE mechanism by global pooling, adaptively adjusts the contribution ratio of the RGB image features and global 3D geometric features, and the formula is as follows:

[0093] A U-Net =σ(FC(GlobalPool(F U-Net )))

[0094] A 3D =σ(FC(GlobalPool(F 3D )))

[0095] And uses the fully connected layer to perform channel-level attention allocation to calculate the final fused features, and the formula is as follows:

[0096] F fused =(A U-Net ·F U-Net )+(A 3D ·F 3D )

[0097] Among them, F U-Net is the initial fusion feature output by U-Net, and F 3D is the global 3D geometric feature extracted by the PointNet++ module. σ is the Sigmoid activation function, FC is the fully connected layer, GlobalPool is the global pooling operation, and A U-Net is the attention weight vector of the initial fusion feature, and A 3D is the attention weight vector of the 3D global feature, and F fused is the final fusion feature.

[0098] (5) Introduce residual learning to correct the initial dense depth map;

[0099] (6) The CSPN++ module performs depth refinement, optimizes the depth completion result through local domain propagation, reduces the edge blur effect, adopts a six-scale pyramid structure, and gradually optimizes the dense depth map at different scales to obtain the final dense depth map.

[0100] The present invention is trained using the PyTorch deep learning framework, and the hyperparameters are set as follows:

[0101] Optimizer: Aadm, with the weight decay set to 0.05.

[0102] The maximum learning rate is 2.5×10 -4 , and the warm-up preheating strategy is adopted. It gradually rises to the maximum learning rate in 10% of the training iterations, and then the learning rate decays according to cosine annealing in the subsequent 90% of the iterations, decaying to 10% of the maximum learning rate.

[0103] Training hyperparameters:

[0104] Number of training epochs: 30 rounds; Batch size: 2.

[0105] 3D feature extraction: The PointNet++ module is used to extract 256-dimensional global geometric features from the point cloud generated by projecting the sparse depth map.

[0106] Loss function: The L2 loss (mean square error) is used for multi-scale depth supervision.

[0107] The performance of the GAC-Net architecture based on geometric perception and channel attention mechanism of the present invention on the official test set of the KITTI depth completion dataset is shown in Table 1, and the comparison between the present invention and the existing state-of-the-art depth completion methods is shown in Table 2.

[0108] Table 1 Performance of the GAC-Net architecture of the present invention on the official test set of the KITTI depth completion dataset

[0109] RMSE MAE iRMSE iMAE Error 680.82 193.85 1.81 0.84

[0110] It can be found that the GAC-Net architecture of the present invention has achieved good results in various indicators. In the KITTI depth completion dataset leaderboard, it currently ranks third. The experimental results show that the present invention has achieved an RMSE of 680.82 mm in the KITTI depth completion benchmark test, which is better than the existing state-of-the-art methods.

[0111] Table 2 Comparison between the present invention and the existing state-of-the-art depth completion methods

[0112] Method RMSE MAE iRMSE iMAE GAC-Net (This invention) 680.82 193.85 1.81 0.84 BP-Net 684.90 194.69 1.82 0.84 GraphCSPN 738.41 199.31 1.96 0.84 CSPN++ 743.69 209.28 2.07 0.90

[0113] Figure 3 It is a comparison chart of the completion effects between the present invention and other state-of-the-art depth completion networks. As Figure 3 shown, it can be found that the present invention leads the current state-of-the-art depth completion methods in the main indicator RMSE and other indicators.

[0114] The present invention optimizes the overall GAC-Net structure, improves the scene geometric perception ability, makes full use of 3D global features, and enhances the adaptability to complex environments; the adaptive fusion based on the attention mechanism effectively optimizes the complementarity of RGB and depth information and improves the completion accuracy; the multi-scale depth optimization mechanism combines CSPN++ to improve the boundary detail recovery ability and reduce noise and blurred areas; the efficient computing optimization only uses 3D features in necessary stages, reduces the computing cost, and improves the network inference speed.

[0115] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0116] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A depth completion method based on geometric perception and channel attention mechanism, characterized in that It includes the following steps: S1. Collect the sparse LiDAR point cloud data and RGB images obtained synchronously; S2. Construct the GAC-Net architecture based on geometric perception and channel attention mechanism. The GAC-Net architecture includes a non-linear propagation model, a PointNet++ module, a U-Net structure, a channel attention feature fusion module, a CSPN++ module, and a six-scale pyramid structure; S3. Use the non-linear propagation model to generate an initial dense depth map; S4. Employ the PointNet++ module to extract the global 3D geometric features of the initial sparse depth map; S5. Based on the U-Net structure, preliminarily fuse the initial dense depth map and the RGB image to obtain initial fusion features; S6. Use the channel attention feature fusion module for multi-modal feature fusion. Based on the SE mechanism, calculate the channel weights of the RGB image, the initial dense depth map, and the global 3D geometric features through global pooling, and use a fully connected layer for channel-level attention allocation to calculate the final fusion features; S7. Introduce residual learning to correct the initial dense depth map; S8. Use the CSPN++ module and the six-scale pyramid structure to refine the depth of the corrected initial dense depth map, and combine the original sparse ground truth in the sparse LiDAR point cloud data to obtain the final dense depth map.

2. A depth completion method based on geometric perception and channel attention mechanism according to claim 1, characterized in that in S1, the sparse LiDAR point cloud data and RGB images collected synchronously are used for depth completion.

3. A depth completion method based on geometric perception and channel attention mechanism according to claim 1, characterized in that in S4, the PointNet++ module is used to extract the global 3D geometric features of the initial sparse depth map, that is, the encoder part is used to extract the scene structure information to generate global 3D geometric features, which are shared at all scales.

4. A depth completion method based on geometric perception and channel attention mechanism according to claim 1, characterized in that in S6, the channel attention feature fusion module is used for multi-modal feature fusion. The specific content of calculating the channel weights of the RGB image, the initial dense depth map, and the global 3D geometric features through global pooling based on the SE mechanism and using a fully connected layer for channel-level attention allocation to calculate the final fusion features is as follows: Based on the SE mechanism, calculate the channel weights of the RGB image, the initial dense depth map, and the global 3D geometric features through global pooling, and adaptively adjust the contribution ratio of the RGB image features and the global 3D geometric features. The formula is as follows: A U-Net = σ(FC(GlobalPool(F U-Net ))) A 3D = σ(FC(GlobalPool(F 3D ))) And use a fully connected layer for channel-level attention allocation to calculate the final fusion features. The formula is as follows: F fused = (A U-Net · F U-Net ) + (A 3D · F 3D ) Among them, F U-Net is the initial fusion feature output by U-Net, F 3D is the global 3D geometric feature extracted by the PointNet++ module, σ is the Sigmoid activation function, FC is the fully connected layer, GlobalPool is the global pooling operation, A U-Net is the attention weight vector of the initial fusion feature, A 3D is the attention weight vector of the 3D global feature, F fused is the final fusion feature.

5. A depth completion method based on geometric perception and channel attention mechanism according to claim 1, characterized in that it also includes using the L2 supervised loss function for multi-scale depth supervision.

Citation Information

Cited By

  • Image restoration method based on multi-view 3D reconstruction and geometric attention

    CN121504771A

  • Image inpainting method based on multi-view 3D reconstruction and geometric attention

    CN121504771B

  • Concrete structure crack detection method based on GP-MSF model

    CN121788856A