A Depth Completion Method Robust to Outliers Based on Segmentation and Regression Networks
Through the segmentation-regression cascade network architecture and robust L1 loss function, the problem of the existing depth completion method being affected by outliers is solved, and the accuracy and robustness of the depth map are improved.
Patent Information
- Application Number
- CN202211532176.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-01
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-12-01
AI Technical Summary
The existing depth completion methods are easily affected by a small number of outliers, resulting in a degradation in performance of most points.
Using a segmentation-regression cascade network architecture, the depth plane is generated by segmentation network and filtering outliers, combined with the regression network to calculate the residual depth map, and finally generate the initial depth map, and supervise it using the L1 loss function that is robust to the outliers.
The performance of most points is improved, the robustness to outliers is enhanced, and the generated depth map accuracy is higher.
Smart Images

Figure CN115810019B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of machine learning and monocular depth completion, and relates to algorithms such as ResNet, ConvNext, and depth planes. Specifically, it relates to a depth completion method that is robust to outliers based on a segmentation and regression network. Background Art
[0002] When a sparse depth map is given, depth completion can obtain a more accurate dense depth map. However, since it often has to fit a small number of outliers composed of noise and depth discontinuity regions, the performance of the vast majority of points deteriorates. To solve this problem, we propose a learning paradigm that is robust to outliers. Specifically, by combining the advantages of a segmentation network and a regression network, a segmentation-regression cascade network architecture is proposed.
[0003] Because an image contains a large amount of information, such as surface, edge, and semantic information, which helps the depth completion process. Therefore, many existing depth completion technologies help the network to perform more accurate depth completion by mining information in the color image. For example, DeepLiDAR: Deep Surface Normal Guided Depth Prediction for Outdoor Scene from Sparse LiDAR uses the relationship between depth and normal to obtain a more accurate depth map. Learning Guided Convolutional Network for Depth Completion proposes an image-guided convolution to fully utilize the correspondence between the color image and the depth map to assist the network in predicting a more accurate depth map. PENet: Towards Precise and Efficient Image Guided Depth Completion combines the information contained in the depth domain and the color domain and fuses the two at different stages and different scales. A Multi-Scale Guided Cascade Hourglass Network for Depth Completion uses a series of cascade hourglass networks of different scales to mine the information contained in the color image at different scales. RigNet: Repetitive Image Guided Network for Depth Completion uses a set of repetitive hourglass networks to mine similar information. The above methods are all based on regression networks for depth completion and are often affected by a small number of outliers, resulting in poor performance of the vast majority of points.
[0004] The depth plane refers to dividing the depth into several one-hot encoded maps according to a set of pre-set depth bins, and each encoded map represents the valid point positions in the corresponding depth bin; thereby, the regression problem can be transformed into a segmentation problem. The Deep Ordinal Regression Network for Monocular Depth Estimation classifies the depth plane using ordinal regression, which transforms the multi-classification problem into several binary classification problems, and each problem indicates whether the actual depth value is greater than the depth value of a specific depth plane. AdaBins: Depth Estimation using Adaptive Bins constructs the depth plane using adaptive depth bins to provide a more accurate depth map. The proposed bin-centered density loss enables the network to generate an optimal depth plane based on the input image. Depth Completion using Plane-Residual Representation proposes to estimate the dense depth map through the form of "from depth plane to residual" representation. This method is similar to the method of the present invention. The difference is that we directly use the predicted depth plane to form a rough segmentation depth map instead of in the way of weighted summation, aiming to protect the segmentation features from outliers in the regression method. Since the loss function used for segmentation and the final predicted categories are unordered attributes, using segmentation to generate the depth map can reduce the problem of accuracy degradation caused by fitting outliers, but the generated depth map often has low accuracy.
[0005] The present invention combines the advantages of both, and uses a depth completion method that is more robust to outliers, including a cascaded network architecture of segmentation-regression and a loss function that is more robust to outliers. Summary of the Invention
[0006] The present invention aims to provide a depth completion method that is more robust to outliers, and solve the problem that the existing depth completion methods are over-fitted to outliers, resulting in poor performance of most points.
[0007] The technical solution of the present invention is as follows:
[0008] A depth completion method based on a segmentation and regression network that is robust to outliers, and its steps are as follows:
[0009] Step 1: With the help of an in-vehicle camera, obtain a color image and the corresponding sparse depth map;
[0010] Step 2: Convert the sparse depth map into depth planes;
[0011] The sparse depth map can be converted into the corresponding depth plane according to the following formula:
[0012]
[0013] where D s represents the sparse depth map, P i represents the i-th depth plane, and b represents the depth segment; the depth segment is calculated by evenly dividing the depth range composed of the maximum depth and the minimum depth in a certain scene into K = 64 depth segments, and the central value of each depth segment is used as the depth value of this depth segment, denoted as b c .
[0014] Step 3: Input the color image, the sparse depth map, and the depth plane into a segmentation network together to obtain the corresponding depth plane prediction result, denoted as the segmented depth map.
[0015] (3.1) Send the above three inputs to an encoder-decoder structure, and the decoder outputs a segmentation guidance information where C s , H s and W s represent the number of channels, height, and width of the feature map respectively. The skip connections between the same encoder and decoder use addition.
[0016] (3.2) After the segmentation guidance information G s passes through a convolution and a Softmax layer, the segmentation probability l of the depth plane is obtained. After upsampling with a scale of 2, the segmentation probability l' with the same size as the input is obtained.
[0017] (3.3) According to the segmentation result l', the segmented depth map S can be calculated as follows:
[0018] S x,y = b c (argmax(l ′ x,y ))
[0019] Step 4: Calculate the residual between the sparse depth map and the segmented depth map, denoted as the sparse residual map R s :
[0020]
[0021] where is used to represent the valid points in the sparse depth map. When the condition in the parentheses is satisfied, its value is 1, otherwise it is 0.
[0022] Step 5: Input the color image, the segmented depth map, and the sparse residual map into a regression network to obtain the residual depth map;
[0023] (5.1) Input the above three images into another encoder-decoder structure, and the output of the decoder is the regression guidance information. where C r , H, and W are the number of channels, length, and width respectively. The skip connections between the same encoder-decoder use addition, and the skip connections between different encoder-decoders are concatenated.
[0024] (5.2) The regression guidance information passes through a convolutional layer to generate the residual depth map R.
[0025] Step 6: Add the segmented depth map and the residual depth map to obtain the initial depth map.
[0026] That is, the final dense depth map D * = S + R.
[0027] Step 7: Use the information provided by the segmentation network to filter out some outliers, and perform L1 loss supervision on the filtered points. This loss is called the L1 loss robust to outliers.
[0028] Given the segmentation probability l′, we first select the maximum probability at the positions of valid points in the ground truth; then we sort these points in descending order according to the previously selected maximum probability and select the top τ = 90% of the points, denoted as the point set Then the L1 loss robust to outliers is:
[0029]
[0030] where D gt and D * are the ground truth and the final predicted dense depth map respectively.
[0031] The final loss function used in this network is:
[0032] L = αL M + βL focal
[0033] where α = β = 1.0, and L focal is the segmentation loss function, and FocalLoss is used.
[0034] Advantages of the present invention:
[0035] (1) Using a cascaded network structure of segmentation-regression, the segmentation and regression tasks are completely decoupled. Compared with a pure regression network with the same number of parameters, it is beneficial to improve the performance for the vast majority of points.
[0036] (2) By leveraging the more robust property of the segmentation task against outliers, the probability results generated by the segmentation network are used as the basis for judging whether a point is an outlier. The outliers are filtered out and the supervision for non-outliers is concentrated, thereby further improving the performance of non-outliers. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram of the segmentation-regression cascaded network structure. DETAILED IMPLEMENTATION MANNER
[0038] The following further describes the detailed implementation manner of the present invention in combination with the drawings and technical solutions.
[0039] Figure 1 It is a schematic diagram of the segmentation-regression cascaded network structure, which is divided into two parts: a segmentation network and a regression network. The segmentation network uses depth planes to predict the depth plane to which each point belongs and generates the corresponding segmentation depth. The input of this network includes three parts: a color image, a sparse depth map, and the depth plane corresponding to the sparse depth map. The encoder of this network consists of several residual modules 1, where each residual module 1 is composed of a downsampling convolution and 3 ConvNext modules; the decoder consists of several decoder modules 1, and each decoder module 1 is composed of a transposed convolution module and a ResBlock module. The skip connection between the same encoder-decoder uses addition. The output of the decoder is the segmentation probability of the corresponding depth plane. Using this probability and the depth value of any depth plane, the corresponding segmentation depth can be obtained, and then through an upsampling, the segmentation depth of the original size is obtained.
[0040] Figure 1 The other part is the regression network, whose purpose is to obtain the residual between the segmentation depth and the ground truth. The input of this network includes three parts: a color image, the segmentation depth, and the residual between the segmentation depth and the sparse depth. The decoder of this network consists of several residual modules 2, and each residual module 2 contains 2 ResBlock modules, where the first ResBlock module can be used for downsampling; the decoder consists of several decoder modules 2, and each decoder module 2 contains a transposed convolution module. The final output of the decoder is the residual depth map of the segmentation depth to the ground truth.
[0041] The densely predicted depth map finally predicted by the network is the element-wise addition of the segmentation depth map and the residual depth map.
[0042] Another part of this method is the outlier-robust L1 loss. Its calculation method utilizes the segmentation probability of the depth plane. Since the segmentation network is less sensitive to outliers, points with too low segmentation probability are more likely to be outliers. First, take the maximum class probability of all valid points on the ground truth, sort all valid points in descending order according to this probability, and take the top specific percentage of points for L1 loss supervision.
[0043] For the segmentation network, the loss function used is Focal loss, aiming to make the network predict the corresponding depth plane more accurately and lay a foundation for the regression branch.
[0044] The segmentation-regression cascaded network is trained on the KITTI outdoor dataset. The data augmentation methods include: random flipping, random cropping, and color jittering. The optimizer is AdamW, where β 1 = 0.9, β 2 = 0.95. The initial learning rate is 0.002. The training process contains 30 epochs, and the learning rate will decay to 0.001 and 0.0002 at the 20th and 25th epochs.
[0045] During the training process, the batch size is 12, the image size is 352×608, and the input pair image size during testing is 352×1216.
Claims
1. A depth completion method robust to outliers based on a segmentation and regression network, characterized in that, the steps are as follows: Step 1: With the help of an in-vehicle camera, obtain a color image and a corresponding sparse depth map; Step 2: Convert the sparse depth map into a depth plane; According to the following formula, convert the sparse depth map into a corresponding depth plane: Among them, D s represents the sparse depth map, P i represents the i-th depth plane, and b represents the depth segment; the calculation of the depth segment is to evenly divide the depth range composed of the maximum depth and the minimum depth in a certain scene into K = 64 depth segments, and the central value of each depth segment is used as the depth value of the depth segment, denoted as b c ; Step 3: Input the color image, the sparse depth map, and the depth plane into a segmentation network together to obtain a corresponding depth plane prediction result, denoted as the segmentation depth map; (3.1) Feed the color image, sparse depth map, and depth plane into an encoder-decoder structure, and the decoder outputs a segmentation guidance information where C s , H s and W s represent the number of channels, height, and width of the feature map respectively; the skip connections between the same encoder-decoder use addition; (3.2) Segmentation guidance information G s After passing through a convolutional layer and a Softmax layer, the segmentation probability l of the depth plane is obtained; after upsampling with a scale of 2, the segmentation probability l' with the same size as the input is obtained. (3.3) According to the segmentation result l′, calculate the segmentation depth map S: S x,y = b C (argmax(l′ x,y )) Step 4: Calculate the residual from the sparse depth map to the segmentation depth map, denoted as the sparse residual map R s : Among them, is used to represent valid points in the sparse depth map. When the conditions in the brackets are met, its value is 1; otherwise, it is 0. Step 5: Input the color image, the segmentation depth map, and the sparse residual map into a regression network to obtain a residual depth map; (5.1) Input the color image, segmented depth map, and sparse residual map into another encoder-decoder structure, and the output of the decoder is the regression guidance information. where C r , H, and W are the number of channels, length, and width respectively; the skip connections between the same encoder-decoders use addition, and the skip connections between different encoder-decoders are concatenated. (5.2) The regression guidance information passes through a convolutional layer to generate a residual depth map R; Step 6: Add the segmented depth map and the residual depth map to obtain the initial depth map, i.e., the final dense depth map D * = S + R; Step 7: Use the information provided by the segmentation network to filter out a part of the outliers, and perform L1 loss supervision on the filtered points. This loss is called the L1 loss robust to outliers; Given the segmentation probability l′, first select the maximum probability of the positions of the valid points in the ground truth; then sort these points in descending order according to the previously selected maximum probability, and select the first τ = 90% of the points, denoted as the point set The L1 loss that is robust to outliers is then: Among them, D gt and D * are the true value and the finally predicted dense depth map respectively; The final loss function used in this network is: L = αL M + βL focal where α = β = 1.0, L focal is the segmentation loss function, and FocalLoss is used.
Citation Information
Patent Citations
Skin lesion segmentation and feature extraction method based on deep residual pyramid
CN113940635A
Monocular unsupervised depth estimation method based on contextual attention mechanism
US20210390723A1