A Real Depth Completion Method and Device Based on Pretraining and Scale Recovery
By using pre-trained visual feature encoder and monocular depth estimation method in depth completion, combined with multi-scale feature fusion and adaptive spatial propagation network, the problems of difficulty in depth completion and inaccurate scale information in the existing technology are solved, and efficient and accurate depth completion effect is achieved.
Patent Information
- Application Number
- CN202510186803.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The existing depth completion methods are difficult to effectively obtain dense absolute depth information, especially when learning is difficult and computationally expensive, and cannot accurately reflect the real scale information.
A pre-trained visual feature encoder is used, combined with a monocular depth estimation method, and through multi-scale feature fusion and weighted pooling, dense absolute depth maps are gradually restored, and fine-tuned through an adaptive spatial propagation network.
It significantly reduces the difficulty and complexity of model learning, can accurately obtain absolute depth information of the real scale, and improves the accuracy and robustness of depth completion.
Smart Images

Figure CN119672360B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a real depth completion method and device based on pre-training and scale recovery. Background Art
[0002] Obtaining dense depth information plays an important role in applications such as autonomous driving, robotics, and scene reconstruction. However, the reliable depth information obtained by common depth sensors such as lidar, depth cameras, iToF, etc. is usually sparse and cannot meet the requirements of practical applications. In addition, scene information can be obtained from high-resolution RGB images to estimate depth. Due to the characteristics of projection, the two-dimensional image lacks real scale information, and the estimated depth only represents the front-back relationship of the depth and cannot reflect the real scale. Therefore, using high-resolution RGB images to extract scene structure information and complete sparse depth is a good method for obtaining dense depth.
[0003] Most of the existing depth completion methods directly regress the depth with absolute scale, and there are large differences between different pixels in the whole image, which is difficult for the model to learn. In terms of network structure, the methods of early fusion or feature fusion are usually adopted, and the network is used to extract features respectively, and the generalization ability for different scenes is poor, and it is necessary to re-learn the feature expression for different categories of scenes. For the initial densification of sparse depth information, it is usually filled gradually by a learning method or directly by a non-learning method, resulting in a large amount of calculation or inaccurate completed depth maps. Summary of the Invention
[0004] The purpose of the present invention is to provide a real depth completion method and device based on pre-training and scale recovery in view of the deficiencies of the prior art.
[0005] The purpose of the present invention is achieved by the following technical solutions:
[0006] The first aspect of the present invention: A real depth completion method based on pre-training and scale recovery, comprising the following steps:
[0007] (1) Obtain the KITTI depth completion dataset, which includes the color image of the current scene collected and the sparse depth map obtained by projection of the depth sensor, and aggregate the sparse depth of consecutive frames to obtain the dense depth ground truth;
[0008] (2) Encode the input image using a pre-trained visual feature encoder on the dataset, and select the outputs of four intermediate layers among them to obtain multi-scale image features;
[0009] (3) Obtain the relative depth map using a monocular depth estimation method;
[0010] (4) Divide the sparse depth map pixel by pixel with the relative depth map at the depth valid points to obtain a sparse scale map, calculate the pooling weights based on the convolutional calculation of multi-scale image features, and downsample the sparse scale map to four different scale sizes in a weighted pooling manner;
[0011] (5) Take the multi-scale image features and the sparse scale map as inputs, perform fusion decoding for scale completion, and output the refined dense scale map;
[0012] (6) Perform a transformation from the scale space to the depth space, multiply the dense scale map pixel by pixel with the relative depth map obtained in step (2) to obtain the final dense absolute depth map;
[0013] (7) Calculate the loss between the dense depth ground truth, the relative depth map obtained in step (2), and the predicted dense absolute depth to supervise the learning of the depth completion method.
[0014] Further, the step (5) is implemented through the following sub-steps:
[0015] (5.1) Perform pre-filling on the sparse scale map, calculate the global single average scale for filling, output the initial dense scale map and obtain the confidence map;
[0016] (5.2) Use a lightweight convolutional network to extract features from the initial dense scale map output in step (5.1) to obtain scale features aligned with the image features;
[0017] (5.3) Perform hierarchical fusion decoding on the scale features and image features at four scales to output the dense scale map;
[0018] (5.4) Use an adaptive spatial propagation network for refinement: for all target pixel points to be updated in the dense scale map output in step (5.3), adaptively calculate the offset to select the neighborhood according to the mixed features and the output scale map, and directly predict the affinity, that is, the correlation weight, between the current pixel point; finally, update the current pixel scale value according to the affinity to obtain the final dense scale map.
[0019] Further, the step (5.1) specifically includes the following: At four scales, for the unknown scale points on the sparse scale map, search for their N nearest neighbor valid points and calculate the pixel distances between them; set a distance threshold. For the unknown scale points whose pixel distances from the nearest neighbors are less than the threshold, calculate the attention scores between the unknown scale points and the image features extracted in step (2) of the neighboring points in a cross-attention manner, and weight the scale values of the known scale points to obtain the scale of the current point; for the unknown scale points whose distances are greater than the threshold, calculate the global single average scale for filling to output the initial dense scale map; at the same time, extract a confidence map from the sparse scale map, and the confidence map indicates whether the measurement values on the depth map are valid.
[0020] Further, the loss calculation in the step (7) includes the supervision of both the depth ground truth and the relative depth. The loss loss expression is as follows: ;
[0021] where represents at the depth ground truth valid points , is the edge supervision of the relative depth with respect to the absolute depth; combine the two losses with the hyperparameter to obtain the total loss.
[0022] Specifically, the expression is as follows: ;
[0023] where, represents the dense depth ground truth, represents the predicted absolute depth.
[0024] Further, the edge supervision of the relative depth with respect to the absolute depth is specifically as follows: First, normalize the absolute depth, and then calculate the partial derivatives of the residual between the relative depth with respect to the x and y directions. The expression is as follows:
[0025] ;
[0026] where, represents the relative depth, represents the normalized predicted absolute depth.
[0027] The second aspect of the present invention: A real-depth completion device based on pre-training and scale recovery, including the following modules:
[0028] Acquisition and Aggregation Module: Obtain the KITTI depth completion dataset, which contains the color images of the current scene and the sparse depth maps obtained by projecting the depth sensor acquisitions, and aggregate the sparse depths of consecutive frames to obtain the dense depth ground truth; Image Encoding Module: Encode the input image using a visual feature encoder pre-trained on the dataset, and select the outputs of four intermediate layers to obtain multi-scale image features;
[0029] Estimation Module: Use a monocular depth estimation method to obtain the relative depth map;
[0030] Pooling and Sampling Module: Divide the sparse depth map by the relative depth map pixel by pixel at the depth valid points to obtain the sparse scale map, and calculate the pooling weights by convolving the multi-scale image features, and downsample the sparse scale map to four different scale sizes in a weighted pooling manner;
[0031] Decoding and Completion Module: Take the multi-scale image features and the sparse scale map as inputs, perform fusion decoding for scale completion, and output the refined dense scale map;
[0032] Pixel Calculation Module: Perform a transformation from the scale space to the depth space, multiply the dense scale map and the relative depth map obtained by the image encoding module pixel by pixel to obtain the final dense absolute depth map;
[0033] Loss Supervision Module: Calculate the loss between the dense depth ground truth and the relative depth map obtained by the image encoding module and the predicted dense absolute depth, and supervise the learning of the depth completion method.
[0034] The third aspect of the present invention: An electronic device, comprising:
[0035] One or more processors;
[0036] A memory for storing one or more programs;
[0037] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned real depth completion method based on pre-training and scale recovery.
[0038] The fourth aspect of the present invention: A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the above-mentioned real depth completion method based on pre-training and scale recovery are implemented.
[0039] The beneficial effects of the present invention are as follows:
[0040] The present invention performs calculations in the scale space through scale decomposition, and uses existing relative depth estimation methods for scale recovery, significantly reducing the difficulty and complexity of model learning and enabling the acquisition of the absolute depth of the true scale. When fusing and decoding image features and the sparse scale map, first, cross-attention or global average scale is used for pre-filling the sparse scale, which reduces the computational cost compared to directly applying convolution iteration on the sparse map and obtains a more accurate dense output compared to non-learning methods; and the image features extracted by the pre-trained encoder are utilized, and feature fusion decoding is performed with the scale features extracted from the scale map, which has better robustness to different scenarios; finally, by adaptively adjusting the neighborhood selection and affinity calculation of the spatial propagation network, iterative optimization of the scale map is carried out, and a more accurate scale result can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is the overall flowchart of the true depth completion method based on pre-training and scale recovery of the present invention;
[0042] Figure 2 is the design structure diagram of the scale recovery decoder of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.
[0044] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0045] The present invention will be described in detail below with reference to the drawings.
[0046] As Figure 1 shown, the present invention is a true depth completion method based on pre-training and scale recovery, including the following steps:
[0047] Step 1: Obtain dataset information: Here, the open-source KITTI depth completion dataset is selected. It includes the color image of the current scene, the sparse depth map obtained by projection of the depth sensor collection, and the dense depth ground truth obtained by aggregating consecutive frame sparse depths;
[0048] Step 2: Use the pre-trained feature encoder of Depth Anything V2 (based on DINOv2) to encode the input image, and select the outputs of four intermediate layers to obtain multi-scale image features: ; where represents four different scales.
[0049] Step 3: Use an existing monocular depth estimation method, such as Depth Anything V2, to obtain a relative depth map: ;
[0050] Step 4: Divide the input sparse depth map (sparse depth map) by the relative depth map pixel by pixel at valid points to obtain a sparse scale map (sparse scale map): ;
[0051] Use weighted pooling to calculate the pooling weights according to the image features at each scale, and downsample the sparse scale map to four different scale sizes: ;
[0052] Step 5: Take the multi-scale image features , the sparse scale map as inputs, perform fusion decoding for scale completion, and output a refined dense scale.
[0053] This step is the core of the present invention. As Figure 2 shown, it is divided into the following sub-steps.
[0054] (5.1) Pre-fill the sparse scale map and obtain a confidence map
[0055] For the invalid points of the scale map, search for its nearest neighbor points. Here, take , and calculate the pixel distance between them, and select the minimum distance . Set a distance threshold . For the unknown scale points whose distance from the nearest neighbor pixel is less than the threshold , use the cross-attention method to calculate the attention score between the image features of the unknown scale point and the neighboring points extracted in Step 2, and weight the scale values of the known scale points to obtain the scale of the current point. Here, the image feature of the target point is used as the query vector Q, The valid value points that are the nearest neighbors of the current point, neighboring points image features Take K and scale values Take V, calculate cross-attention, and obtain the pre-filled dense scale map For distances greater than the threshold For unknown scale points, calculate the global single average scale Perform filling and output the initial dense scale map , and the expression is as follows: ; At the same time, for the sparse scale map Extract the confidence map, and the confidence map indicates whether the measurement values on the sparse depth map are valid. The expression is as follows: ;
[0056] (5.2) Extract scale features
[0057] The scale input here is the scale map, which needs to be semantically aligned with the image features. Concatenate the initial dense scale map output in step (5.1) and the confidence map Perform feature concatenation, and use a lightweight convolutional network for feature extraction to obtain scale features aligned with the image features , and the expression is as follows: ;
[0058] (5.3) Perform feature fusion and decoding
[0059] The information here includes the image features , scale features , and the hybrid features output at the previous scale . Concatenate the three features through convolution one by one and output the hybrid features of the current scale The expression is as follows: ; Among them represents three convolutional residual networks with the same structure and non-shared parameters.
[0060] Output the dense scale map from the last-level hybrid features through the prediction head : ;
[0061] (5.4) Use the adaptive spatial propagation network for fine-tuning
[0062] At the same time, utilize the hybrid features and the features extracted from the scale map to guide the spatial propagation process: Use the hybrid features Concatenate with the features extracted from the scale map, and perform decoding in two branches to predict the offsets and affinity weights of the neighborhood pixels of the current point. When the neighborhood kernel size is k, the predicted output of the offset here is , where the number of channels represents the offsets in the x and y directions of each neighborhood point. The predicted output of the weight is , representing the weighted weights of each neighborhood point. The spatial propagation outputs a dense residual scale map and are added together to obtain the final dense scale output . The expressions for the predicted output of the offset and the predicted output of the weight are as follows: ; ;
[0063] Step 6: Perform a transformation from the scale space to the depth space, multiply the dense scale map and the relative depth map obtained in Step 3 pixel by pixel to obtain the final dense absolute depth map , and the expression is as follows: .
[0064] Step 7: Calculate the loss between the dense depth ground truth and the relative depth map obtained in Step 3 , and the predicted dense absolute depth in Step 6 to supervise the learning of the depth completion method. The total loss is as follows: ; where represents at the valid points of the depth ground truth . Since the depth ground truth is usually distributed inside the object, this loss is used to supervise the flat surface of the object: ; Since the edges of the relative depth are better, it can be used to supervise the edges of the absolute depth. First, normalize the absolute depth to obtain , and then calculate the partial derivatives of the residual between the absolute depth and the relative depth with respect to the x and y directions: ; Combine the two losses with the hyperparameter to obtain the total loss.
[0065] The present invention also discloses a real depth completion device based on pre-training and scale recovery, including the following modules:
[0066] Acquisition and Aggregation Module: Obtain the KITTI depth completion dataset, which contains the color images of the current scene collected and the sparse depth maps obtained by projecting the depth sensor. Aggregate the sparse depths of consecutive frames to obtain the dense depth ground truth; Image Encoding Module: Encode the input image using a pre-trained visual feature encoder on the dataset, and select the outputs of four intermediate layers to obtain multi-scale image features;
[0067] Estimation Module: Use a monocular depth estimation method to obtain the relative depth map;
[0068] Pooling and Sampling Module: Divide the sparse depth map by the relative depth map pixel by pixel at the depth valid points to obtain the sparse scale map, and calculate the pooling weights by convolving the multi-scale image features. Downsample the sparse scale map to four different scale sizes in a weighted pooling manner;
[0069] Decoding and Completion Module: Take the multi-scale image features and the sparse scale map as inputs, perform fusion decoding for scale completion, and output the fine-tuned dense scale map;
[0070] Pixel Calculation Module: Perform a transformation from the scale space to the depth space, multiply the dense scale map and the relative depth map obtained by the image encoding module pixel by pixel to obtain the final dense absolute depth map;
[0071] Loss Supervision Module: Calculate the loss between the dense depth ground truth and the relative depth map obtained by the image encoding module and the predicted dense absolute depth to supervise the learning of the depth completion method.
[0072] The present invention also discloses an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the above-mentioned real depth completion method based on pre-training and scale recovery; and discloses a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the above-mentioned real depth completion method based on pre-training and scale recovery are implemented.
[0073] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily think of other embodiments of the present application. The present application aims to cover any variations, uses, or adaptive changes of the present application, and these variations, uses, or adaptive changes follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0074] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A true depth completion method based on pre-training and scale recovery, characterized in that: The following steps are involved: (1) Obtain the KITTI depth completion dataset, which contains the color image of the current scene and the sparse depth map obtained by the depth sensor projection. Aggregate the sparse depth of consecutive frames to obtain the dense depth truth value; (2) Use the pre-trained visual feature encoder on the dataset to encode the input image, select the outputs of four intermediate layers, and obtain multi-scale image features; (3) Obtain relative depth map using monocular depth estimation method; (4) Divide the sparse depth map by the relative depth map pixel by pixel at the effective depth point to obtain a sparse scale map. Calculate the pooling weights based on the multi-scale image feature convolution and downsample the sparse scale map to four different scales using weighted pooling. (5) Taking multi-scale image features and sparse scale maps as input, fusion decoding is performed to complete the scale, and a fine-tuned dense scale map is output; It includes the following sub-steps: (5.1) Pre-fill the sparse scale map, calculate the global single average scale for filling, output the initial dense scale map and obtain the confidence map; specifically: at four scales, for the unknown scale point on the sparse scale map, search its N nearest valid neighbor points and calculate the pixel distance between them; set a distance threshold, for the unknown scale point whose distance to the nearest neighbor pixel is less than the threshold, use the cross-attention method to calculate the attention score between the unknown scale point and the image features extracted in step (2) of the neighboring point, and weight the scale value of the known scale point to obtain the scale of the current point; for the unknown scale point whose distance is greater than the threshold, calculate the global single average scale for filling, and output the initial dense scale map; at the same time, extract the confidence map for the sparse scale map, and the confidence map indicates whether the measurement value on the depth map is valid; (5.2) Using a lightweight convolutional network, extract features from the initial dense scale map output in step (5.1) to obtain scale features that are aligned with the image features; (5.3) Perform level-by-level fusion decoding on the scale features and image features at the four scales and output a dense scale map; (5.4) Fine-tune using the adaptive spatial propagation network: For all target pixels to be updated in the scale map output by step (5.3), adaptively calculate the offset to select the neighborhood based on the mixed features and the output scale map, and directly predict the affinity between the network and the current pixel, that is, the correlation weight; finally, perform a weighted update on the current pixel scale value based on the affinity to obtain the final dense scale map; (6) Transform from scale space to depth space, multiply the dense scale map and the relative depth map obtained in step (3) pixel by pixel to obtain the final dense absolute depth map; (7) Use the dense depth truth and the relative depth map obtained in step (3) to calculate the loss between the predicted dense absolute depth to supervise the learning of the depth completion method.
2. The depth completion method according to claim 1, characterized in that: The loss calculation in step (7) includes the supervision of the true depth value and the relative depth. The loss expression is as follows: ; in Represents the depth at the valid point , It is the edge supervision of relative depth to absolute depth; the two losses are used as hyperparameters Combined together, it is the total loss.
3. The depth completion method according to claim 2, characterized in that: Said The expression is as follows: ; in, represents the dense depth truth, Represents the absolute depth of the prediction.
4. The depth completion method according to claim 2, characterized in that: The edge supervision of the relative depth on the absolute depth is specifically as follows: first, the absolute depth is normalized, and then the partial derivatives of the residual between the relative depth and the absolute depth are calculated in the x and y directions. The expression is as follows: ; in, Indicates the relative depth. represents the absolute depth of the normalized prediction.
5. A real depth completion device based on pre-training and scale recovery, characterized in that: Includes the following modules: Acquisition and aggregation module: obtains the KITTI depth completion dataset, which contains the color image of the current scene and the sparse depth map obtained by the depth sensor acquisition projection, and aggregates the sparse depth of consecutive frames to obtain the dense depth truth value; Image encoding module: uses the visual feature encoder pre-trained on the dataset to encode the input image, selects the output of four intermediate layers, and obtains multi-scale image features; Estimation module: Use monocular depth estimation method to obtain relative depth map; Pooling sampling module: The sparse depth map is divided pixel by pixel with the relative depth map at the effective depth point to obtain a sparse scale map, and the pooling weight is calculated according to the multi-scale image feature convolution, and the sparse scale map is downsampled to four different scales using weighted pooling; Decoding and completion module: takes multi-scale image features and sparse scale maps as input, performs fusion decoding for scale completion, and outputs a fine-tuned dense scale map; specifically includes the following contents: pre-fills the sparse scale map, calculates the global single average scale for filling, outputs the initial dense scale map and obtains the confidence map; specifically: at four scales, for unknown scale points on the sparse scale map, searches for its N nearest valid neighbor points and calculates the pixel distance between them; sets a distance threshold, and for unknown scale points whose distance to the nearest neighbor pixel is less than the threshold, uses a cross-attention method to calculate the attention score between the image features extracted from the unknown scale point and the neighboring point, and weights the scale value of the known scale point to obtain the scale of the current point; for unknown scale points whose distance is greater than the threshold, calculates Calculate the global single average scale for filling and output the initial dense scale map; extract the confidence map for the sparse scale map at the same time, the confidence map indicates whether the measurement value on the depth map is valid; use the lightweight convolutional network to extract features from the output initial dense scale map to obtain scale features aligned with the image features; perform step-by-step fusion decoding on the scale features and image features at four scales to output the dense scale map; use the adaptive spatial propagation network for fine tuning: for all target pixels to be updated in the output scale map, according to the mixed features and the output scale map, adaptively calculate the offset to select the neighborhood, and directly predict the affinity between the current pixel and the network, that is, the correlation weight; finally, perform weighted update on the current pixel scale value according to the affinity to obtain the final dense scale map; Pixel calculation module: performs the transformation from scale space to depth space, multiplies the dense scale map and the relative depth map obtained by the image encoding module pixel by pixel, and obtains the final dense absolute depth map; Loss supervision module: Use the relative depth map obtained by the dense depth truth and image encoding module to calculate the loss between the predicted dense absolute depth to supervise the learning of the depth completion method.
6. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the true depth completion method based on pre-training and scale recovery as described in any one of claims 1-4.
7. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by the processor, the steps of the true depth completion method based on pre-training and scale recovery as described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Multi-level multi-attention target detection method based on radar and camera
CN119445301A
Method for 3D scene dense reconstruction based on monocular visual slam
US20200273190A1