A training method for a monocular depth estimation model based on cross-distillation

By introducing cross-distillation and uncertainty mapping in the monocular depth estimation model, the problem of difficulty in capturing long-distance correlation and local information in the prior art is solved, and higher depth estimation accuracy and lower computational cost are achieved.

CN116188904BActive Publication Date: 2025-06-13HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310115401.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-15
Publication Date
2025-06-13
Estimated Expiration
2043-02-15

AI Technical Summary

Technical Problem

Existing depth estimation methods based on CNN are difficult to effectively capture long-distance correlation and local information, resulting in low depth estimation accuracy.

Method used

The monocular depth estimation model training method based on cross-distillation was used, and the long-distance dependence and short-distance consistency were modeled respectively by Transformer and CNN, and the model was optimized through cross-distillation loss and uncertainty mapping.

Benefits of technology

Improve the accuracy of depth estimation, avoid the additional computational burden in the evaluation stage, and enable more efficient use of depth cues in the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188904B_ABST
    Figure CN116188904B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for training a monocular depth estimation model based on cross-distillation, which relates to the field of deep learning. The present invention includes the following steps: inputting urban RGB images into the monocular depth estimation model; the monocular depth estimation model generating a first depth prediction and a second depth prediction for the urban RGB images; using the two predicted depths as pseudo-labels and denoising the pseudo-labels; and optimizing the monocular depth estimation model by using the denoised pseudo-labels. The present invention can effectively improve the accuracy and has excellent performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and more specifically to a method for training a monocular depth estimation model based on cross-distillation. Background Art

[0002] Depth estimation is a fundamental research topic in the field of computer vision, and its applications range from scene understanding, 3D reconstruction to augmented reality. Benefiting from the progress of convolutional neural networks, recent research has achieved good depth results. Due to the lack of depth cues, making full use of long-range correlations (i.e., the distance relationship between objects) and local information (i.e., the consistency within objects) is crucial for accurate depth estimation. However, it is difficult for convolutional operators with limited receptive fields to capture long-range correlations, which becomes a potential bottleneck for current CNN-based depth estimation methods.

[0003] A large amount of work has been devoted to alleviating the above-mentioned defects of CNNs, which can be roughly divided into two categories: manipulating convolutional operations and incorporating attention mechanisms. The former uses dilated spatial pyramid pooling, coarse-to-fine fusion, and dense connections to enhance the efficacy of convolutional operators. The latter sets up attention modules to establish long-range dependencies in the feature maps. In addition, some general methods utilize both strategies. Although there has been a considerable improvement in depth accuracy, the dilemma still exists.

[0004] Recently, Transformer has been proven to be promising to replace CNN. Based on the attention mechanism, Transformer with a global receptive field is better at capturing long-range correlations. However, due to the lack of spatial inductive bias, local feature details are easily ignored by it, resulting in unsatisfactory performance. Some depth estimation methods that overcome the shortcomings of Transformer utilize additional CNN branches. However, their frameworks also rely on CNN branches during the evaluation phase, increasing the computational cost during inference. Summary of the Invention

[0005] In view of this, the present invention provides a method for training a monocular depth estimation model based on cross-distillation to solve the problems existing in the background art.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for training a monocular depth estimation model based on cross-distillation, comprising the following steps:

[0008] Inputting a city RGB image into a monocular depth estimation model;

[0009] The monocular depth estimation model generates a first depth prediction and a second depth prediction for the city RGB image;

[0010] Use the first depth prediction and the second depth prediction as pseudo - labels and denoise the pseudo - labels;

[0011] Optimize the monocular depth estimation model using the denoised pseudo - labels.

[0012] Optionally, the expressions of the first depth prediction and the second depth prediction are respectively:

[0013]

[0014] where, represents the prediction from the Transformer branch and represents the prediction from the CNN branch , p represents a pixel point, and r n (p) represents the RGB image.

[0015] Optionally, use the cross - distillation loss based on uncertainty to denoise the pseudo - labels. The definition formula of the cross - distillation loss is as follows:

[0016]

[0017] where, - is the gradient stop operation, is the element - wise multiplication, is the uncertainty map from the Transformer branch is the uncertainty map from the CNN branch , and p represents a pixel point.

[0018] Optionally, model the uncertainty map. The modeling formula is as follows:

[0019]

[0020] where, d n (p) represents the predicted depth map, represents the ground - truth depth map, b is a coefficient controlling the error tolerance, T represents the set of pixels with valid ground - truth depth, and p represents a pixel point.

[0021] Optionally, apply L u to predict the uncertainty to approximate

[0022]

[0023] where, represents the uncertainty map predicted by the Transformer branch, represents the uncertainty map predicted by the CNN branch, represents the ground - truth of the uncertainty map of the Transformer branch Represents the ground truth of the CNN branch uncertainty map, and p represents a pixel point.

[0024] Optionally, the formula for optimizing the monocular depth estimation model is as follows:

[0025] L total = L sl + λ 1 L urcd + λ 2 L u ,

[0026]

[0027] Where L sl is the silog loss, λ 1 , λ 2 , κ, η are hyperparameters used to balance the weights of each term in the loss, |T| represents the number of pixels with valid values, d n (p) represents the predicted depth map, represents the ground truth depth map, represents the depth error of the Transformer branch in the log space, represents the depth error of the CNN branch in the log space.

[0028] According to the above technical solutions, compared with the prior art, the present invention discloses a method for training a monocular depth estimation model based on cross-distillation, which has the following beneficial effects:

[0029] 1. A new monocular depth estimation model is introduced, which contains cross-distillation based on uncertainty correction to utilize long-distance correlations and local information. In addition, due to the cross-distillation paradigm, the model has no additional computational burden during the evaluation phase.

[0030] 2. A simple and effective data augmentation strategy is designed, which enables the model to focus on more valuable depth estimation cues rather than just cues related to the vertical image position.

[0031] 3. It can effectively improve the accuracy and has excellent performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0033] Figure 1 This is a schematic structural diagram of the present invention. Specific implementation manners

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] An embodiment of the present invention discloses a method for training a monocular depth estimation model based on cross-distillation, which is applied to the field of autonomous driving and is used for processing urban traffic images. As Figure 1 shown, it includes the following steps:

[0036] Input the RGB images collected in the urban scene into the monocular depth estimation model;

[0037] Due to the lack of depth cues, making full use of long-range dependencies and short-range consistencies is crucial for accurate urban scene depth estimation.

[0038] In view of this, Transformer and CNN are used to model long-range dependencies and short-range consistencies respectively. Since directly integrating Transformer and CNN will bring a huge increase in computational complexity, we generate a first depth prediction and a second depth prediction for the RGB images through the monocular depth estimation model;

[0039] Use the first depth prediction and the second depth prediction as pseudo-labels and denoise the pseudo-labels;

[0040] Use the denoised pseudo-labels to optimize the monocular depth estimation model.

[0041] The expressions of the first depth prediction and the second depth prediction are respectively:

[0042]

[0043] where represents the prediction from the Transformer branch and represents the prediction from the CNN branch respectively.

[0044] Use the cross-distillation loss based on uncertainty to denoise the pseudo-labels. The definition formula of the cross-distillation loss is as follows:

[0045]

[0046] Among them, - is the gradient stopping operation, and is the element-wise multiplication, is from the Transformer branch is from the CNN branch is the uncertainty map, and p represents the pixel point.

[0047] For modeling the uncertainty map, the modeling formula is as follows:

[0048]

[0049] Among them, d n (p) represents the predicted depth map, represents the ground truth depth map, b is a coefficient for controlling the error tolerance, and T represents the set of pixels with valid ground truth depth.

[0050] Apply L u to predict the uncertainty to approximate

[0051]

[0052] The formula for optimizing the monocular depth estimation model is as follows:

[0053] L total = L sl + λ 1 L urcd + λ 2 L u ,

[0054]

[0055] Among them, L sl is the silog loss, λ 1 , λ 2 , κ, η are hyperparameters used to balance the weights of each term in the loss, |T| represents the number of pixels with values, and d n (p) represents the predicted depth map, represents the ground truth depth map, represents the depth error of the Transformer branch in the log space, represents the depth error of the CNN branch in the log space.

[0056] The embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0057] The foregoing description of the disclosed embodiments enables those skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for training a monocular depth estimation model based on cross-distillation, characterized in that, it includes the following steps: Input the urban RGB image into the monocular depth estimation model; The monocular depth estimation model generates a first depth prediction and a second depth prediction for the urban RGB image; Use the first depth prediction and the second depth prediction as pseudo-labels and denoise the pseudo-labels; Optimize the monocular depth estimation model using the denoised pseudo-labels; The expressions of the first depth prediction and the second depth prediction are respectively: Among them, represents the prediction from the Transformer branch , represents the prediction from the CNN branch , p represents a pixel point, and r n (p) represents the urban RGB image; Use the cross-distillation loss based on uncertainty to denoise the pseudo-labels, and the definition formula of the cross-distillation loss is as follows: Among them, - is the gradient stopping operation, ⊙ is the element-wise multiplication, is the uncertainty map from the Transformer branch and is the uncertainty map from the CNN branch where p represents a pixel point; The formula for optimizing the monocular depth estimation model is as follows: Among them, is the silog loss, λ 1 , λ 2 , κ, η are hyperparameters used to balance the weights of each term in the loss. |T| represents the number of pixels with values, and d n (p) represents the predicted depth map, represents the ground-truth depth map, represents the depth error of the Transformer branch in the log space, represents the depth error of the CNN branch in the log space.

2. The method for training a monocular depth estimation model based on cross-distillation according to claim 1, characterized in that, Model the uncertainty map, and the modeling formula is as follows: where d n (p) represents the predicted depth map, denotes the ground truth depth map, b is a coefficient for controlling the error tolerance, T represents the set of pixels with valid ground truth depth, and p represents a pixel point.

3. The method for training a monocular depth estimation model based on cross-distillation according to claim 2, characterized in that, Application to predict uncertainty for approximation Among them, represents the uncertainty graph of Transformer branch prediction, represents the uncertainty graph of CNN branch prediction, represents the true value of the Transformer branch uncertainty graph, represents the true value of the CNN branch uncertainty graph, where p represents a pixel point.

Citation Information

Patent Citations

  • A morphology-based scene dense depth map acquisition method, system and equipment

    CN112927251A

  • Monocular video depth estimation method based on deep convolutional network

    CN113570658A