A self-supervised monocular depth estimation method and device

Through the joint training of the self-training mechanism and the self-improvement mechanism, the teacher model and the student model generate high-precision depth maps and uncertainty maps, solving the problem of low interpretability of the monocular depth estimation algorithm in real scenarios, and improving the accuracy and usability of depth estimation.

CN114022799BActive Publication Date: 2025-07-11NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111117413.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-23
Publication Date
2025-07-11
Estimated Expiration
2041-09-23

AI Technical Summary

Technical Problem

The existing monocular depth estimation algorithm has a low explanatory problem when deploying in real scenarios. How to apply uncertainty in the monocular self-supervised depth estimation algorithm is still a problem to be solved.

Method used

The self-training mechanism and self-improvement mechanism are adopted, through the joint training of the teacher model and the student model, the self-supervised training is used to use the unlabeled video data to generate high-precision depth maps and depth uncertainty maps, optimize the teacher model and constrain the training process of the student model.

Benefits of technology

It improves the accuracy and availability of depth estimation, can effectively block noise, improves the convergence ability of the model, and enhances the application effect in real environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114022799B_ABST
    Figure CN114022799B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of depth estimation. More specifically, this application relates to a self-supervised monocular depth estimation method and device. The method includes: obtaining video data; inputting the video data into a trained teacher model to obtain a first depth map; inputting the video data into a trained student model to obtain a second depth map and a first depth uncertainty map; wherein, the training method of the teacher model is a self-supervised training method, the training method of the student model is a supervised training method, and the teacher model and the student model are jointly trained. This application can effectively estimate the depth map, can perceive and mask the noise existing in the depth estimation result, enable the model to achieve better estimation accuracy, and bring obvious performance improvement. This application evaluates the magnitude of the noise in the form of a depth uncertainty map, improving the usability of the depth estimation method in various application scenarios such as unmanned driving in the real environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of depth estimation technology. More specifically, this application relates to a self-supervised monocular depth estimation method and device. Background Art

[0002] Depth estimation is an advanced application for almost all mobile robots, such as autonomous driving, etc. Although existing traditional methods have more or less solved this problem through sensors such as binocular cameras, lidar, and millimeter-wave radars, these devices are often expensive and difficult to deploy. Therefore, people have gradually become interested in using monocular cameras with low cost, simple deployment, and high resolution to achieve depth estimation.

[0003] Nowadays, deep learning-based methods have shown powerful performance in many image processing tasks. Neural networks can directly recover depth information from a single image through supervised learning methods. However, these methods require a large number of depth maps with accurate annotations as labels, thus limiting their generalization ability. Existing work combines the depth estimation task with the pose estimation task under the assumption of photometric invariance of images, and proposes a novel self-supervised training paradigm. This self-supervised paradigm uses continuous image data as input, and takes the difference (i.e., photometric error) between the target frame and the reconstructed new image as the supervision signal, achieving an accuracy similar to that of supervised methods.

[0004] However, the inherent low interpretability problem of deep learning still exists, hindering its deployment and application in real scenarios. In other words, how to apply uncertainty in monocular self-supervised depth estimation algorithms remains an unsolved problem. Summary of the Invention

[0005] Based on the above technical problems, the present invention aims to provide a self-supervised monocular depth estimation method and device with a self-training mechanism and a self-improving mechanism. The teacher model adopts the self-training mechanism, the student model adopts the self-improving mechanism, the teacher model and the student model are jointly trained, the trained teacher model can predict a high-precision depth map, and the trained student model can predict a high-precision depth map and a depth uncertainty map.

[0006] The first aspect of the present invention provides a self-supervised monocular depth estimation method, including:

[0007] Obtain video data;

[0008] Input the video data into the trained teacher model to obtain a first depth map;

[0009] Input the video data into the trained student model to obtain a second depth map and a first depth uncertainty map;

[0010] Among them, the training method of the teacher model is a self-supervised training method, the training method of the student model is a supervised training method, and the teacher model and the student model are jointly trained.

[0011] Specifically, the joint training of the teacher model and the student model includes:

[0012] Load unannotated video data into the teacher model;

[0013] Train the teacher model using a self-supervised training method to predict the third depth map;

[0014] Create a depth estimation task dataset with pseudo-labels based on the third depth map;

[0015] Use the depth estimation task dataset with pseudo-labels to perform supervised training on the student model.

[0016] Further, after using the depth estimation task dataset with pseudo-labels to perform supervised training on the student model, it further includes:

[0017] Determine whether the teacher model and the student model converge. If they converge, end the training;

[0018] If they do not converge, predict the second depth uncertainty map and calculate the depth uncertainty mask based on the second depth uncertainty map;

[0019] Use the depth uncertainty mask to optimize the teacher model.

[0020] Preferably, after loading the unannotated video data into the teacher model, it further includes:

[0021] Determine whether there is a depth uncertainty mask in the teacher model. If so, load the depth uncertainty mask. If not, load a blank mask.

[0022] Furthermore, the calculation formula for calculating the depth uncertainty mask based on the second depth uncertainty map is:

[0023]

[0024] Among them, ∑ s represents the second depth uncertainty map, and P 95% represents the 95th percentile in the depth uncertainty map.

[0025] More specifically, the teacher model includes a depth estimation network, an image feature extraction network, and a pose estimation network. Inputting the video data into the trained teacher model to obtain the first depth map includes:

[0026] Inputting the video data into the trained teacher model;

[0027] The depth estimation network outputs a first depth map.

[0028] Preferably, the teacher model is also trained and optimized by backpropagation, specifically including:

[0029] Inputting two consecutive frames of training images into the teacher model, where the two consecutive frames of training images include a target frame and a reference frame;

[0030] Sending the target frame to the depth estimation network to obtain a fourth depth map;

[0031] After processing the target frame and the reference frame through an image feature extraction network, splicing them together by channels and then sending them into the pose estimation network to obtain a pose transformation matrix;

[0032] Using the fourth depth map, the pose transformation matrix and the reference frame to obtain a reconstructed frame of the target frame through backprojection and bilinear interpolation;

[0033] Training and optimizing the teacher model by backpropagation based on the difference between the reconstructed frame and the target frame.

[0034] Additionally preferably, constraining the second depth map and the first depth uncertainty map of the student model based on the first depth map of the teacher model, and the constructed loss function is:

[0035]

[0036] where V represents all pixel points in the image, D t represents the first depth map, D s represents the second depth map, and ∑ s represents the first depth uncertainty map.

[0037] The second aspect of the present invention provides a self-supervised monocular depth estimation device, and the device includes:

[0038] An acquisition module for acquiring video data;

[0039] A depth map obtaining module for inputting the video data into the trained teacher model to obtain a first depth map;

[0040] An uncertainty map obtaining module for inputting the video data into the trained student model to obtain a second depth map and a first depth uncertainty map;

[0041] where the training method of the teacher model is a self-supervised training method, the training method of the student model is a supervised training method, and the teacher model and the student model are jointly trained.

[0042] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to perform the following steps:

[0043] Obtain video data;

[0044] Input the video data into the trained teacher model to obtain a first depth map;

[0045] Input the video data into the trained student model to obtain a second depth map and a first depth uncertainty map;

[0046] Wherein, the training method of the teacher model is a self-supervised training method, the training method of the student model is a supervised training method, and the teacher model and the student model are jointly trained.

[0047] The beneficial effects of the present application are as follows: The method described in the present application can effectively estimate the depth map, can sense and mask the noise existing in the depth estimation result, so that the model reaches a better estimation accuracy, bringing an obvious performance improvement. The present application evaluates the magnitude of the noise in the form of a depth uncertainty map, improving the usability of the depth estimation method in various application scenarios such as unmanned driving in the real environment. The teacher model optimized by the depth uncertainty mask can predict a more reasonable result, and the depth map predicted by the teacher model is used to constrain the depth map and the depth uncertainty map predicted by the student model, so that the student model further improves the accuracy of the depth estimation result. In addition, the present application can also exclude the interference of inaccurate depth estimation results on the self-supervised signal, can effectively solve the problem that the photometric error and the geometric consistency error are sensitive to the depth estimation noise when the model approaches convergence, and improves the convergence ability of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings forming a part of the specification illustrate embodiments of the present application and, together with the description, are used to explain the principles of the present application.

[0049] Referring to the drawings, the present application can be more clearly understood according to the following detailed description, wherein:

[0050] Figure 1 The schematic diagram of the method steps of the exemplary embodiment of the present application is shown;

[0051] Figure 2 The flowchart of the method in the exemplary embodiment of the present application is shown;

[0052] Figure 3 The schematic diagram of the working process of the teacher model in the exemplary embodiment of the present application is shown;

[0053] Figure 4Shows a schematic diagram of the working process of the self-lifting structure in an exemplary embodiment of the present application;

[0054] Figure 5 Shows a schematic diagram of depth error in an exemplary embodiment of the present application;

[0055] Figure 6 Shows a schematic diagram of the device structure in an exemplary embodiment of the present application;

[0056] Figure 7 Shows a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of the present application;

[0057] Figure 8 Shows a schematic diagram of a storage medium provided in an exemplary embodiment of the present application. Detailed implementation manners

[0058] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application. It is obvious to those skilled in the art that the present application can be implemented without one or more of these details. In other examples, some well-known technical features in the art are not described to avoid confusion with the present application.

[0059] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0060] Now, exemplary embodiments according to the present application will be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments set forth herein. The drawings are not drawn to scale, and some details may be enlarged for the purpose of clear expression, and some details may be omitted. The shapes of various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are only exemplary. In practice, there may be deviations due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0061] The following is combined with the specification appendixFigure 1-8 Several embodiments are given to describe the exemplary embodiments according to the present application. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application can be applied to any applicable scenario.

[0062] Embodiment 1:

[0063] This embodiment implements a self-supervised monocular depth estimation method, as Figure 1 shown, including:

[0064] S1. Obtain video data;

[0065] S2. Input the video data into the trained teacher model to obtain the first depth map;

[0066] S3. Input the video data into the trained student model to obtain the second depth map and the first depth uncertainty map;

[0067] Among them, the training method of the teacher model is a self-supervised training method, the training method of the student model is a supervised training method, and the teacher model and the student model are jointly trained.

[0068] Specifically, the joint training of the teacher model and the student model includes:

[0069] Load unlabeled video data into the teacher model;

[0070] Train the teacher model by adopting a self-supervised training method to predict the third depth map;

[0071] Create a depth estimation task dataset with pseudo-labels based on the third depth map;

[0072] Use the depth estimation task dataset with pseudo-labels to conduct supervised training on the student model.

[0073] Furthermore, after using the depth estimation task dataset with pseudo-labels to conduct supervised training on the student model, it further includes:

[0074] Judge whether the teacher model and the student model converge. If they converge, end the training;

[0075] If they do not converge, predict the second depth uncertainty map, and calculate the depth uncertainty mask based on the second depth uncertainty map;

[0076] Use the depth uncertainty mask to optimize the teacher model.

[0077] Preferably, after loading the unannotated video data into the teacher model, the following steps are further included:

[0078] Determine whether there is a depth uncertainty mask in the teacher model. If so, load the depth uncertainty mask; if not, load a blank mask.

[0079] Furthermore, the calculation formula for calculating the depth uncertainty mask based on the second depth uncertainty map is:

[0080]

[0081] where ∑ s represents the second depth uncertainty map, and P 95% represents the 95th percentile in the depth uncertainty map. It should be noted here that the depth map is a two-dimensional matrix with the same length and width as the image. Each element in this matrix represents the depth predicted by the model at the corresponding image coordinate, that is, the distance from the object in the real world to the camera. In actual application scenarios, such as the driverless scenario, we can judge the distance of the obstacles in front through the depth map and construct a three-dimensional map of the scene, so as to guide the driverless vehicle to adjust the path in time to avoid collisions. However, there is no index or method to represent the quality of the depth map. In some scenarios, the neural network may estimate completely wrong depths. In this case, the driverless vehicle may not be able to make the correct decision, resulting in accidents. Therefore, this application uses the depth uncertainty map to represent the quality of the depth map. The depth uncertainty map is a two-dimensional matrix with the same length and width as the depth map, and each element in it represents the quality of the depth value at the corresponding coordinate of the depth map. If the depth uncertainty of a depth value is very high, it can be considered that this depth value is probably inaccurate; if the depth uncertainty of a depth value is very low, it can be considered that this depth value is probably accurate. The 95th percentile can be understood as 100% minus 5%. It is to arrange all the elements in the uncertainty map according to their sizes (the sizes of the uncertainty values). If this element is in the largest 5%, the value of the depth uncertainty mask is 0, otherwise it is 1.

[0082] More specifically, the teacher model includes a depth estimation network, an image feature extraction network, and a pose estimation network. The step of inputting the video data into the trained teacher model to obtain the first depth map includes: inputting the video data into the trained teacher model; the depth estimation network outputs the first depth map.

[0083] Preferably, the teacher model is also trained and optimized by backpropagation, which specifically includes: inputting two consecutive frames of training images into the teacher model, where the two consecutive frames of training images include a target frame and a reference frame; sending the target frame to the depth estimation network to obtain a fourth depth map; processing the target frame and the reference frame through an image feature extraction network and then splicing them together by channels and sending them into the pose estimation network to obtain a pose transformation matrix; using the fourth depth map, the pose transformation matrix, and the reference frame to obtain a reconstructed frame of the target frame through back-projection and bilinear interpolation; and training and optimizing the teacher model by backpropagation based on the difference between the reconstructed frame and the target frame.

[0084] Additionally preferably, the second depth map and the first depth uncertainty map of the student model are constrained based on the first depth map of the teacher model, and the constructed loss function is:

[0085]

[0086] where V represents all pixel points in the image, D t represents the first depth map, D s represents the second depth map, and ∑ s represents the first depth uncertainty map.

[0087] The method of the present application can effectively estimate the depth map, can sense and mask the noise existing in the depth estimation result, enabling the model to achieve better estimation accuracy and bringing an obvious performance improvement. The present application evaluates the magnitude of the noise in the form of a depth uncertainty map, improving the usability of the depth estimation method in various application scenarios such as unmanned driving in a real environment. The teacher model optimized by the depth uncertainty mask can predict more reasonable results. Constraining the predicted depth map and depth uncertainty map of the student model based on the depth map predicted by the teacher model further improves the accuracy of the depth estimation result of the student model. In addition, the present application can also exclude the interference of inaccurate depth estimation results on the self-supervised signal, can effectively solve the problem that the photometric error and geometric consistency error are sensitive to depth estimation noise when the model approaches convergence, and improves the convergence ability of the algorithm.

[0088] Embodiment 2:

[0089] This embodiment provides a self-supervised monocular depth estimation method, as Figure 2As shown in the figure, first, a teacher model is constructed. Then, unannotated video data is loaded into the teacher model. Next, it is determined whether there is a depth uncertainty mask. If there is, the depth uncertainty mask is loaded; if not, a blank mask is loaded. After that, the teacher model is self-supervised trained and a dense depth map is predicted. A pseudo-dataset is created using this dense depth map. A student model is constructed and supervised training is performed on the student model. It is determined whether the teacher model and the student model converge. If they do not converge, a depth uncertainty map is predicted and a depth uncertainty mask is calculated based on the depth uncertainty map. If they converge, the process ends. Additionally, the calculated uncertainty mask is used for the optimization of the teacher model. The training is iterative until both the teacher model and the student model are well-trained.

[0090] Among them, the schematic diagram of the working process of the constructed teacher model is as Figure 3 shown, including a depth estimation network, an image feature extraction network, and a pose estimation network. The depth estimation network obtains a dense depth map.

[0091] Preferably, the teacher model is also trained and optimized through backpropagation, specifically including:

[0092] Two consecutive training images are input into the teacher model. Among them, the two consecutive training images include a target frame and a reference frame. The target frame is sent to the depth estimation network to obtain a fourth depth map. The target frame and the reference frame are processed by the image feature extraction network and then concatenated together by channels and sent into the pose estimation network to obtain a pose transformation matrix. Using the fourth depth map, the pose transformation matrix, and the reference frame, the reconstructed frame of the target frame is obtained through back-projection and bilinear interpolation. Based on the difference between the reconstructed frame and the target frame, backpropagation is used to train and optimize the teacher model.

[0093] In a specific implementation, as Figure 3 shown, given two consecutive pictures <I i , I i+1 >, where i (i > 0) is the time index, I i is the target frame, and I i+1 is the reference frame. I i is sent into the depth estimation network to predict a dense depth map I i and I i+1 are processed by the image feature extraction network and then concatenated together by channels and sent into the pose estimation network to predict the pose transformation matrix of the camera from time i to time i + 1

[0094] Since the illumination changes, exposure conditions, and camera position differences between two consecutive frames in the video stream are relatively small, it can be considered that they satisfy the photometric invariance assumption, that is, the pixel values of the corresponding pixels in the target frame and the reference frame are the same. Therefore, the depth map can be combined Pose transformation matrix and the image I at time i+1 i+1 , and the image at time i is reconstructed through back-projection and bilinear interpolation. The reconstructed image is called I i+1。 Then, the error between the real image I i and the reconstructed image I′ i is used as the supervision signal to train and optimize the teacher model using backpropagation. Among them, the loss function used includes the photometric error as:

[0095]

[0096] where V represents all the pixels in the target frame that project onto the reference frame, and |V| represents the number of these pixels. Since the pixels outside the image do not provide meaningful gradients, the present invention only selects the pixels located inside the reference frame after projection to calculate the pixel gradient. The calculation process of F is:

[0097] F(I i , I′ i ) = α||I i - I′ i ||1 + (1 - α)SSIM(I i , I′ i )

[0098] where SSIM represents the structural similarity index, which is used to calculate the difference between each small block of the two pictures of I i and I′ i . Preferably, α is set to 0.15.

[0099] The depth map predicted by the model should have similar object boundaries to the original image. Preferably, a smoothness loss function is also adopted:

[0100]

[0101] where Δ represents the pixel gradient corresponding to the image coordinate p, including the horizontal gradient and the vertical gradient.

[0102] In a specific implementation, in order to constrain the scale of the estimation result, a geometric consistency loss is also adopted, and its calculation process is:

[0103]

[0104] where D i→i′Denote the depth map after projecting the target frame onto the reference frame as D′i, and the depth map of the target frame after applying bilinear interpolation. This loss requires the model to predict results of the same scale for data at every two adjacent moments, thus ensuring the consistency of the prediction results for the entire input sequence.

[0105] Combining the above loss terms yields the optimization objective of the baseline model, i.e.:

[0106] L baseline = λ1L p + λ2L s + λ3L CG

[0107] where λ1, λ2, and λ3 represent the weights of different loss functions, which are 1.0, 0.1, and 0.5 respectively. By optimizing this baseline loss, the model can predict better results.

[0108] As a transformable implementation manner, in an actual application scenario, it is possible to determine whether to adopt a depth map through the depth uncertainty map proposed in this application. For example, in a downstream application, the decision control system of an autonomous vehicle can set a threshold by itself, discard depth values with uncertainty higher than the threshold, and only adopt depth estimation results with lower uncertainty for decision-making.

[0109] Embodiment 3:

[0110] This embodiment provides a self-supervised monocular depth estimation method, including: obtaining video data; inputting the video data into a trained teacher model to obtain a first depth map; inputting the video data into a trained student model to obtain a second depth map and a first depth uncertainty map; wherein, the training method of the teacher model is a self-supervised training method, the training method of the student model is a supervised training method, and the teacher model and the student model are jointly trained.

[0111] Specifically, the joint training adopts a cyclic iterative self-boosting structure in the self-boosting mechanism. Figure 4 is a schematic diagram of the working process of the self-boosting structure, as shown in Figure 4As shown, the working process of the structure includes a training step and an optimization step. The training step adopts a self-training mechanism, uses the prediction results of the teacher model as data labels, and trains the student model in a supervised manner, enabling the student model to output corresponding uncertainty estimates while predicting the depth; the optimization step further optimizes the teacher model using the obtained depth uncertainty. The teacher model is continuously trained in a self-supervised manner. At this time, the depth uncertainty map is used as a mask to shield pixels that may violate the photometric invariance assumption, and the photometric error and geometric consistency error brought by these pixels are not calculated. The training step and the optimization step are repeated to gradually improve the performance of the teacher model and the student model, and more accurate estimation results are obtained. After the teacher model and the student model are fully converged, the model weights are saved for downstream applications.

[0112] During the training process, the training step includes: loading unlabeled video data into the teacher model; training the teacher model in a self-supervised training manner to predict a third depth map; creating a depth estimation task dataset with pseudo-labels based on the third depth map; taking supervised training of the student model using the depth estimation task dataset with pseudo-labels; determining whether the teacher model and the student model converge. If they converge, the training ends; if they do not converge, a second depth uncertainty map is predicted, and a depth uncertainty mask is calculated based on the second depth uncertainty map; the teacher model is optimized using the depth uncertainty mask.

[0113] Further, after taking supervised training of the student model using the depth estimation task dataset with pseudo-labels, it also includes: after loading unlabeled video data into the teacher model, it also includes: determining whether there is a depth uncertainty mask in the teacher model. If so, the depth uncertainty mask is loaded; if not, a blank mask is loaded.

[0114] The calculation formula for calculating the depth uncertainty mask based on the second depth uncertainty map is:

[0115]

[0116] where, ∑ s represents the second depth uncertainty map, and P 95% represents the 95th percentile in the depth uncertainty map. Here, all elements in the uncertainty map are arranged in size. If this element is the largest 5%, the value of the depth uncertainty mask is 0; otherwise, it is 1.

[0117] The depth uncertainty mask is applied to the photometric error loss and geometric consistency loss of the teacher model. The updated photometric error loss is:

[0118]

[0119] The updated geometric consistency loss is as follows:

[0120]

[0121] Furthermore, the optimization step can be based on the first depth map of the teacher model to constrain the second depth map and the first depth uncertainty map of the student model, and the constructed loss function is:

[0122]

[0123] where V represents all pixel points in the image, D t represents the first depth map, D s represents the second depth map, and ∑ s represents the first depth uncertainty map.

[0124] Figure 5 As shown in the depth error schematic diagram, existing self-supervised methods consider the reprojection error of each pixel when calculating the loss. This approach performs excellently under ideal circumstances, but when the pixel depth estimated by the model differs significantly from the true depth, it may lead to inaccurate data association, resulting in extremely high photometric error penalties. As Figure 5 shown, I t and I s respectively represent the target frame and the reference frame. During the reprojection process, for a certain pixel point p on the target frame, when the depth value estimated by the model is close to the true depth value D(P), it can be reprojected to I s to obtain the corresponding matching point p′ g , providing meaningful pixel gradients. However, when the depth value estimated by the model differs significantly from the true depth value, point p may be reprojected to p′ b , pointing the model in the wrong optimization direction. In addition, since self-supervised methods are troubled by scenarios that violate the photometric invariance assumption (such as moving objects, occlusions, and artifacts), the above-mentioned regions with depth estimation errors will always exist and prevent the model from further converging. Therefore, this application uses the first depth map of the teacher model to constrain the second depth map and the first depth uncertainty map of the student model, constructs a loss function to train and optimize the student model, in order to achieve better accuracy.

[0125] In comparison with other self-supervised depth estimation methods, the experimental device is a desktop computer equipped with an Intel Core i7-9700 8-core 8-thread processor with a main frequency of 3 GHz; the memory size is 64 GB and the frequency is 3200 MHz. All experiments were completed on the Ubuntu 18.04 64-bit operating system. We trained and evaluated the proposed self-supervised depth estimation algorithm on the KITTI dataset. This dataset is one of the widely used datasets in the field of odometry, providing depth information collected by lidar and trajectory information obtained by GPS / IMU fusion.

[0126] In this application, the KITTI depth estimation dataset was used to evaluate the depth estimation results. After comparing with the original results, it was found that the teacher model optimized with depth uncertainty masks could predict more reasonable results. In addition, by detecting and eliminating the noise in the estimation results of the teacher model, the student model further improved the accuracy of the depth estimation results.

[0127] In this invention, seven metrics for evaluating depth estimation results were calculated when quantitatively evaluating the depth estimation results. Among them, AbsRel, Sq Rel, RMSE, and RMSE(log) are the errors between the model prediction results and the true depth, and the smaller this value is, the better; Acc.1, Acc.2, and Acc.3 represent the accuracy of the model prediction results, and the larger this value is, the better. The specific calculation formulas are as follows:

[0128]

[0129]

[0130]

[0131]

[0132]

[0133] Among them, represents the true depth value of a certain pixel in the test set, d i represents the predicted value of the depth of this pixel by the model, and N represents the total number of all pixels. For the Acc metric, Acc.1, Acc.2, and Acc.3 correspond to thresholds thr equal to 1.25, 1.25 2 and 1.25 3 .

[0134] As can be intuitively seen from Table 1, both the teacher model and the student model have achieved the optimal performance in the self-supervised method and can obtain an accuracy similar to that of the supervised training and the self-supervised training method using binocular constraints. This proves the effectiveness of the provided self-supervised monocular depth estimation method based on the self-training mechanism and the self-improving mechanism.

[0135] In addition to depth, we also additionally evaluated the ability of the student model to estimate depth uncertainty. Table 2 shows the performance comparison between the student model and other uncertainty estimation methods after self-improving training. Among them, Table 2(a) evaluates the impact of each method on the depth estimation result, and Table 2(b) evaluates the quality of the uncertainty estimation of each method. For the depth estimation result, we used the same metrics as in the table; for the uncertainty estimation result, we calculated the area under the sparsification error (AUSE) and the area under the random gain (AURG) for the three metrics of Abs Rel, RMSE, and 1 - Acc.1 respectively. Both AUSE and AURG are metrics derived from sparsification plots. It should be noted that for an uncertainty estimation model, the smaller the AUSE and the larger the AURG, the better its performance. As can be seen from Table 2, our student model has achieved the best depth estimation accuracy while obtaining good uncertainty estimation results.

[0136] Table 1 Evaluation Results of Depth Estimation Accuracy

[0137]

[0138] Table 2(a) Comparison of Uncertainty Estimation Method Results Table 1

[0139]

[0140]

[0141] Table 2(b) Comparison of Uncertainty Estimation Method Results Table 2

[0142]

[0143] Example 4:

[0144] This example provides a self-supervised monocular depth estimation device, as Figure 6 shown, including:

[0145] An acquisition module 601, configured to acquire video data;

[0146] A depth map obtaining module 602, configured to input the video data into a trained teacher model to obtain a first depth map;

[0147] An uncertainty map obtaining module 603, configured to input the video data into a trained student model to obtain a second depth map and a first depth uncertainty map;

[0148] Wherein, the training method of the teacher model is a self-supervised training method, the training method of the student model is a supervised training method, and the teacher model and the student model are jointly trained, that is, joint training and optimization based on a self-training mechanism and a self-improving mechanism.

[0149] The device can effectively estimate the depth map, can sense and mask the noise existing in the depth estimation result, enable the model to achieve better estimation accuracy, and bring obvious performance improvement. The device evaluates the magnitude of the noise in the form of a depth uncertainty map, improving the usability of the depth estimation method in various application scenarios such as driverless in the real environment.

[0150] Please refer to the following Figure 7 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. As Figure 7 shown, the electronic device 2 includes: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202; a computer program that can run on the processor 200 is stored in the memory 201. When the processor 200 runs the computer program, it executes the self-supervised monocular depth estimation method provided by any one of the foregoing embodiments of the present application. The electronic device may be an electronic device with a touch-sensitive display.

[0151] Wherein, the memory 201 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 203 (which may be wired or wireless), a communication connection between the system network element and at least one other network element is realized, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.

[0152] The bus 202 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store a program. After receiving an execution instruction, the processor 200 executes the program. The self-supervised monocular depth estimation method disclosed in any one of the foregoing embodiments of the present application may be applied to the processor 200 or implemented by the processor 200.

[0153] The processor 200 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 200 or the instructions in the form of software. The above-mentioned processor 200 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.

[0154] The electronic device provided in the embodiments of the present application and the self-supervised monocular depth estimation method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by it.

[0155] The embodiments of the present application also provide a computer-readable storage medium corresponding to the self-supervised monocular depth estimation method provided in the foregoing embodiments. Please refer to Figure 8 , Figure 8 The computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the self-supervised monocular depth estimation method provided in any of the foregoing embodiments.

[0156] In addition, examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.

[0157] The computer-readable storage medium provided by the above embodiments of the present application and the method for allocating quantum key distribution channels in the space-division multiplexing optical network provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0158] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor implements the steps of the self-supervised monocular depth estimation method provided by any of the foregoing embodiments. The steps of the method include: acquiring video data; inputting the video data into a trained teacher model to obtain a first depth map; inputting the video data into a trained student model to obtain a second depth map and a first depth uncertainty map; wherein, the training method of the teacher model is a self-supervised training method, the training method of the student model is a supervised training method, and the teacher model and the student model are jointly trained.

[0159] It should be noted that: The algorithms and displays provided herein are not inherently related to any particular computer, virtual device or other equipment. Various general-purpose devices can also be used in conjunction with the teachings herein. Based on the above description, the structures required to construct such devices are obvious. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of the specific language above is to disclose the best implementation mode of the present application. In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0160] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure or description thereof. However, the disclosed method should not be construed as reflecting the intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.

[0161] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and set in one or more devices different from those in the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification and all the processes or units of any method or device thus disclosed. Unless otherwise explicitly stated, each feature disclosed in this specification can be replaced by an alternative feature that provides the same, equivalent or similar purpose.

[0162] Finally, it should be emphasized that the embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0163] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that in practice, a microprocessor or a Digital Signal Processor (DSP) can be used to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of the present application. The present application can also be implemented as a device or device program for executing part or all of the methods described herein. The program for implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0164] As mentioned above, the above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A self-supervised monocular depth estimation method, characterized in that, Including: Obtain video data; Input the video data into a trained teacher model to obtain a first depth map; Input the video data into a trained student model to obtain a second depth map and a first depth uncertainty map; Among them, the training method of the teacher model is a self-supervised training method, the training method of the student model is a supervised training method, and the teacher model and the student model are jointly trained; The joint training of the teacher model and the student model includes: Load unlabeled video data into the teacher model; Train the teacher model using a self-supervised training method to predict a third depth map; Create a depth estimation task dataset with pseudo-labels based on the third depth map; Use the depth estimation task dataset with pseudo-labels to perform supervised training on the student model; After using the depth estimation task dataset with pseudo-labels to perform supervised training on the student model, it further includes: Determine whether the teacher model and the student model converge. If they converge, end the training; If they do not converge, predict a second depth uncertainty map and calculate a depth uncertainty mask based on the second depth uncertainty map; Use the depth uncertainty mask to optimize the teacher model.

2. The self-supervised monocular depth estimation method according to claim 1, characterized in that After loading unlabeled video data into the teacher model, it further includes: Determine whether there is a depth uncertainty mask in the teacher model. If so, load the depth uncertainty mask. If not, load a blank mask.

3. The self-supervised monocular depth estimation method according to claim 1, wherein The calculation formula for calculating the depth uncertainty mask based on the second depth uncertainty map is: Among them, ∑ s represents the second depth uncertainty map, and P 95% represents the 95th percentile in the depth uncertainty map.

4. The self-supervised monocular depth estimation method according to claim 1, characterized in that, The method further includes: Constraining the second depth map and the first depth uncertainty map of the student model based on the first depth map of the teacher model. The constructed loss function is: Among them, V represents all pixel points in the image, D t represents the first depth map, D s represents the second depth map, ∑ s represents the first depth uncertainty map.

5. The self-supervised monocular depth estimation method according to claim 1, wherein The teacher model includes a depth estimation network, an image feature extraction network, and a pose estimation network. The step of inputting the video data into the trained teacher model to obtain a first depth map includes: Input the video data into the trained teacher model; The depth estimation network outputs a first depth map.

6. The self-supervised monocular depth estimation method according to claim 5, wherein The teacher model is also trained and optimized through backpropagation, specifically including: Input two consecutive frames of training images into the teacher model, where the two consecutive frames of training images include a target frame and a reference frame; Send the target frame to the depth estimation network to obtain a fourth depth map; After processing the target frame and the reference frame through the image feature extraction network, splice them together by channel and send them into the pose estimation network to obtain a pose transformation matrix; Use the fourth depth map, the pose transformation matrix, and the reference frame to obtain a reconstructed frame of the target frame through backprojection and bilinear interpolation; Train and optimize the teacher model through backpropagation based on the difference between the reconstructed frame and the target frame.

7. A self-supervised monocular depth estimation device, characterized in that The device includes: An acquisition module for acquiring video data; A depth map acquisition module for inputting the video data into a trained teacher model to obtain a first depth map; An uncertainty map acquisition module for inputting the video data into a trained student model to obtain a second depth map and a first depth uncertainty map; Among them, the training method of the teacher model is a self-supervised training method, the training method of the student model is a supervised training method, and the teacher model and the student model are jointly trained; The joint training of the teacher model and the student model includes: Loading unannotated video data into the teacher model; Training the teacher model in a self-supervised training method to predict a third depth map; Creating a depth estimation task dataset with pseudo-labels based on the third depth map; Taking supervised training on the student model using the depth estimation task dataset with pseudo-labels; After taking supervised training on the student model using the depth estimation task dataset with pseudo-labels, it further includes: Judging whether the teacher model and the student model converge. If they converge, the training ends; If they do not converge, predicting a second depth uncertainty map and calculating a depth uncertainty mask based on the second depth uncertainty map; Optimizing the teacher model using the depth uncertainty mask.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Optical flow estimation method and system based on asynchronous event flow and grayscale image fusion

    CN113269699A

  • Foresight scene depth estimation method based on self-supervised learning

    CN113313732A