A monocular image depth estimation method and system
Patent Information
- Application Number
- CN202311521956.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-11-15
AI Technical Summary
[0003]单目深度图像估计训练方法通常分为三种:第一种方法是直接利用高精度雷达或者是深度相机采集相关深度数据,使用深度估计网络对深度信息进行回归,学习场景的深度信息,但是,在真实世界中,深度标签的获取相对来说较为困难,限制了算法的使用
[0056]本发明具有的有益效果如下:本发明可以用于自动驾驶过程中的障碍物探测、深度估计不稳定区域的标注以及尺度恢复,考虑深度估计误差和尺度不确定性,在提升单目深度估计精度的同时解决了单目图像深度估计本身的不稳定性以及尺度问题。
Smart Images

Figure CN117437274B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of monocular image depth estimation, and in particular to a monocular image depth estimation method and system that considers depth estimation error and scale uncertainty. Background Technology
[0002] In recent years, the main application areas of autonomous mobile robots can be divided into commercial and research fields, primarily warehousing and logistics and autonomous driving, and civilian fields, primarily small sweeping robots. They are also widely used in smart agriculture and unmanned delivery. These tasks from different fields have placed higher demands on the technology of autonomous mobile robots, such as reducing reliance on sensors, improving real-time system performance, and increasing system stability. Applying depth estimation algorithms to the obstacle detection stage of autonomous mobile robots can effectively promote the development and application of obstacle avoidance technology.
[0003] Monocular depth image estimation training methods are generally divided into three types: The first method directly uses high-precision radar or depth cameras to collect relevant depth data, and uses a depth estimation network to regress the depth information to learn the depth information of the scene. However, in the real world, obtaining depth labels is relatively difficult, limiting the use of the algorithm. The second method uses stereo images and camera pose transformation, but due to the baseline problem of stereo cameras, it cannot be guaranteed that the camera is applicable to all depth cameras. The third method uses continuous video frames and simultaneously estimates the camera pose and depth to complete the training. This method is simple to implement and has low dependence on the dataset. However, due to the error in the camera pose estimation network, it cannot achieve good results. The current encoder-decoder architecture monocular depth estimation algorithm is still the mainstream algorithm. The network is trained by inter-frame transformation, constructing photometric reprojection, and depth smoothing loss, but due to continuous downsampling and the loss of detail information, the accuracy of depth estimation is affected. Furthermore, due to the inherent problems of monocular depth estimation, its true scale needs to be manually recovered, which is difficult. At the same time, the depth estimation algorithm itself has certain instability issues and needs further consideration. Summary of the Invention
[0004] To address the shortcomings of the existing technologies, this invention provides a monocular image depth estimation method and system that considers depth estimation error and scale uncertainty. This method improves the accuracy of monocular depth estimation while taking into account the instability of monocular image depth estimation itself and the scale problem of depth estimation.
[0005] To achieve the objectives of the invention described above, the present invention adopts the following technical solution:
[0006] By adding a dilated convolution downsampling module to the ResNet18 residual network encoder and using short skip connections to fuse and encode coarse-grained edge information from the upper layer, the depth estimation details of monocular images are enhanced, the loss of downsampling information is reduced, and a depth feature encoder is obtained.
[0007] A decoder neural network is constructed using a convolutional upsampling module, and long skip connections are used in each layer of the decoder neural network to fuse the semantic features of the corresponding layers of the encoder to obtain a deep decoder; a depth convolution module is used to obtain depth maps at different scales.
[0008] The ResNet18 feature encoder is used as the pose encoder.
[0009] The depth estimation network for the image and the pose estimation network for the camera are constructed using the depth feature encoder, depth decoder and pose encoder.
[0010] Establish loss functions that minimize photometric reprojection error, depth smoothing, geometric consistency, and feature point reprojection, and train depth estimation and pose estimation networks.
[0011] The images between adjacent frames are re-stitched on the channel, the pose between the two frames is recalculated, and the pixel-level geometric consistency loss function is recalculated to mark the unstable regions of the depth estimation network.
[0012] For the output depth map, the depth map is back-projected into a pseudo point cloud using camera intrinsics, and the ground point cloud is segmented using an SVM classifier pre-trained on the road dataset to obtain ground point cloud information.
[0013] The image depth map scale information is reconstructed using camera altitude, camera intrinsic parameters, camera and ground point cloud information.
[0014] Furthermore, the specific steps for establishing the loss function that minimizes the photometric reprojection error, the depth smoothing loss function, the geometric consistency loss function, and the feature point reprojection loss function, and training the depth estimation network and the pose estimation network, are as follows:
[0015] Using continuous frame video stream images from the Kitti public dataset, the images are normalized, randomly cropped, horizontally flipped, and photometrically varied. At the same time, the camera's intrinsic parameter matrix is modified accordingly to adapt to the reprojection process.
[0016] Assume there exists a -1 frame f -1 The image consists of three consecutive frames: frame 0 (f0) and frame 1 (f1). Frame 0 (f0) is fed into a depth estimation network to obtain the depth map D0. Simultaneously, a camera pose estimation network is used to calculate the depth maps from f0 to f1. -1 The relationship between camera poses P 0→-1 and f0 to f-1 Camera pose P between 0→1 At the same time, using P 0→-1 Depth map D0 and -1 frame f -1 Reconstruct frame 0 (f0) and obtain the reconstructed frame 0 (f′). -1→0 Similarly, the reconstructed graph f′ of frame f1 from frame f0 is obtained. 1→0 Simultaneously calculate f′ -1→0 The photometric reprojection error between f0 and f′ 1→0 The photometric reprojection error between f0 and f0 is calculated, and the minimum reprojection error is obtained at each pixel to construct a loss function that minimizes the photometric reprojection error.
[0017] A depth smoothing loss function is constructed using the depth map output by the depth estimation network and the corresponding input RGB image.
[0018] The depth map of the source frame is reconstructed into the depth map of the target frame using the depth map of the source frame, the depth map of the target frame, and the geometric consistency loss function is constructed using the geometric information of the camera pose.
[0019] The LoFTR feature point matching network is used to obtain the feature points that match the source image and the target image. The depth value is sampled on the depth image of the target image. The projection points of the feature points on the source image are calculated by using the depth value and the camera pose. The feature point reprojection loss function is constructed.
[0020] After training, the RGB image is normalized and then input into the depth estimation network, which outputs the depth map of the monocular image. The pose estimation network is then used to obtain the pose relationship between the two frames by inputting two consecutively output images.
[0021] Furthermore, the depth convolution downsampling module consists of dilated convolution operations, batch normalization operations, point convolution operations, and the GELU activation function. The downsampling method using dilated convolution is point convolution and batch normalization operations for the channel space, as well as spatial grouping dilated convolution and the GELU activation function.
[0022] The dilated convolution has an expansion rate of 2, a stride of 2, a kernel size of 3*3, and the number of groups is consistent with the number of channels of the input feature; the point convolution has a stride of 1, a kernel size of 1*1, and a number of groups of 1.
[0023] Furthermore, a multi-scale feature fusion module is added at the end of the deep feature encoder. This module consists of a dilated spatial convolutional pooling pyramid module and a compression excitation module, used to further extract features and details at different scales. The specific extraction process includes:
[0024] In the dilated spatial convolution pooling pyramid module, dilated convolutions with dilation rates of 1, 3, 6, 9, and 12 are used to extract multi-scale features and concatenate the features along the channel dimension. After concatenation, 1x1 point convolutions are used to fuse and compress the features along the channel. Then, the compression activation module is used to calculate the weights of the feature blocks along the channel dimension. Finally, the features are weighted along the channel dimension using these weights.
[0025] Furthermore, the loss function for minimizing photometric reprojection error is:
[0026]
[0027] In the formula, L p To minimize photometric reprojection error, I t 、I′ t These represent the current frame and the reconstructed frame, respectively; Pe is the photometric reprojection error, which is calculated as follows:
[0028]
[0029] Where λ is an optional hyperparameter;
[0030]
[0031] In the formula, x and y represent the windows at the same location corresponding to the two graphs, and μ x and μ y This represents the average value within the corresponding window. Let x be the variance of all data in window x. Let σ be the variance of all data in window y. xy Let c1 = (k1, L) be the covariance of the x and y data in the window. 2 c2 = (k2, L) 2 k1 = 0.01, k2 = 0.03, L = 1.0, and the window size is 3*3.
[0032] Furthermore, the depth smoothing loss function is:
[0033]
[0034] Among them, f p For the input frame image, f p x is the mean gradient value of pixels in the horizontal direction of the input frame image, f p y represents the mean gradient value of pixels in the vertical direction of the input frame image, and d represents the mean gradient value of pixels in the vertical direction. p The depth map estimated by the depth network corresponding to the input frame, d p x represents the gradient value of the pixels in the horizontal direction of the output depth image, and d py represents the gradient value of the pixels in the vertical direction of the output depth map; v represents the total number of pixels in the map;
[0035] The geometric consistency loss function is:
[0036]
[0037] in,
[0038]
[0039] Among them, D diff (p) represents the depth consistency error for each point, D′ f0 (p) is the estimated depth map. This is a depth map reconstructed from the depth map of the previous frame.
[0040] Furthermore, the feature point reprojection loss function is:
[0041]
[0042] Where, x r y is the x-coordinate of the reprojection coordinates of the feature points in the target image onto the source image. r y is the ordinate value of the reprojection coordinates of the feature points of the target image to the source image; v is the total number of pixels in the image.
[0043] Furthermore, the pre-training process of the SVM classifier is as follows: Classes 40, 44, 48, and 49 of the Kitti_Semantic road point cloud classification dataset are designated as positive ground point cloud classes, and the rest as negative classes; after completing the dataset partitioning, the point cloud classification data is input into the SVM classifier, a 2-class cross-entropy loss function is constructed, and the SVM classifier is trained using the SGD optimizer to obtain an SVM classifier capable of classifying ground point clouds;
[0044] The method for restoring the scale information of the image depth map is as follows: Obtain the segmented ground point cloud (Point). grund Projecting these point cloud data back onto the pixel plane to obtain the y-coordinates of these points in the pixel coordinate system. ground After obtaining the data, the relative height h of the camera is calculated using the similarity theorem. relative :
[0045]
[0046] Among them, f y D is the vertical focal length of the camera; relative The output of the depth estimation network;
[0047] Each Point ground Both obtain an h through the above formularelative Then calculate h relative The median h of relative_m h absolute The absolute height of the camera, used to determine the scaling ratio. Let the absolute depth be denoted as D. absolute =αD relative .
[0048] Furthermore, using the geometric consistency error construction method, the pixel-level geometric consistency error is recalculated using the current frame depth map, the previous frame depth map, the current frame RGB image, the previous frame RGB image pair, and the camera pose between the two frames. The positions of the top 20% of pixels with the largest error are set to 1, and the remaining pixels are set to 0, thus binarizing the error.
[0049] The present invention also provides a monocular image depth estimation system that considers depth estimation error and scale uncertainty, comprising:
[0050] Encoder module: The ResNet18 residual network encoder is enhanced by adding a dilated convolution downsampling module and short skip connections to fuse and encode coarse-grained edge information from the upper layer, thereby improving the details of depth estimation and obtaining a depth feature encoder.
[0051] Decoder module: Constructs a decoder neural network using a convolutional upsampling module, and fuses the semantic features of the corresponding encoder layers in each layer of the decoder neural network using long skip connections to obtain a deep decoder; uses a depth convolution module to obtain depth maps at different scales;
[0052] Estimation network construction module: Construct a depth estimation network using the aforementioned depth feature encoder and depth decoder; Construct a camera pose estimation network using the ResNet18 feature encoder as a pose encoder;
[0053] Loss function module: Establish loss functions for minimizing photometric reprojection error, depth smoothing, geometric consistency, and feature point reprojection, and train depth estimation network and pose estimation network;
[0054] Unsupervised monocular depth estimation training module: Utilizing continuous frame video stream images from the Kitti public dataset, the images are normalized, randomly cropped, horizontally flipped, and subjected to photometric variations. Simultaneously, the camera's intrinsic parameter matrix is modified accordingly to adapt to the reprojection process. A depth-minimizing reprojection error is constructed using the continuous frame RGB images, depth estimation network, camera pose estimation network-obtained depth, and camera pose. A depth smoothing loss function is constructed using the depth map output by the depth estimation network and the corresponding input RGB image. The source frame depth map is reconstructed to the target frame depth map using the source frame depth map, target frame depth map, and camera pose. A geometric consistency loss function is constructed using the geometric information of the camera pose. The LoFTR feature point matching network is used to obtain matching feature points between the source and target images. Depth values are sampled on the target image depth map. The projection points of the feature points onto the source image are calculated using the depth values and camera pose, thus constructing a feature point reprojection loss function.
[0055] Application Module: After training, the RGB images are normalized and input into the depth estimation network, outputting a depth map of the monocular image. The pose estimation network obtains the pose relationship between two frames by inputting two consecutively output images. The images between adjacent frames are re-stitched on the channel, the pose between the two frames is recalculated, and the pixel-level geometric consistency loss function is recalculated to mark unstable regions in the depth estimation. For the output depth map, the depth map is back-projected into a pseudo-point cloud using camera intrinsics, and the ground point cloud is segmented using an SVM pre-trained on a road dataset to obtain ground point cloud information. The scale information of the image depth map is restored using camera height, camera intrinsics, camera and ground point cloud information.
[0056] The beneficial effects of this invention are as follows: This invention can be used for obstacle detection, labeling of unstable depth estimation regions, and scale recovery in the process of autonomous driving. Considering depth estimation error and scale uncertainty, it solves the instability and scale problem of monocular image depth estimation itself while improving the accuracy of monocular depth estimation. Attached Figure Description
[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 This is a structural diagram of the depth estimation network in a specific embodiment of the present invention;
[0059] Figure 2This is a projection diagram of a point in three-dimensional space of a pinhole camera in a specific embodiment of the present invention;
[0060] Figure 3 This is a structural diagram of the dilated convolution downsampling module in a specific embodiment of the present invention;
[0061] Figure 4 This is a structural diagram of the convolutional upsampling module in a specific embodiment of the present invention;
[0062] Figure 5 This is a structural diagram of the depthwise convolution module in a specific embodiment of the present invention;
[0063] Figure 6 This is a structural diagram of the multi-scale feature fusion module in a specific embodiment of the present invention;
[0064] Figure 7 This is an image showing the SVM ground point cloud segmentation results tested on a standard dataset in a specific embodiment of the present invention;
[0065] Figure 8 This is a flowchart illustrating the calculation of a binarized depth uncertainty mask in a specific embodiment of the present invention;
[0066] Figure 9 This is a flowchart illustrating the process of recovering depth scale from ground clouds and camera altitude in a specific embodiment of the present invention. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0068] This embodiment provides a monocular image depth estimation method that considers depth estimation error and scale uncertainty, and its steps are as follows:
[0069] (1) A dilated convolutional downsampling module is added to the ResNet18 residual network encoder. This module, through short skip connections, fuses and encodes coarse-grained edge information from the upper layer, enhancing the details of monocular image depth estimation and obtaining a depth feature encoder. The dilated convolutional downsampling module is described as follows: Figure 3 As shown, dilated convolution has a larger receptive field than traditional convolutional downsampling, which can effectively reduce feature loss. At the same time, a multi-scale feature fusion module consisting of dilated spatial convolution pooling pyramid and compression excitation module is used at the end of the decoder to extract feature information at different scales and further fuse them.
[0070] (2) A decoder neural network is constructed using a convolutional upsampling module, and long skip connections are used in each layer of the decoder neural network to fuse the semantic features of the corresponding encoder layers to obtain a deep decoder. The convolutional upsampling module is as follows: Figure 4 As shown; finally, depth maps of different scales are obtained by using a depthwise convolution module, as shown in the figure. Figure 5 As shown;
[0071] (3) Use ResNet18 feature encoder as pose encoder;
[0072] (4) Construct an image depth estimation network and a camera pose estimation network using the depth feature encoder, depth decoder, and pose encoder; the depth estimation network constructed using the depth feature encoder and depth decoder has the following structure: Figure 1 As shown;
[0073] (5) Establish the loss function that minimizes photometric reprojection error, depth smoothing loss function, geometric consistency loss function and feature point reprojection loss function, and train the depth estimation network and pose estimation network;
[0074] (6) Re-stitch the adjacent frame images on the channel, recalculate the pose between the two depth frames, and recalculate the pixel-level geometric consistency loss function to mark the unstable regions of the depth estimation network.
[0075] (7) For the output depth map, the depth map is back-projected into a pseudo point cloud using camera intrinsics, and the ground point cloud is segmented using an SVM classifier pre-trained on the road dataset to obtain ground point cloud information.
[0076] (8) Use camera height, camera intrinsic parameters and ground point cloud information to reconstruct the image depth map scale information.
[0077] In step (1), the depthwise convolution downsampling module consists of dilated convolution, batch normalization, pointwise convolution, and the GELU activation function. The downsampling method using dilated convolution is pointwise convolution and batch normalization for the channel space, as well as spatial grouping dilated convolution and the GELU activation function, to reduce the loss of detailed features during downsampling. The specific process includes:
[0078] Feature extraction is performed using convolutions with a stride of 2, an expansion rate of 2, and a kernel size of 3x3. Batch normalization is then applied to the extracted features. After normalization, 1x1 point convolutions are used to fuse the features across channels. Finally, the GELU activation function is used to perform a non-linear transformation on the output features, as detailed below. Figure 3 As shown;
[0079] In step (1), a multi-scale feature fusion module is added at the end of the deep feature encoder. The multi-scale feature fusion module consists of a dilated spatial convolutional pooling pyramid module and a compression excitation module, which are used to further extract features and details at different scales. The specific extraction process includes:
[0080] In the dilated spatial convolutional pooling pyramid module, dilated convolutions with dilation rates of 1, 3, 6, 9, and 12 are used to extract multi-scale features and concatenate them along the channel dimension. After concatenation, 1x1 pointwise convolutions are used to fuse and compress the features along the channel dimension. Then, a compression activation module is used to calculate the weights of the feature blocks along the channel dimension. Finally, these weights are used to weight the features along the channel dimension, such as... Figure 6 As shown.
[0081] The specific details of step (5) are as follows:
[0082] Using continuous frame video stream images from the Kitti public dataset, the images are normalized, randomly cropped, horizontally flipped, and photometrically varied. At the same time, the camera's intrinsic parameter matrix is modified accordingly to adapt to the reprojection process.
[0083] Assume there exists a -1 frame f -1 The image consists of three consecutive frames: frame 0 (f0) and frame 1 (f1). Frame 0 (f0) is fed into a depth estimation network to obtain the depth map D0. Simultaneously, a camera pose estimation network is used to calculate the depth maps from f0 to f1. -1 The relationship between camera poses P 0→-1 and f0 to f -1 Camera pose P between 0→1 At the same time, using P 0→-1 Depth map D0 and -1 frame f -1 Reconstruct frame 0 (f0) and obtain the reconstructed frame 0 (f′). -1→0 Similarly, the reconstructed graph f′ of frame f1 from frame f0 is obtained. 1→0 Simultaneously calculate f′ -1→0 The photometric reprojection error between f0 and f′ 1→0 The photometric reprojection error between f0 and f0 is calculated, and the minimum reprojection error is obtained at each pixel to construct a loss function that minimizes the photometric reprojection error.
[0084] A depth smoothing loss function is constructed using the depth map output by the depth estimation network and the corresponding input RGB image.
[0085] The depth map of the source frame is reconstructed into the depth map of the target frame using the depth map of the source frame, the depth map of the target frame, and the geometric consistency loss function is constructed using the geometric information of the camera pose.
[0086] The LoFTR feature point matching network is used to obtain the feature points that match the source image and the target image. The depth value is sampled on the depth image of the target image. The projection points of the feature points on the source image are calculated by using the depth value and the camera pose. The feature point reprojection loss function is constructed.
[0087] After training, the RGB images are normalized and then input into the depth estimation network, which outputs the depth map of the monocular image. The pose estimation network obtains the pose relationship between the two frames by inputting the two consecutively output images.
[0088] Consider a simplified geometric model of a basic pinhole camera, with specific images as follows: Figure 2 As shown, point E = (X, Y, Z) is a point in 3D space, e = (x, y) is a point on the camera plane, and the line connecting point E and the optical center of the camera passes through point e. Using similar triangles, we can easily obtain the following formula:
[0089]
[0090] Where x and y are the coordinates in the Z plane.
[0091] If we convert points in space to homogeneous coordinates, we can write the points in space as standard homogeneous coordinates, such as E = (X, Y, Z, 1). Considering the principal point offset, the formula can be rewritten as:
[0092]
[0093] The above formula expresses how to project a point represented by homogeneous coordinates in three-dimensional space onto a pixel homogeneous coordinate system, where fX, fY, and p... x p y This refers to the camera's intrinsic parameters, specifically the focal length and offset.
[0094] The camera pose estimation network outputs a 6×1 vector, where the first three dimensions represent camera displacement and the last three dimensions represent camera rotation. In this invention, Euler angles are used for rotation in three-dimensional space. x r y r z These are represented as rotations around the x-axis, y-axis, and z-axis, respectively. The obtained rotations and displacements are then expressed as a pose transformation matrix (R|T), with the specific transformation formula as follows:
[0095] Rotation matrix in the X-axis direction:
[0096]
[0097] Rotation matrix in the Y-axis direction:
[0098]
[0099] Rotation matrix in the Z-axis direction:
[0100]
[0101] This invention assumes the rotation direction is Z→Y→X, and the overall rotation matrix of the camera is:
[0102] R zyx (r z r y r x ) = R z (r z )R y (r y )R z (r z ),
[0103] For the displacement matrix T xyz Then we can directly use the following representation:
[0104] T xyz =[T x T y T z ] T ,
[0105] The complete camera pose transformation matrix is represented as follows:
[0106] (R|T)=(R zyx |T xyz ).
[0107] The change of a point in three-dimensional space can be represented as:
[0108]
[0109] X′, Y′, and Z′ are points in the three-dimensional space after the camera pose transformation.
[0110] The reprojection process is as follows:
[0111] The data for each pixel in the corresponding depth map D can be represented as (x, y, Z), where x and y are the pixel coordinates of a point, and Z is the depth corresponding to that point. From the above equation...
[0112]
[0113] It can solve for the point (X, Y, Z) in the three-dimensional space corresponding to that point.
[0114] For a point in the above three-dimensional space, considering the camera pose transformation between two frames, we can obtain (X′, Y′, Z′);
[0115] Reconsidering the camera projection equation, (X′, Y′, Z′) can be reprojected as (x′, y′, 1).
[0116] The above (x′, y′) are the pixel coordinates corresponding to the point with pixel coordinates (x, y) in the first frame in the second frame.
[0117] The aforementioned loss function consists of four parts: minimizing the reprojection loss function, the depth smoothing loss function, the geometric consistency loss function, and the feature point reprojection loss function. These loss functions, by constructing sparse and dense photometric and geometric loss functions, ensure geometric consistency between adjacent frames, constraining the pose estimation network and depth estimation network to learn more accurate depth and pose information.
[0118] 1. Minimize reprojection error
[0119] The loss function for minimizing photometric reprojection error is:
[0120]
[0121] In the formula, L p To minimize photometric reprojection error, I t 、I′ t These represent the current frame and the reconstructed frame, respectively; Pe is the photometric reprojection error, which is calculated as follows:
[0122]
[0123] Where λ is an optional hyperparameter;
[0124]
[0125] In the formula, x and y represent the windows at the same location corresponding to the two graphs, and μ x and μ y This represents the average value within the corresponding window. Let x be the variance of all data in window x. Let σ be the variance of all data in window y. xy Let c1 = (k1, L) be the covariance of the x and y data in the window. 2 c2 = (k2, L) 2 k1 = 0.01, k2 = 0.03, L = 1.0, and the window size is 3*3.
[0126] This loss function uses the minimum value of photometric reprojection loss between multiple frames as the optimization objective, which further improves the robustness of the training process to the problem of occlusion between adjacent frames.
[0127] 2. Deep smoothing loss function:
[0128]
[0129] Among them, f p For the input frame image, f p x is the mean gradient value of pixels in the horizontal direction of the input frame image, f p y represents the mean gradient value of pixels in the vertical direction of the input frame image, and d represents the mean gradient value of pixels in the vertical direction. p The depth map estimated by the depth network corresponding to the input frame, d p x represents the gradient value of the pixels in the horizontal direction of the output depth image, and d p y represents the gradient value of the pixels in the vertical direction of the output depth map; v represents the total number of pixels in the map.
[0130] The depth smoothing loss function encourages smoothness in depth estimation within objects by constraining the consistency between depth variations and color variations in RGB images.
[0131] 3. Geometric consistency loss:
[0132]
[0133] in,
[0134]
[0135] Among them, D diff (p) represents the depth consistency error for each point, D′ f0 (p) is the estimated depth map. This is a depth map reconstructed from the depth map of the previous frame.
[0136] By forcing the output depth between frames to satisfy the geometric relationship in three-dimensional space, the scale consistency of the network output depth is guaranteed.
[0137] 4. Point reprojection loss function:
[0138]
[0139] Where, x r y is the x-coordinate of the reprojection coordinates of the feature points in the target image onto the source image. r y is the ordinate value of the reprojection coordinates of the feature points in the target image to the source image; v is the total number of feature points in the image.
[0140] Compared to dense photometric loss, sparse feature point reprojection loss can improve the estimation accuracy of pose estimation networks.
[0141] Traditional monocular depth estimation algorithms cannot provide absolute scale information, as the absolute scale ratio is usually given manually. This leads to unsupervised monocular depth estimation algorithms failing to generalize well in different environments. This invention provides an absolute scale recovery algorithm for recovering absolute scale information.
[0142] The pre-training process of the SVM classifier is as follows: Classes 40, 44, 48, and 49 of the Kitti_Semantic road point cloud classification dataset are designated as positive ground point cloud classes, and the rest as negative classes. After the dataset is divided, the point cloud classification data is input into the SVM classifier, a binary cross-entropy loss function is constructed, and the SVM classifier is trained using the SGD optimizer to obtain an SVM classifier capable of classifying ground point clouds.
[0143] The method for restoring the scale information of the image depth map is as follows: Obtain the segmented ground point cloud (Point). grund These point cloud data are then projected back onto the pixel plane to obtain the coordinates y of these points in the pixel coordinate system. ground After obtaining the data, the relative height h of the camera is calculated using the similarity theorem. relative :
[0144]
[0145] Among them, f y D is the vertical focal length of the camera; relative The output of the depth estimation network;
[0146] Each Point ground Both obtain an h through the above formula relative Then calculate h relative The median h of relative_m h absolute The absolute height of the camera, used to determine the scaling ratio. Let the absolute depth be denoted as D. absolute =αD relative The process of reconstructing depth scale from ground clouds and camera altitude is as follows: Figure 9 As shown.
[0147] Monocular depth estimation algorithms are often not very stable and are easily affected by the environment. This invention proposes a binarized mask for labeling uncertain regions, and the specific implementation is as follows:
[0148] After training, the RGB images are normalized and then input into the depth estimation network to obtain the depth map of the monocular image. The images of adjacent frames are re-stitched along the channels, the pose between the two frames is recalculated, and the geometric consistency loss function is recalculated. Then, the geometric consistency losses are sorted, with the top 20% of the loss regions set to 1 and the rest set to 0, serving as a pixel uncertainty mask to label depth inconsistencies. The specific process is as follows... Figure 8 As shown.
[0149] Ground point cloud acquisition algorithm:
[0150] To recover the absolute depth scale, ground point cloud information is needed. This embodiment uses an SVM classifier with a linear kernel to segment the ground point cloud. The specific training process is as follows: Classes 40, 44, 48, and 49 of the Kitti_Semantic road point cloud classification dataset are designated as positive ground point clouds, and the rest are classified as negative. After completing the dataset segmentation, the point cloud data is input into the SVM classifier, a binary cross-entropy loss function is constructed, and the SGD optimizer is used to train the linear kernel SVM ground point cloud to obtain an SVM classifier capable of segmenting the ground point cloud. The classification is then tested on the corresponding test set, and the final classification results are as follows. Figure 7 As shown in the table below, the accuracy rates of the classification tests are as follows.
[0151] time 16ms 0.9ms 2ms accuracy 97% 67% 86%
[0152] Specifically, the binary cross-entropy loss function mentioned above is:
[0153] L(y,t)=-t*log(y)-(1-t)*log(1-y)
[0154] Where L is the loss function, y is the model output, and t is the corresponding true label.
[0155] This embodiment also provides a monocular image depth estimation system, which includes:
[0156] Encoder module: The ResNet18 residual network encoder is augmented with dilated convolutions and short skip connections to fuse and encode coarse-grained edge information from the upper layer, thereby enhancing the details of depth estimation and obtaining a depth feature encoder;
[0157] Decoder module: Constructs a decoder neural network using convolution and upsampling, and fuses the semantic features of the corresponding encoder layers in each layer of the decoder neural network using long skip connections to obtain a deep decoder;
[0158] Estimation network construction module: Construct a depth estimation network using the aforementioned depth feature encoder and depth decoder; Construct a camera pose estimation network using the ResNet18 feature encoder as a pose encoder;
[0159] Loss function module: Establish loss functions for minimizing photometric reprojection error, depth smoothing, geometric consistency, and feature point reprojection, and train depth estimation network and pose estimation network;
[0160] Unsupervised monocular depth estimation training module: Utilizing continuous frame video stream images from the Kitti public dataset, the images are normalized, randomly cropped, horizontally flipped, and subjected to photometric variations. Simultaneously, the camera's intrinsic parameter matrix is modified accordingly to adapt to the reprojection process. A depth-minimizing reprojection error is constructed using the continuous frame RGB images, depth estimation network, camera pose estimation network-obtained depth, and camera pose. A depth smoothing loss function is constructed using the depth map output by the depth estimation network and the corresponding input RGB image. The source frame depth map is reconstructed to the target frame depth map using the source frame depth map, target frame depth map, and camera pose. A geometric consistency loss function is constructed using the geometric information of the camera pose. The LoFTR feature point matching network is used to obtain matching feature points between the source and target images. Depth values are sampled on the target image depth map. The projection points of the feature points onto the source image are calculated using the depth values and camera pose, thus constructing a feature point reprojection loss function.
[0161] Application Module: After training, the RGB images are normalized and input into the depth estimation network, outputting a depth map of the monocular image. The pose estimation network obtains the pose relationship between two frames by inputting two consecutively output images. The images between adjacent frames are re-stitched on the channel, the pose between the two frames is recalculated, and the pixel-level geometric consistency loss function is recalculated to mark unstable regions in the depth estimation. For the output depth map, the depth map is back-projected into a pseudo-point cloud using camera intrinsics, and the ground point cloud is segmented using an SVM pre-trained on a road dataset to obtain ground point cloud information. The scale information of the image depth map is restored using camera height, camera intrinsics, camera and ground point cloud information.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for depth estimation in monocular images, characterized in that, include: By adding a dilated convolutional downsampling module to the ResNet18 residual network encoder and using short skip connections to fuse and encode coarse-grained edge information from the upper layer, the depth estimation details of monocular images are enhanced, and a depth feature encoder is obtained. A decoder neural network is constructed using a convolutional upsampling module, and long skip connections are used in each layer of the decoder neural network to fuse the semantic features of the corresponding layers of the encoder to obtain a deep decoder; a depth convolution module is used to obtain depth maps at different scales. The ResNet18 feature encoder is used as the pose encoder. The depth estimation network for the image and the pose estimation network for the camera are constructed using the depth feature encoder, depth decoder and pose encoder. Establish loss functions that minimize photometric reprojection error, depth smoothing, geometric consistency, and feature point reprojection, and train depth estimation and pose estimation networks. The images between adjacent frames are re-stitched on the channel, the pose between the two frames is recalculated, and the pixel-level geometric consistency loss function is recalculated to mark the unstable regions of the depth estimation network. For the output depth map, the depth map is back-projected into a pseudo point cloud using camera intrinsics, and the ground point cloud is segmented using an SVM classifier pre-trained on the road dataset to obtain ground point cloud information. Using camera altitude, camera intrinsic parameters, camera and ground point cloud information, the image depth map scale information is reconstructed; The pre-training process of the SVM classifier is as follows: Classes 40, 44, 48, and 49 of the Kitti_Semantic road point cloud classification dataset are designated as positive ground point cloud classes, and the rest as negative classes. After the dataset is divided, the point cloud classification data is input into the SVM classifier, a binary cross-entropy loss function is constructed, and the SVM classifier is trained using the SGD optimizer to obtain an SVM classifier capable of classifying ground point clouds. The method for restoring the scale information of the image depth map is as follows: Obtain the segmented ground point cloud. Projecting these point cloud data back onto the pixel plane to obtain the coordinates of these points in the pixel coordinate system After obtaining the data, the relative height of the camera is calculated using the similarity theorem. : in, The vertical focal length of the camera; The output of the depth estimation network; Each All are obtained through the above formula Then seek the median of , The absolute height of the camera, used to determine the scaling ratio. Denote the absolute depth as ; Using the geometric consistency loss function construction method, the pixel-level geometric consistency error is recalculated using the current frame depth map, the previous frame depth map, the current frame RGB image, the previous frame RGB image, and the camera pose between the two frames. The first 20% of pixels with the largest error are set to 1, and the remaining pixels are set to 0, thus binarizing the error.
2. The monocular image depth estimation method according to claim 1, characterized in that, The specific steps for establishing the loss function that minimizes the photometric reprojection error, the depth smoothing loss function, the geometric consistency loss function, and the feature point reprojection loss function, and training the depth estimation network and the pose estimation network are as follows: Using continuous frame video stream images from the Kitti public dataset, the images are normalized, randomly cropped, horizontally flipped, and photometrically varied. At the same time, the camera's intrinsic parameter matrix is modified accordingly to adapt to the reprojection process. Assuming there is a -1 frame 0 frames And 1 frame Three consecutive frames of images, frame 0 The image is fed into a depth estimation network to obtain a depth map. Simultaneously, the camera pose estimation network is used to calculate... arrive Camera pose relationship between ,as well as arrive Camera pose between At the same time, utilize Depth map and -1 frame For 0 frames Perform reconstruction and obtain the reconstructed 0 frame. Similarly, 1 frame is obtained. Frame to 0 frames Frame reconstruction graph Simultaneously calculate and The photometric reprojection error between and and The photometric reprojection error between pixels is calculated, and the minimum reprojection error is found at each pixel to construct a loss function that minimizes the photometric reprojection error. A depth smoothing loss function is constructed using the depth map output by the depth estimation network and the corresponding input RGB image. The depth map of the source frame is reconstructed into the depth map of the target frame using the depth map of the source frame, the depth map of the target frame, and the geometric consistency loss function is constructed using the geometric information of the camera pose. The LoFTR feature point matching network is used to obtain the feature points that match the source image and the target image. The depth value is sampled on the depth image of the target image. The projection points of the feature points on the source image are calculated by using the depth value and the camera pose. The feature point reprojection loss function is constructed. After training, the RGB images are normalized and then input into the depth estimation network, which outputs the depth map of the monocular image. The pose estimation network obtains the pose relationship between the two frames by inputting the two consecutively output images.
3. The monocular image depth estimation method according to claim 1 or 2, characterized in that, The depthwise convolution downsampling module consists of dilated convolution operations, batch normalization operations, point convolution operations, and the GELU activation function. The downsampling method using dilated convolution is point convolution and batch normalization operations for the channel space, and spatial grouping dilated convolution and the GELU activation function. The dilated convolution has an expansion rate of 2, a stride of 2, a kernel size of 3×3, and the number of groups is consistent with the number of channels of the input feature; the point convolution has a stride of 1, a kernel size of 1×1, and a number of groups of 1.
4. The monocular image depth estimation method according to claim 1, characterized in that, A multi-scale feature fusion module is added at the end of the deep feature encoder. This module consists of a dilated spatial convolutional pooling pyramid module and a compression excitation module, and is used to further extract features and details at different scales. The specific extraction process includes: In the dilated spatial convolution pooling pyramid module, dilated convolutions with dilation rates of 1, 3, 6, 9, and 12 are used to extract multi-scale features and concatenate the features along the channel dimension. After concatenation, 1x1 point convolutions are used to fuse and compress the features along the channel. Then, the compression activation module is used to calculate the weights of the feature blocks along the channel dimension. Finally, the features are weighted along the channel dimension using these weights.
5. The monocular image depth estimation method according to claim 1, characterized in that, The loss function for minimizing photometric reprojection error is: In the formula, To minimize photometric reprojection error, , These are the current frame and the reconstructed frame, respectively. The photometric reprojection error is calculated as follows: in, These are optional hyperparameters; In the formula, This refers to the window at the same location corresponding to both images. and This represents the average value within the corresponding window. For window The variance of all data in the dataset. For window The variance of all data in the dataset. In the window and window covariance of the data The window size is 3×3.
6. The monocular image depth estimation method according to claim 1, characterized in that, The depth smoothing loss function is: in, For the input frame image, The average gradient values of pixels in the horizontal direction of the input frame image. The average gradient values of pixels in the vertical direction of the input frame image. This is the depth map estimated by the depth network corresponding to the input frame. To output the gradient values of pixels in the horizontal direction of the depth image, To output the gradient values of pixels in the vertical direction of the depth map; This represents the total number of pixels in the image. The geometric consistency loss function is: in, in, For the depth consistency error of each point. For the estimated depth map, This is a depth map reconstructed from the depth map of the previous frame.
7. The monocular image depth estimation method according to claim 1, characterized in that, The feature point reprojection loss function is: in, The x-coordinate value of the reprojection coordinates of the feature points in the target image onto the source image. The ordinate value of the reprojection coordinates of the feature points in the target image to the source image; This represents the total number of pixels in the image.
8. A monocular image depth estimation system, characterized in that, A method for implementing the monocular image depth estimation method according to any one of claims 1-7 includes: Encoder module: The ResNet18 residual network encoder is further enhanced with dilated convolutions and short skip connections to fuse and encode the coarse-grained edge information of the upper layer, thereby enhancing the details of depth estimation and obtaining a depth feature encoder; Decoder module: Constructs a decoder neural network using convolution and upsampling, and fuses the semantic features of the corresponding encoder layers in each layer of the decoder neural network using long skip connections to obtain a deep decoder; Estimation network construction module: Construct a depth estimation network using the aforementioned depth feature encoder and depth decoder; Construct a camera pose estimation network using the ResNet18 feature encoder as a pose encoder; Loss function module: Establish loss functions for minimizing photometric reprojection error, depth smoothing, geometric consistency, and feature point reprojection, and train depth estimation network and pose estimation network; Unsupervised monocular depth estimation training module: Utilizing continuous frame video stream images from the Kitti public dataset, the images are normalized, randomly cropped, horizontally flipped, and subjected to photometric variations. Simultaneously, the camera's intrinsic parameter matrix is modified accordingly to adapt to the reprojection process. A depth-minimizing reprojection error is constructed using the continuous frame RGB images, depth estimation network, camera pose estimation network-obtained depth, and camera pose. A depth smoothing loss function is constructed using the depth map output by the depth estimation network and the corresponding input RGB image. The source frame depth map is reconstructed to the target frame depth map using the source frame depth map, target frame depth map, and camera pose. A geometric consistency loss function is constructed using the geometric information of the camera pose. The LoFTR feature point matching network is used to obtain matching feature points between the source and target images. Depth values are sampled on the target image depth map. The projection points of the feature points onto the source image are calculated using the depth values and camera pose, thus constructing a feature point reprojection loss function. Application Module: After training, the RGB images are normalized and input into the depth estimation network, outputting a depth map of the monocular image. The pose estimation network obtains the pose relationship between two frames by inputting two consecutively output images. The images between adjacent frames are re-stitched on the channel, the pose between the two frames is recalculated, and the pixel-level geometric consistency loss function is recalculated to mark unstable regions in the depth estimation. For the output depth map, the depth map is back-projected into a pseudo-point cloud using camera intrinsics, and the ground point cloud is segmented using an SVM pre-trained on a road dataset to obtain ground point cloud information. The scale information of the image depth map is restored using camera height, camera intrinsics, camera and ground point cloud information.
Citation Information
Patent Citations
Automatic crack detection method based on cavity convolution
CN111179244A
Target positioning method based on monocular depth estimation and scale recovery
CN116402870A