Complex weather scene image monocular depth estimation method based on self-supervised learning
By introducing a multi-scale attention mechanism and designing the DSR module and DE module in the self-supervised monocular depth estimation model, the problem of degradation of depth prediction accuracy in complex weather scenarios is solved, and higher quality depth map generation is achieved.
Patent Information
- Application Number
- CN202510007012.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
The existing self-supervised monocular depth estimation method has significantly reduced prediction accuracy in complex weather scenarios such as rainy and foggy days. It is mainly due to the distortion of the relationship between pixels caused by noise and environmental factors, and the supervision signal is no longer reliable.
Using a multi-scale attention self-supervised monocular depth estimation model, the multi-scale attention mechanism is introduced into the encoder to enhance the learning of the global and local features of the scene, and the DSR module and DE module are designed to filter out pseudo-depth and refine edges, thereby generating an accurate and clear depth map.
The accuracy and quality of depth prediction are significantly improved in complex weather scenarios, the efficiency and ability of images are enhanced, the problem of blurred object boundaries is solved, and excellent results are also shown in standard scenarios.
Smart Images

Figure CN119941820A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a monocular depth estimation method for complex weather scene images based on self-supervised learning. Background Art
[0002] Accurate 3D geometric information is essential for various computer vision tasks, such as 3D reconstruction positioning and mapping, mobile robots and autonomous driving. How to accurately estimate the scene depth from images is a hot topic in current research. At present, monocular depth estimation methods can be roughly divided into supervised learning-based methods and self-supervised learning-based methods. Supervised monocular depth estimation obtains real ground depth information based on sensors such as LiDAR and RGB-D cameras. Such methods have high accuracy, but the problem of poor robustness in complex weather scenes limits the practical value of such algorithms. Self-supervised monocular depth estimation is a method that uses image reconstruction as a supervisory signal to train a depth estimation model in the absence of ground truth depth data. Due to its low cost, it has become a research hotspot in the field of depth estimation in recent years. Specifically, the model inputs a set of stereo pairs or monocular sequence images. The pose estimation network (PoseNet) and the depth estimation network (DepthNet) are trained to generate the camera's self-motion and depth predictions, and these predictions are used to synthesize the current frame from adjacent frames. After that, by constraining the photometric consistency between the synthesized image and the real image, the model is able to predict the depth information in a self-supervised manner. Based on this paradigm, the existing self-supervised monocular depth estimation can provide relatively accurate depth maps under standard conditions. However, the prediction accuracy of this paradigm drops significantly in complex weather conditions such as rainy and foggy days. This is mainly because the noise introduced in the complex environment causes the relationship between pixels to be distorted, which conflicts with the ideal site preset by the algorithm, making the supervision signal no longer reliable. In rainy and foggy days, the details and textures in the environment are blocked, resulting in blurred and distorted targets; the reflection of light on slippery roads further confuses the depth information. In addition, the scattering effect of rain and fog reduces the visibility of distant targets. These factors together lead to the difficulty of extracting image depth information. In night conditions, insufficient light makes target recognition more difficult, especially small objects are more dependent on the background. When the background is blurred or the object is occluded, it is difficult for the algorithm to accurately extract the features of these objects, thus affecting the accuracy of depth estimation. Therefore, under these complex conditions, the complexity and uncertainty of self-supervised monocular depth estimation increase significantly.
[0003] At present, some researchers realize depth estimation of complex weather scenes by generating single image prior depth information from standard samples. Although this method can solve the depth prediction problem of such scenes, there is still a gap between the predicted performance and normal samples.
[0004] To address the above problems, a monocular depth estimation method for complex weather scene images based on self-supervised learning is provided. Summary of the invention
[0005] The purpose of the present invention is to overcome the existing defects and provide a monocular depth estimation method for complex weather scene images based on self-supervised learning to achieve prediction of complex weather scenes.
[0006] The technical solution to achieve the above purpose is:
[0007] A monocular depth estimation method for complex weather scene images based on self-supervised learning, including:
[0008] Step S1, obtaining a data set, dividing it into a training set, a validation set, and a test set in proportion, and preprocessing the data set;
[0009] Step S2, constructing a monocular depth estimation model based on multi-scale attention self-supervision;
[0010] Step S3, training and validating a self-supervised monocular depth estimation model based on the preprocessed data set;
[0011] Step S4, using the trained self-supervised monocular depth estimation model to test the test set and output an accurate and clear depth map.
[0012] Preferably, in step S1, the data set includes:
[0013] KITTI: is a widely used outdoor depth estimation benchmark that contains continuous stereo images and sparse points collected by sensors mounted on vehicles, including 19,905 training images and 2,212 validation images;
[0014] KITTI-C: is a comprehensive benchmark for evaluating the robustness of monocular depth estimation. This benchmark shares the same original images as the KITTI test set, simulating a variety of challenging scenarios, including lighting conditions, sensor failure or movement, and noise in data processing. The benchmark contains 18 perturbations.
[0015] NuScnece: is a challenging large-scale dataset containing 1,000 different scenes, covering different weather scenarios, including but not limited to daytime, rainy days, and nights; following the official split of R4Dyn, there are 15,129 training images and 6,019 validation images.
[0016] Preferably, in step S1, the preprocessing includes performing data enhancement and normalization processing on the data set, wherein the data enhancement method includes increasing brightness, reducing brightness and adding Gaussian noise.
[0017] Preferably, in step S2, based on the multi-scale attention self-supervised monocular depth estimation model, a monocular video training depth network and a posture network are used to train a single frame I t The purpose of monocular depth estimation is to predict its corresponding depth map D t , in self-supervised training, model supervision comes from adjacent frames I t ∈{I t-1 ,I t+1}, the self-supervised monocular depth estimation model uses the target depth estimation D t =DepthNet(I t ) predicts the depth of the current frame and uses the camera pose to estimate P t→t′ =PoseNet(I t ,I′ t ) estimates the camera’s ego-motion to the next timestamp, and then projects it onto the current timestamp to generate the corresponding synthetic object:
[0018] I t′→t =I t′ <proj(D t ,P t→t′ ,K)>;
[0019] Where K is the essence of the camera, proj(·) is the table reprojected to I′ t The camera returns to D t 2D coordinates, and (·) is the pixel sampling operator, I′ t is the adjacent frame, I t′ is the target frame;
[0020] If the prediction is correct, then I t′→t Should be with I t Same, therefore, using the minimum luminosity loss L p To constrain D t and P t→t′ The generated I a and I b The twisted flow between:
[0021] I p =min pe(I t ,I t′→t );
[0022]
[0023] In the formula, I a is the target frame, I b is the reconstructed target frame, pe(·) is SSIM and L pThe former is used to calculate the structural similarity, and the latter is used to calculate the photometric loss. θ is the weight parameter, and SSIM(·) is the image similarity index.
[0024] In addition, the predicted depth map is regularized using edge-aware smoothness loss:
[0025]
[0026] Where ω(D) is the normalized inverse depth of D, ξ x and y are the horizontal and vertical gradients respectively, D is the depth map, and I is a given single frame.
[0027] Preferably, in step S2, based on the multi-scale attention self-supervised monocular depth estimation model training, complex samples and corresponding simple samples are required, and the disturbance space is constructed from weak to strong by setting critical values;
[0028] The model trains an image translation model Translation, for images I in standard scenes t ,Use the image translation model to obtain complex samples:
[0029] T=Translation(I t );
[0030] Where Translation(·) includes different weather or lighting conditions, sensor failure or motion, and noise in the data processing process;
[0031] After that, the deep network and posture network are trained. The input data of the deep network is not just standard scene samples, so the data is normalized. Since the enhancement operations used by the self-supervised monocular depth estimation model do not change the underlying depth map, the fully trained network is used to obtain the depth map on the original image, and then the depth map is used as the label in the supervised regression loss.
[0032] Preferably, in step S2, the multi-scale attention self-supervised monocular depth estimation model is divided into three parts, namely feature extraction, depth prior and baseline refinement, including:
[0033] Step S21, embedding the multi-scale attention module into the self-supervised monocular depth estimation model to perform feature extraction;
[0034] Step S22, filtering out inaccurate pixels in the pseudo depth through the consistency output strategy DSR module in the self-supervised monocular depth estimation model;
[0035] Step S23, the predicted edge is refined by the DE edge refinement module in the self-supervised monocular depth estimation model, so as to predict an accurate and clear depth map.
[0036] Preferably, in step S21, the self-supervised monocular depth estimation model is a Unet structure, which is divided into an encoder, a skip connection and a decoder;
[0037] Based on ResNet18, a multi-scale attention mechanism is introduced as the encoder;
[0038] After extracting the feature encoder output, the multi-scale attention module is used to enhance the learning of the scene's global and local features;
[0039] The decoder makes a jump connection with the feature part of the encoder through multi-scale depth estimation, concatenates features of the same resolution, and upsamples them step by step to restore the depth map;
[0040] The multi-scale attention module contains two convolution kernels, which are placed in parallel models.
[0041] Among them, the 1x1 convolution kernel enables the model to capture local cross-channel interactions and shares similarities with channel convolution, while the 3x3 convolution branch captures multi-scale feature representations;
[0042] Feature map X∈R C×H×W The feature map is divided into G sub-features through the group branch, where the group style is:
[0043] X=[X0,X1,......,X G-1 ],X i ∈R C / G×H×W ;
[0044] In the formula, R C×H×W is the feature dimension of the original input feature map, C is the number of input channels, H is the height of the input image, and W is the dimension of the input feature;
[0045] At the same time, two parallel convolution branches are used, namely 1x1 convolution branch and 3x3 convolution branch;
[0046] The 1x1 convolution branch is mainly used to capture the dependencies between channels and generate channels through global average pooling.
[0047] The 3x3 convolution branch is responsible for extracting multi-scale spatial information, which can better capture local features and contextual information.
[0048] After that, one-dimensional global average pooling is performed through parallel 1x1 and 3x3 convolution branches;
[0049] The two branches of the 1x1 convolution perform average pooling along the horizontal and vertical dimensions respectively, thereby retaining the precise position information in different dimensional directions;
[0050] Among them, the global average pooling of the coded information along the horizontal dimension in the channel C at height H can be expressed as:
[0051]
[0052] Where, X c is the input feature of the cth channel, and W is the dimension of the input feature;
[0053] After that, the 2D global average pooling encodes the global spatial information, and the different branch outputs are converted to the corresponding dimensions before the joint activation mechanism of the channel features, i.e. and is the feature dimension of the feature map of the average pooling of the 1x1 convolution branch, The input feature dimension of the feature grouping 3x3 convolution branch feature map, The feature dimension of the feature map of the feature grouping 3x3 convolution branch average pooling, The input feature dimension of the feature grouping 1x1 convolution branch feature map, the two-dimensional global pooling operation can be expressed as:
[0054]
[0055] The two-dimensional encoding vector output by the above operation is multiplied by the matrix dot product operation to obtain the spatial attention feature map.
[0056] Preferably, in step S22, for the current frame I t and adjacent frame I t―1 or I t+1 The depth map D of the current frame is obtained through the image depth prior model deep network and posture network t and camera pose transformation P t+1→t ;
[0057] Afterwards, reconstruct the current frame I t+1→t The cost function is constructed by calculating the minimum photometric error with the target frame to filter out inaccurate pixels in the pseudo depth. The consistent output cost function is as follows:
[0058] D mask =min t+1→t pe(I t ,I t+1→t )≤σ g ;
[0059] In the formula, σ gis a predefined threshold for the reprojection loss.
[0060] Preferably, in step S23, the depth network and the posture network are trained under standard scenarios, so the predicted pseudo-depth has good smoothness, and then based on the surface normal extracted from the predicted depth and pseudo-depth, a DE edge refinement module is proposed to refine the predicted edge, that is, by sampling the surface normal vector of the pseudo-depth and the depth edge predicted by the network model, and at the same time, applying the relative normal angle loss to constrain the depth estimation of the target boundary area, the normal matching loss is:
[0061]
[0062] Where A is the surface normal of the predicted depth of complex samples by the self-supervised monocular depth estimation model, B is the surface normal of the pseudo depth, and N is the total number of pixels;
[0063] By edge-guided sampling, the sampling pair is obtained<A,B> , thus obtaining the edge-aware normal loss, namely:
[0064]
[0065] In the formula, and are the normal vectors of the sampling points from the predicted depth, and These are the normal vectors of the sampling points of pseudo depth.
[0066] The beneficial effects of the present invention are as follows: the present invention integrates a multi-scale attention mechanism, which processes different resolutions to capture richer contextual information due to the loss of feature information in local areas caused by problems such as image models, low contrast and color distortion, and encodes global information while maintaining network depth to recalibrate the channel weights in each parallel branch and further aggregate the output features of the two parallel branches through cross-dimensional interactions, thereby capturing pixel-level pairwise relationships, thereby enhancing the efficiency and capabilities of the image; at the same time, a DSR module is proposed for the pseudo depth information generated by the self-supervised monocular depth estimation model based on geometric consistency to filter out unreliable pixels and provide reliable depth information; finally, a DE edge refinement module is proposed to solve the problem of blurred object boundaries, thereby improving the quality of depth prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is a flow chart of a monocular depth estimation method for complex weather scene images based on self-supervised learning of the present invention;
[0068] Figure 2 It is a specific flow chart of constructing a multi-scale attention self-supervised monocular depth estimation model in the present invention;
[0069] Figure 3 It is a quantitative result diagram of TalentDepth of the present invention and other models in NuScene;
[0070] Figure 4 This is a quantitative result diagram of TalentDepth of the present invention and other models in KITTI;
[0071] Figure 5 This is a graph showing the test results of TalentDepth of the present invention and other models on KITTI-C;
[0072] Figure 6 This is a graph showing the results of an ablation experiment performed by TalentDepth on the NuScnece dataset. DETAILED DESCRIPTION
[0073] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. In the description of the present invention, it should be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, which are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.
[0074] The present invention will be further described below in conjunction with the accompanying drawings.
[0075] like Figure 1 As shown in FIG. 1 , a monocular depth estimation method for complex weather scene images based on self-supervised learning includes:
[0076] Step S1, obtain a data set, divide it into a training set, a validation set and a test set in proportion, and preprocess the data set.
[0077] In an embodiment, the data set includes:
[0078] KITTI: is a widely used outdoor depth estimation benchmark that contains continuous stereo images and sparse points collected by sensors mounted on vehicles, including 19,905 training images and 2,212 validation images;
[0079] KITTI-C: is a comprehensive benchmark for evaluating the robustness of monocular depth estimation. This benchmark shares the same original images as the KITTI test set, simulating a variety of challenging scenarios, including lighting conditions, sensor failure or movement, and noise in data processing. The benchmark contains 18 perturbations.
[0080] NuScnece: is a challenging large-scale dataset containing 1,000 different scenes, covering different weather scenarios, including but not limited to daytime, rainy days, and nights; following the official split of R4Dyn, there are 15,129 training images (with synchronized sensors) and 6,019 validation images (4,449 during the day, 602 at night, and 1,088 in the rain).
[0081] In an embodiment, the preprocessing includes performing data enhancement and normalization processing on the data set, wherein the data enhancement method includes increasing brightness, reducing brightness, and adding Gaussian noise.
[0082] Step S2: construct a monocular depth estimation model based on multi-scale attention self-supervision.
[0083] In the embodiment, based on the multi-scale attention self-supervised monocular depth estimation model, a monocular video training deep network and a posture network are used. t The purpose of monocular depth estimation is to predict its corresponding depth map D t , in self-supervised training, model supervision comes from adjacent frames I t ∈{I t-1 ,I t+1}, the self-supervised monocular depth estimation model uses the target depth estimation D t =DepthNet(I t ) predicts the depth of the current frame and uses the camera pose to estimate P t→t′ =PoseNet(I t ,I′ t ) estimates the camera’s ego-motion to the next timestamp, and then projects it onto the current timestamp to generate the corresponding synthetic object:
[0084] I t′→t =I t′ <proj(D t ,P t→t′ ,K)>;
[0085] Where K is the essence of the camera, proj(·) is the table reprojected to I′ t The camera returns to D t 2D coordinates, and (·) is the pixel sampling operator, I′ t is the adjacent frame, I t′ is the target frame;
[0086] If the prediction is correct, then I t′→t Should be with I t Same, therefore, using the minimum luminosity loss L p To constrain Dt and P t→t′ The generated I a and I b The twisted flow between:
[0087] L p =min pe(I t ,I t′→t );
[0088]
[0089] In the formula, I a is the target frame, I b is the reconstructed target frame, pe(·) is SSIM and L p The former is used to calculate the structural similarity, and the latter is used to calculate the photometric loss. θ is the weight parameter, and SSIM(·) is the image similarity index.
[0090] In addition, the predicted depth map is regularized using edge-aware smoothness loss:
[0091]
[0092] Where ω(D) is the normalized inverse depth of D, ξ x and y are the horizontal and vertical gradients respectively, D is the depth map, and I is a given single frame.
[0093] In the embodiment, the training of the monocular depth estimation model based on multi-scale attention self-supervision requires complex samples and corresponding simple samples, and the perturbation space is constructed from weak to strong by setting critical values;
[0094] The model trains an image translation model Translation, for images I in standard scenes t ,Use the image translation model to obtain complex samples:
[0095] Y=Translation(I t );
[0096] Where Translation(·) includes different weather or lighting conditions, sensor failure or motion, and noise in the data processing process;
[0097] After that, the deep network and posture network are trained. The input data of the deep network is not just standard scene samples, so the data is normalized. Since the enhancement operations used by the self-supervised monocular depth estimation model do not change the underlying depth map, the fully trained network is used to obtain the depth map on the original image, and then the depth map is used as the label in the supervised regression loss.
[0098] like Figure 2 As shown in the figure, the multi-scale attention self-supervised monocular depth estimation model is divided into three parts, namely feature extraction, depth prior and baseline refinement, including:
[0099] Step S21, embedding the multi-scale attention module into the self-supervised monocular depth estimation model for feature extraction.
[0100] In the embodiment, the self-supervised monocular depth estimation model is a Unet structure, which is divided into an encoder, a skip connection and a decoder;
[0101] Based on ResNet18, a multi-scale attention mechanism is introduced as the encoder;
[0102] After extracting the feature encoder output, the multi-scale attention module is used to enhance the learning of the scene's global and local features;
[0103] The decoder makes a jump connection with the feature part of the encoder through multi-scale depth estimation, concatenates features of the same resolution, and upsamples them step by step to restore the depth map;
[0104] The multi-scale attention module contains two convolution kernels, which are placed in parallel models.
[0105] Among them, the 1x1 convolution kernel enables the model to capture local cross-channel interactions and shares similarities with channel convolution, while the 3x3 convolution branch captures multi-scale feature representations;
[0106] Feature map X∈R C×H×W The feature map is divided into G sub-features through the group branch, where the group style is:
[0107] X=[X0,X1,......,X G-1 ],X i ∈R C / G×H×W ;
[0108] In the formula, R C×H×W is the feature dimension of the original input feature map, C is the number of input channels, H is the height of the input image, and W is the dimension of the input feature;
[0109] At the same time, two parallel convolution branches are used, namely 1x1 convolution branch and 3x3 convolution branch;
[0110] The 1x1 convolution branch is mainly used to capture the dependencies between channels and generate channels through global average pooling.
[0111] The 3x3 convolution branch is responsible for extracting multi-scale spatial information, which can better capture local features and contextual information.
[0112] After that, one-dimensional global average pooling is performed through parallel 1x1 and 3x3 convolution branches;
[0113] The two branches of the 1x1 convolution perform average pooling along the horizontal and vertical dimensions respectively, thereby retaining the precise position information in different dimensional directions;
[0114] Among them, the global average pooling of the coded information along the horizontal dimension in the channel C at height H can be expressed as:
[0115]
[0116] Where, X c is the input feature of the cth channel, and W is the dimension of the input feature;
[0117] After that, the 2D global average pooling encodes the global spatial information, and the different branch outputs are converted to the corresponding dimensions before the joint activation mechanism of the channel features, i.e. and is the feature dimension of the feature map of the average pooling of the 1x1 convolution branch, The input feature dimension of the feature grouping 3x3 convolution branch feature map, The feature dimension of the feature map of the feature grouping 3x3 convolution branch average pooling, The input feature dimension of the feature grouping 1x1 convolution branch feature map, the two-dimensional global pooling operation can be expressed as:
[0118]
[0119] The two-dimensional encoding vector output by the above operation is multiplied by the matrix dot product operation to obtain the spatial attention feature map.
[0120] Step S22, filtering out inaccurate pixels in the pseudo depth through the consistency output strategy DSR module in the self-supervised monocular depth estimation model.
[0121] In the embodiment, for the current frame I t and adjacent frame I t-1 or I t+1 The depth map D of the current frame is obtained through the image depth prior model deep network and posture network t and camera pose transformation P t+1→t ;
[0122] After this, the current frame I is reconstructed t+1→t The cost function is constructed by calculating the minimum photometric error with the target frame to filter out inaccurate pixels in the pseudo depth. The consistent output cost function is as follows:
[0123] D mask =min t+1→t pe(I t ,I t+1→t )≤σ g ;
[0124] In the formula, σ g is a predefined threshold for the reprojection loss.
[0125] Step S23, the predicted edge is refined by the DE edge refinement module in the self-supervised monocular depth estimation model, so as to predict an accurate and clear depth map.
[0126] In the embodiment, the depth network and the posture network are trained in a standard scenario, so the predicted pseudo depth has good smoothness. Then, based on the surface normal extracted from the predicted depth and pseudo depth, a DE edge refinement module is proposed to refine the predicted edge, that is, the surface normal vector is sampled for the pseudo depth and the depth edge predicted by the network model. At the same time, the relative normal angle loss is applied to constrain the depth estimation of the target boundary area. The normal matching loss is:
[0127]
[0128] Where A is the surface normal of the predicted depth of complex samples by the self-supervised monocular depth estimation model, B is the surface normal of the pseudo depth, and N is the total number of pixels;
[0129] By edge-guided sampling, the sampling pair is obtained<A,B> , thus obtaining the edge-aware normal loss, namely:
[0130]
[0131] In the formula, and are the normal vectors of the sampling points from the predicted depth, and These are the normal vectors of the sampling points of pseudo depth.
[0132] Step S3: training and verifying a self-supervised monocular depth estimation model based on the preprocessed data set.
[0133] Step S4, using the trained self-supervised monocular depth estimation model to test the test set and output an accurate and clear depth map.
[0134] The present invention evaluates TalentDepth (a self-supervised monocular depth estimation model) on the NuScenece dataset and compares it with several representative methods. Figure 3As shown in the figure, the absolute relative difference (absRel), root mean square error (RMSE) and accuracy δ1, it can be seen that TalentDepth has achieved good results in terms of error and accuracy under different depths of field; however, its performance is poor at night. This is due to the strong noise level and reflection of night scenes in TalentDepth training. Compared with rainy scenes, night scenes cannot obtain contextual clues in feature extraction, so the effect of rainy scenes is significantly better than that of night. However, compared with other existing methods, TalentDepth's δ1 in night scenes is still improved by more than 10%. This is mainly due to the extraction of more robust features in complex samples.
[0135] This paper evaluates the TalentDepth model on the KITTI dataset and the KITTI-C dataset, and also compares it with representative methods. Figure 4 and Figure 5 As shown, the absolute relative difference (absRel), square relative difference (sqRel), root mean square error (RMSE) and its logarithmic variation (RMSL), as well as the accuracy δ1, δ2, δ3. Since the KITTI-C dataset is more complex than the KITTI dataset, the depth estimation results of most methods in the KITTI-C dataset are relatively poor. However, TalentDepth still obtains better results than existing methods, and even outperforms the results of other methods in standard scenarios. Although the main contribution of TalentDepth is to improve the effect of self-supervised monocular depth estimation in complex weather scenes, experimental results show that it also shows excellent results in standard scenes.
[0136] In order to verify the performance of each module in TalentDepth in depth estimation, this paper conducts ablation experiments on the NuScnece dataset, mainly targeting the attention mechanism, consistent output, and edge refinement to set three ablation methods: ① Whether to embed multi-scale attention on the encoder side; ② Whether to perform pseudo-depth consistency output; ③ Whether to refine the edge.
[0137] Evaluation results such as Figure 6 As shown in Figure 2, the performance without attention mechanism, edge refinement and consistency output is the lowest. The detailed results are analyzed as follows:
[0138] (1) Multi-scale Attention Mechanism Embedding Analysis
[0139] In the encoder network of depth estimation, this paper designs a fusion multi-scale attention mechanism to increase the ability of feature extraction. In order to verify the effect of the method proposed in this paper in the encoder network of this paper, a comparative experiment is carried out with ResNet as the baseline of the encoder. Figure 6It shows that TalentDepth's fusion multi-scale attention mechanism method reduces the error of pixel depth information and makes the predicted depth information more accurate. Specifically, the error absRel, RMSE and RMSL decreased by 0.3%, 17% and 4.7% respectively, and the accuracy δ1, δ2 and δ3 increased by 0.2%, 0.3% and 0.4% respectively.
[0140] (2)DSR module analysis
[0141] This paper outputs the single image depth prior consistently to ensure the reliability of depth clues. Therefore, ablation experiments are also conducted on the DSR module to explore the impact of pseudo-depth consistency output on accuracy. Figure 6 It shows that the errors absRel, sqRel, RMSE and RMSL decreased by 0.2%, 1.2%, 13.9% and 0.6% respectively, and the accuracy δ1, δ2 and δ3 increased by 0.5%, 0.3% and 0.5% respectively.
[0142] (3) DE edge refinement module analysis
[0143] In the process of resolving blurred object boundaries with deep edge refinement, ablation experiments were conducted on the DE edge refinement module of scene objects. At the same time, a multi-scale attention mechanism was integrated as a benchmark for comparative experiments. The results show that the addition of the DE module helps to restore higher quality images. The error sqRel, RMSE and RMSL indicators decreased by 0.2%, 2.9% and 1.5% respectively, and the accuracy δ2 and δ3 increased by 0.4% and 0.2% respectively.
[0144] In summary, the present invention proposes TalentDepth, a self-supervised monocular depth estimation method based on a multi-scale attention mechanism for complex weather scenes. A multi-scale attention mechanism is introduced in the encoder to obtain more robust features; at the same time, a DSR module and a DE module are designed to obtain more detailed depth information. Experimental results on the Nuscence, KITTI, and KITTI-C datasets show that TalentDepth has a certain improvement over other methods, and the advanced performance of an accuracy of 98.2% is achieved on the KITTC-C dataset. In addition, ablation experiments also verify the effectiveness and rationality of the model.
[0145] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments may still be modified, or some or all of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A monocular depth estimation method for complex weather scene images based on self-supervised learning, characterized in that: include: Step S1, obtaining a data set, dividing it into a training set, a validation set, and a test set in proportion, and preprocessing the data set; Step S2, constructing a monocular depth estimation model based on multi-scale attention self-supervision; Step S3, training and validating a self-supervised monocular depth estimation model based on the preprocessed data set; Step S4, using the trained self-supervised monocular depth estimation model to test the test set and output an accurate and clear depth map.
2. The monocular depth estimation method for complex weather scene images based on self-supervised learning according to claim 1 is characterized in that: In step S1, the data set includes: KITTI: is a widely used outdoor depth estimation benchmark that contains continuous stereo images and sparse points collected by sensors mounted on vehicles, including 19,905 training images and 2,212 validation images; KITTI-C: is a comprehensive benchmark for evaluating the robustness of monocular depth estimation. This benchmark shares the same original images as the KITTI test set, simulating a variety of challenging scenarios, including lighting conditions, sensor failure or movement, and noise in data processing. The benchmark contains 18 perturbations. NuScnece: is a challenging large-scale dataset containing 1,000 different scenes, covering different weather scenarios, including but not limited to daytime, rainy days, and nights; following the official split of R4Dyn, there are 15,129 training images and 6,019 validation images.
3. The monocular depth estimation method for complex weather scene images based on self-supervised learning according to claim 1 is characterized in that: In the step S1, the preprocessing includes performing data enhancement and normalization processing on the data set, wherein the data enhancement method includes increasing brightness, reducing brightness and adding Gaussian noise.
4. The monocular depth estimation method for complex weather scene images based on self-supervised learning according to claim 1, characterized in that: In step S2, based on the multi-scale attention self-supervised monocular depth estimation model, a monocular video training deep network and a posture network are used to train a single frame I t The purpose of monocular depth estimation is to predict its corresponding depth map D t , in self-supervised training, model supervision comes from adjacent frames I t ∈{I t-1 ,I t+1 }, the self-supervised monocular depth estimation model uses the target depth estimation D t =DepthNet(I t ) predicts the depth of the current frame and uses the camera pose to estimate P t→t′ =PoseNet(I t ,I′ t ) estimates the camera’s ego-motion to the next timestamp, and then projects it onto the current timestamp to generate the corresponding synthetic object: I t′→t =I t′ <proj(D t ,P t→t′ ,K)>; Where K is the essence of the camera, proj(·) is the table reprojected to I ′t The camera returns to D t 2D coordinates, and (·) is the pixel sampling operator, I ′t is the adjacent frame, I t′ is the target frame; If the prediction is correct, then I t′ →t should be equal to I t Same, therefore, using the minimum luminosity loss L p To constrain D t and P t→t′ The generated I a and I b The twisted flow between: L p =my pe(I t ,IN t′→t ); In the formula, I a is the target frame, I b is the reconstructed target frame, pe(·) is SSIM and L p The former is used to calculate the structural similarity, and the latter is used to calculate the photometric loss. θ is the weight parameter, and SSIM(·) is the image similarity index. In addition, the predicted depth map is regularized using edge-aware smoothness loss: Where ω(D) is the normalized inverse depth of D, ξ x and y are the horizontal and vertical gradients respectively, D is the depth map, and I is a given single frame.
5. The monocular depth estimation method for complex weather scene images based on self-supervised learning according to claim 4 is characterized in that: In the step S2, based on the multi-scale attention self-supervised monocular depth estimation model training, complex samples and corresponding simple samples are required, and the disturbance space is constructed from weak to strong by setting critical values; The model trains an image translation model Translation, for images I in standard scenes t ,Use the image translation model to obtain complex samples: T=Translation(I t ); Where Translation(·) includes different weather or lighting conditions, sensor failure or motion, and noise in the data processing process; After that, the deep network and posture network are trained. The input data of the deep network is not just standard scene samples, so the data is normalized. Since the enhancement operations used by the self-supervised monocular depth estimation model do not change the underlying depth map, the fully trained network is used to obtain the depth map on the original image, and then the depth map is used as the label in the supervised regression loss.
6. The method for monocular depth estimation of complex weather scene images based on self-supervised learning according to claim 5, characterized in that: In step S2, the multi-scale attention self-supervised monocular depth estimation model is divided into three parts, namely feature extraction, depth prior and baseline refinement, including: Step S21, embedding the multi-scale attention module into the self-supervised monocular depth estimation model to perform feature extraction; Step S22, filtering out inaccurate pixels in the pseudo depth through the consistency output strategy DSR module in the self-supervised monocular depth estimation model; Step S23, the predicted edge is refined by the DE edge refinement module in the self-supervised monocular depth estimation model, so as to predict an accurate and clear depth map.
7. The method for monocular depth estimation of complex weather scene images based on self-supervised learning according to claim 6, characterized in that: In step S21, the self-supervised monocular depth estimation model is a Unet structure, which is divided into an encoder, a skip connection and a decoder; Based on ResNet18, a multi-scale attention mechanism is introduced as the encoder; After extracting the feature encoder output, the multi-scale attention module is used to enhance the learning of the scene's global and local features; The decoder makes a jump connection with the feature part of the encoder through multi-scale depth estimation, concatenates features of the same resolution, and upsamples them step by step to restore the depth map; The multi-scale attention module contains two convolution kernels, which are placed in parallel models. Among them, the 1x1 convolution kernel enables the model to capture local cross-channel interactions and shares similarities with channel convolution, while the 3x3 convolution branch captures multi-scale feature representations; Feature map X∈R C×H×W The feature map is divided into G sub-features through the group branch, where the group style is: X=[X0,X1,......,X G-1 ],X i ∈R C / G×H×W ; In the formula, R C×H×W is the feature dimension of the original input feature map, C is the number of input channels, H is the height of the input image, and W is the dimension of the input feature; At the same time, two parallel convolution branches are used, namely 1x1 convolution branch and 3x3 convolution branch; The 1x1 convolution branch is mainly used to capture the dependencies between channels and generate channels through global average pooling. The 3x3 convolution branch is responsible for extracting multi-scale spatial information, which can better capture local features and contextual information. After that, one-dimensional global average pooling is performed through parallel 1x1 and 3x3 convolution branches; The two branches of the 1x1 convolution perform average pooling along the horizontal and vertical dimensions respectively, thereby retaining the precise position information in different dimensional directions; Among them, the global average pooling of the coded information along the horizontal dimension in the channel C at height H can be expressed as: Where, X c is the input feature of the cth channel, and W is the dimension of the input feature; After that, the 2D global average pooling encodes the global spatial information, and the different branch outputs are converted to the corresponding dimensions before the joint activation mechanism of the channel features, i.e. and is the feature dimension of the feature map of the average pooling of the 1x1 convolution branch, The input feature dimension of the feature grouping 3x3 convolution branch feature map, The feature dimension of the feature map of the feature grouping 3x3 convolution branch average pooling, The input feature dimension of the feature grouping 1x1 convolution branch feature map, the two-dimensional global pooling operation can be expressed as: The two-dimensional encoding vector output by the above operation is multiplied by the matrix dot product operation to obtain the spatial attention feature map.
8. The method for monocular depth estimation of complex weather scene images based on self-supervised learning according to claim 7, characterized in that: In step S22, for the current frame I t and adjacent frame I t-1 or I t+1 The depth map D of the current frame is obtained through the image depth prior model deep network and posture network t and camera pose transformation P t+1→t ; Afterwards, reconstruct the current frame I t+1→t The cost function is constructed by calculating the minimum photometric error with the target frame to filter out inaccurate pixels in the pseudo depth. The consistent output cost function is as follows: D mask =my t+1→t pe(I t ,IN t+1→t )≤σ g ; In the formula, σ g is a predefined threshold for the reprojection loss.
9. The monocular depth estimation method for complex weather scene images based on self-supervised learning according to claim 8, characterized in that: In step S23, the depth network and the posture network are both trained under standard scenarios, so the predicted pseudo-depth has good smoothness. Based on the surface normal extracted from the predicted depth and pseudo-depth, a DE edge refinement module is proposed to refine the predicted edge, that is, by sampling the surface normal vector of the pseudo-depth and the depth edge predicted by the network model, and at the same time, applying the relative normal angle loss to constrain the depth estimation of the target boundary area. The normal matching loss is: Where A is the surface normal of the predicted depth of complex samples by the self-supervised monocular depth estimation model, B is the surface normal of the pseudo depth, and N is the total number of pixels; By edge-guided sampling, the sampling pair is obtained<A,B> , thus obtaining the edge-aware normal loss, namely: In the formula, and are the normal vectors of the sampling points from the predicted depth, and These are the normal vectors of the sampling points of pseudo depth.
Citation Information
Cited By
Vehicle front target monocular distance measurement method based on EdgeCAM-Depth network
CN121170739A