Long-distance monocular depth estimation method for various illumination driving scenes

By adopting multi-scale decomposition, Laplace's high-frequency information enhancement and gradual introduction of residual diffusion models in monocular depth estimation, the problem of insufficient long-distance target depth estimation accuracy in different lighting environments is solved, and higher depth estimation accuracy and robustness are achieved.

CN119991762APending Publication Date: 2025-05-13ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510075384.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has insufficient depth estimation accuracy for long-distance targets under different lighting environments, especially in low lighting environments, and it is difficult for the model to accurately predict depth.

Method used

Using a multi-scale decomposition method based on Gaussian smoothing and downsampling, key details in the image are captured through the Laplace high-frequency information enhancement module, and depth features are generated by combining a potential encoder and a diffusion model that introduces residuals.

Benefits of technology

The model's depth estimation accuracy in different lighting environments is significantly improved, especially in the depth measurement of long-distance targets, and the model's robustness and detection accuracy are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991762A_ABST
    Figure CN119991762A_ABST
Patent Text Reader

Abstract

The invention provides a long-distance monocular depth estimation method for various illumination driving scenes, and the method comprises the steps: carrying out the multi-scale decomposition of an input driving scene image based on Gaussian smoothing and downsampling, and obtaining a multi-layer Gaussian image; performing high-frequency information enhancement on the multi-layer Laplacian image to obtain a multi-layer Laplacian image, and performing image reconstruction on the multi-layer Laplacian image to obtain a reconstructed high-frequency information enhancement image; a potential encoder is adopted to encode the reconstructed high-frequency information enhancement graph, and low-dimensional potential space representation is obtained; taking the low-dimensional potential space representation as input, and generating depth features by adopting a diffusion model which gradually introduces residual errors; and decoding the depth features by using a potential decoder to obtain a final depth map. According to the invention, the distance of a distant object can be measured more accurately, and accurate distance measurement is provided especially in different illumination environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing technology, and in particular to a monocular depth estimation method. Background Art

[0002] In the study of monocular depth estimation tasks, especially in autonomous driving scenarios, this task can measure the distance between the autonomous driving vehicle and the object in front to ensure the safe driving of the vehicle on the road. Traditional distance measurement methods rely on sensors such as lidar, but this method has limitations in terms of cost, power consumption and system complexity. Therefore, distance measurement technology based on computer vision and deep learning technology has gradually become a method that has attracted much attention. Using only a single color image, the depth of the scene can be predicted and the distance of the object can be measured. It has the advantages of simple deployment and strong adaptability, which provide new possibilities for the development of depth estimation tasks.

[0003] In an autonomous driving environment, illumination changes and accurate depth prediction of distant targets are two key challenges in object distance measurement. First, complex lighting conditions, such as strong lighting, shadowed environments, and rainy days, can significantly affect the stability and accuracy of depth estimation, especially in low-light environments such as rainy or cloudy days. The contrast and detail information in the image are significantly reduced, making it difficult for the model to accurately predict the depth. Second, due to the small proportion of the target in the image and the low information density, it is difficult for the model to effectively extract the depth features of distant targets, making it difficult to guarantee the accuracy of depth estimation of distant targets.

[0004] The invention patent with publication number CN112785636 A discloses a multi-scale enhanced monocular depth estimation method, including the following steps: input a single RGB image, and then use the context and receptive field enhanced high-resolution network CRE-HRNet to perform multi-scale feature extraction on the RGB image to obtain a high-resolution first image; use the residual dilation convolution unit of the receptive field enhancement module to dilate the first depth image to obtain a second image; use the weighted non-local neighborhood module to capture the long-distance pixels of the second depth image to obtain a depth image. The method of this invention can make its monocular depth estimation accuracy high on the basis of obtaining the feature information of the intermediate layer. However, the measurement accuracy of this method is insufficient under different lighting environments. Summary of the invention

[0005] In view of the technical problem that the prior art has insufficient depth estimation accuracy of distant targets under different lighting driving scenarios, the present invention proposes a long-distance monocular depth estimation method for multiple lighting driving scenarios. The present invention can measure the distance of distant objects more accurately, especially provide accurate distance measurement under different lighting environments.

[0006] In order to achieve the above object, the technical solution of the present invention is achieved as follows:

[0007] A long-distance monocular depth estimation method for driving scenes with multiple lighting conditions comprises the following steps:

[0008] S1: Perform multi-scale decomposition on the input driving scene image based on Gaussian smoothing and downsampling to obtain a multi-layer Gaussian image;

[0009] S2: Perform high-frequency information enhancement on the multi-layer Gaussian image to obtain a multi-layer Laplacian image, and perform image reconstruction on the multi-layer Laplacian image to obtain a reconstructed high-frequency information enhancement image;

[0010] S3: Use a latent encoder to encode the reconstructed high-frequency information enhancement map to obtain a low-dimensional latent space representation;

[0011] S4: Taking the low-dimensional latent space representation as input, a diffusion model with gradually introduced residuals is used to generate deep features;

[0012] S5: Use the potential decoder to decode the depth features and obtain the final depth map.

[0013] Furthermore, the method for performing multi-scale decomposition of the input driving scene image based on Gaussian smoothing and downsampling is as follows: recursively performing Gaussian smoothing and downsampling operations to generate multiple layers of Gaussian images layer by layer, wherein at each layer, Gaussian smoothing is performed on the input image to obtain a blurred image B k :

[0014] B k =G(σ)*G (k) ;

[0015] Where * represents the convolution operation, k is the number of layers, σ is the standard deviation of the Gaussian filter, and G(·) represents the Gaussian function;

[0016] For the blurred image B l Downsample to get the Gaussian image G of the next layer (k+1) :

[0017] G (k+1) = downsample(B k );

[0018] Among them, downsample represents the downsampling operation.

[0019] Furthermore, the method for enhancing high-frequency information of a multi-layer Gaussian image is as follows: starting from the highest layer, processing downward layer by layer, and performing high-frequency information enhancement on the Gaussian image G of the k+1th layer. k+1 After upsampling and adaptive blurring operations, the Gaussian image G of the k+1th layer is transformed intok+1 The upsampling and adaptive blurring results are compared with the k-th layer Gaussian image G k Subtract and get the Laplacian image L of the kth layer k , the Gaussian image of the highest level is directly used as the Laplacian image of the highest level:

[0020] L k =G k -AGB(Up(G (k+1) )),k=1,2,k,…K;

[0021] Among them, k represents the number of layers of the Laplace pyramid, K represents the highest layer, G k represents the Gaussian image of the kth layer, Up(·) represents the upsampling operation, AGB(·) represents the adaptive blur operation, and L k represents the Laplacian image of the kth layer.

[0022] Further, the method of the adaptive fuzzy operation is:

[0023] Compute the Gaussian function at each position in the image:

[0024]

[0025] Where (x, y) is the current position, G(x) and G(y) are the one-dimensional Gaussian functions of x and y respectively, and σ is the standard deviation;

[0026] Convert a one-dimensional Gaussian function to a two-dimensional Gaussian function by using the outer product:

[0027] G2(x,y)=G(x)·G(y);

[0028] Among them, G2(x,y) is a two-dimensional Gaussian function;

[0029] Convolve the image with a two-dimensional Gaussian function:

[0030]

[0031] Among them, I(x,y) is the original image pixel, i is the offset, and I ′ (x, y) is the blurred image, G2(i, j) is the two-dimensional Gaussian function of the offset, and pad is the padding size.

[0032] Furthermore, the method for reconstructing a multi-layer Laplacian image is as follows: starting from the highest layer, processing downward layer by layer, the Kth layer Laplacian image L K After bilinear interpolation upsampling and multi-scale high-pass filtering, the filtering result is consistent with the K-1 layer Laplacian image L K-1Add together to get the reconstructed high-frequency information enhancement map S of the K-1th layer K-1 ; Then, the reconstructed high-frequency information enhancement map S of the K-1th layer K-1 After bilinear interpolation upsampling and multi-scale high-pass filtering, the filtering result is consistent with the K-2 layer Laplacian image L K-2 Add together to get the reconstructed high-frequency information enhancement map S of the K-2th layer K-2 , processing layer by layer until the first layer, the final reconstructed high-frequency information enhancement Figure S1 is obtained:

[0033] S k =L k +H(Up(L k+1 )),k=1,2,3,K;

[0034] Among them, k represents the number of layers of the Laplace pyramid, L k represents the Laplacian image of the kth layer, Up(·) represents the upsampling operation, H(·) represents the multi-scale high-pass filter, S k Represents the reconstructed high-frequency information enhancement map of the kth layer.

[0035] Furthermore, the method of encoding the reconstructed high-frequency information enhancement image using a potential encoder is:

[0036] The reconstructed high-frequency information enhancement map S1 is mapped to the latent space z, and the mapping z in the latent space is convolved and split into mean mean and logarithmic variance logvar. The mean mean is scaled to obtain the low-dimensional latent space representation z r , the calculation formula is:

[0037] mean,logvar=chunk(Conv(E(S1)),2);

[0038] Where E(·) is the mapping function of the latent space z, Conv(·) is the convolution function, and chunk(·) is the splitting function;

[0039] z r =mean×scale_factor;

[0040] Among them, scale_factor is the scaling factor.

[0041] Furthermore, the method of generating deep features by using a diffusion model with gradually introduced residuals is: z is represented by a low-dimensional latent space r As input, the low-dimensional latent space represents z r It is spliced ​​with the output of the previous time step, and the noise residual of the splicing result is predicted using the pre-trained Unet neural network. The deep features of the current time step are obtained based on the predicted noise residual. According to the depth feature of the current time step Perform step-by-step residual introduction calculations to obtain the final deep features

[0042] Furthermore, the calculation of the stepwise residual introduction is performed to obtain the final depth feature The method is: take the depth feature z of the current time step df As input, the residual r of the previous time step is t-1 As the prior residual of the current time step The prior residual for the current time step and the depth feature of the current time step Perform weighted averaging to obtain the smoothed residual of the current time step Use the deep features of the current time step Subtract the smoothed residual at the current time step Get the residual r of the current time step t ; Calculate the weight factor α for the current time step t , based on the residual r of the current time step t and the weight factor α of the current time step t Update the depth feature and judge the current number of steps. If the current number of steps is less than the total number of steps, the diffusion process is not over and the next cycle continues; if the current number of steps is equal to the total number of steps, the diffusion process ends and the final depth feature is output.

[0043] Further, the depth feature of the current time step is obtained The calculation formula is:

[0044]

[0045] Among them, ∈ t is the predicted noise residual, t is the current time step, represents the potential representation of the depth map at the current time step, and scheduler(·) represents the operation of simulating the reverse of the diffusion process.

[0046] The weight factor α for calculating the current time step is t The formula is:

[0047] α t =compute(t,T,r t );

[0048] Where t is the current step number, T is the total number of steps; compute(·) represents the calculation formula for the weight of residual adjustment in the generated diffusion model;

[0049] The residual r based on the current time stept and the weight factor α of the current time step t The method to update the deep features is:

[0050] Determine the size relationship between the current number of steps and half of the total number of steps. If the current number of steps is less than half of the total number of steps, the residual introduction is strengthened by updating the formula; if the current number of steps is greater than or equal to half of the total number of steps, the residual introduction is gradually weakened by updating the formula;

[0051] The update formula is:

[0052]

[0053] in, is the output of the current time step.

[0054] Beneficial effects of the present invention:

[0055] The present invention adopts the Laplace high-frequency information enhancement module, and the model can better capture the key details in the image, especially the edge and texture features, and improve the feature extraction ability of the model; the present invention adopts a diffusion model that gradually introduces residuals, and incrementally introduces high-frequency and edge information in the depth feature at each step of the diffusion, preventing feature distortion caused by adding a large amount of noise at one time, generating a clearer depth map, and obtaining a more accurate distance, thereby enhancing the robustness of the model and optimizing the detection accuracy. Compared with other monocular depth estimation methods, the present invention can more accurately measure the distance of distant objects, especially providing accurate distance measurement under different lighting environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0057] Figure 1 The figure is a flow chart of the method of the present invention.

[0058] Figure 2 It is the overall structure diagram of the present invention.

[0059] Figure 3 This is a structural diagram of the high-frequency information enhancement module of the present invention.

[0060] Figure 4 This is a module structure diagram for gradually introducing residuals according to the present invention.

[0061] Figure 5 and Figure 6This is a comparison chart of the prediction results of the monocular depth estimation method based on the KITTI dataset compared with other methods of the present invention; wherein, the first row of pictures are original images, the second row of pictures are depth maps predicted by the Monodepth2 method, the third row of pictures are depth maps predicted by the Lite-mono method, the fourth row of pictures are depth maps predicted by the Manydepth method, the fifth row of pictures are depth maps predicted by the Marigold method, and the sixth row of pictures are depth maps predicted by the method of the present invention.

[0062] Figure 7 This is a comparison chart of the prediction results of the monocular depth estimation method based on the DIODE dataset compared with other methods. The first row is the original image, the second row is the depth map predicted by the Marigold method, and the third row is the depth map predicted by the method of the present invention.

[0063] Figure 8 This is a comparison chart of the method of the present invention based on the Cityscapes dataset and the ground truth, wherein the first row of pictures uses the depth map predicted by the method of the present invention, and the second row of pictures is the benchmark true value.

[0064] Fig. 9 This is a comparison chart between the method of the present invention based on the nuScenes dataset and the ground truth, wherein the first row of pictures are depth maps predicted by the method of the present invention, and the second row of pictures are benchmark true values. DETAILED DESCRIPTION

[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] A long-range monocular depth estimation method for driving scenarios with various lighting conditions, such as Figure 1 As shown, the steps include:

[0067] S1: Based on Gaussian smoothing and downsampling, the input driving scene image is decomposed into multiple scales to obtain a multi-layer Gaussian image.

[0068] Gaussian smoothing and downsampling are used for recursive operations to generate multi-layer Gaussian images layer by layer. In each layer, first, the input image is Gaussian smoothed to obtain a blurred image B. l :

[0069] B k =G(σ)*G k ;

[0070] Where * represents the convolution operation, σ is the standard deviation of the Gaussian filter, and G(·) represents the Gaussian function.

[0071] Furthermore, for the blurred image B l Downsample to get the image G of the next layer k+1 :

[0072] G k+1 = downsample(B k );

[0073] Here, downsample represents a downsampling operation, which here means reducing the size of the image by half.

[0074] S2: Perform high-frequency information enhancement on the multi-layer Gaussian image to obtain a multi-layer Laplace image (i.e., a high-frequency information enhanced image), and perform image reconstruction on the multi-layer Laplace image to obtain a reconstructed high-frequency information enhanced image.

[0075] Furthermore, the method for enhancing high-frequency information of a multi-layer Gaussian image is as follows: starting from the highest layer (lowest scale), processing downward layer by layer, and performing high-frequency information enhancement on the Gaussian image G of the k+1th layer. k+1 After upsampling and adaptive blurring operations, the Gaussian image G of the k+1th layer is transformed into k+1 The upsampling and adaptive blurring results are compared with the k-th layer Gaussian image G k Subtract and get the Laplacian image L of the kth layer k , the Gaussian image of the highest layer is directly used as the Laplacian image of the highest layer.

[0076] Adaptive blur is applied during upsampling. This adaptive blur operation is to adjust the degree of high-frequency information retention at different levels to ensure that the details at each scale are properly captured and emphasized. The present invention uses a Laplacian pyramid structure to achieve high-frequency information enhancement. The Laplacian pyramid can effectively separate high-frequency information (such as edges, detailed textures, etc.) at each scale, and this information is retained in the corresponding Laplacian image. Such processing can significantly improve the details of the image, allowing the model to more clearly identify the edges and details of distant objects, which is very beneficial for subsequent depth estimation tasks.

[0077] like Figure 2 As shown:

[0078] The Gaussian image G2 of the second layer is upsampled, and an adaptive blur operation is used in the process. The result is subtracted from the first layer image G1 to obtain the Laplacian image L1.

[0079] The Gaussian image G3 of the third layer is upsampled, and an adaptive blur operation is used in the process. The result is subtracted from the second layer image G2 to obtain the Laplacian image L2.

[0080] The Gaussian image G4 of the fourth layer is upsampled, and an adaptive blur operation is used in the process. The result is subtracted from the third layer image G3 to obtain the Laplacian image L3.

[0081] The Gaussian image G4 of the fourth layer is directly used as the Laplacian image L4.

[0082] Further, the calculation expression is:

[0083] L k =G k -AGB(Up(G (k+1) )),k=1,2,k,…K;

[0084] Among them, k represents the number of layers of the Laplace pyramid, K represents the highest layer, G k represents the Gaussian image of the kth layer, Up(·) represents the upsampling operation, AGB(·) represents the adaptive blur operation, and L k represents the Laplacian image of the kth layer.

[0085] Further, the process of adaptive fuzzy operation is as follows:

[0086] First, calculate the Gaussian function at each position of the image:

[0087]

[0088] Where (x, y) is the current position, G(x) and G(y) are the one-dimensional Gaussian functions of x and y respectively, indicating the Gaussian weight of each position, and σ is the standard deviation, which controls the degree of blur. Here, σ is an adjustable parameter. When it is smaller, the Gaussian function will be more concentrated and the blurring effect will be weaker; when it is larger, the influence range of the Gaussian function will be expanded and the blurring effect will be stronger.

[0089] Furthermore, the one-dimensional Gaussian function is converted to a two-dimensional Gaussian function by using the outer product:

[0090] G2(x,y)=G(x)·G(y);

[0091] Here G2(x,y) is a two-dimensional Gaussian function. Finally, the above two-dimensional Gaussian function is convolved with the image to achieve adaptive blur operation. The convolution process is as follows:

[0092]

[0093] Here, I(x,y) is the original image pixel, i is the offset, and I′ (x, y) is the blurred image, G2(i, j) is the two-dimensional Gaussian function of the offset, and pad is the padding size to ensure that the image boundary information is not lost during convolution.

[0094] Furthermore, the method for reconstructing a multi-layer Laplacian image is as follows: starting from the highest layer (lowest scale), processing downward layer by layer, the Kth layer Laplacian image L K After bilinear interpolation upsampling and multi-scale high-pass filtering, the filtering result is consistent with the K-1 layer Laplacian image L K-1 Add together to get the reconstructed high-frequency information enhancement map S of the K-1th layer K-1 ; Then, the reconstructed high-frequency information enhancement map S of the K-1th layer K-1 After bilinear interpolation upsampling and multi-scale high-pass filtering, the filtering result is consistent with the K-2 layer Laplacian image L K-2 Add together to get the reconstructed high-frequency information enhancement map S of the K-2th layer K-2 , layer by layer until the first layer, and the final reconstructed high-frequency information enhanced image S1 is obtained as the input of the subsequent steps.

[0095] like Figure 2 As shown:

[0096] In the reconstruction process, the Laplacian image L4 is firstly subjected to bilinear interpolation upsampling operation, and then filtered by a multi-scale high-pass filter. The filtering result is added to the Laplacian image L3 to obtain the reconstructed high-frequency information enhanced image S3.

[0097] Then, a bilinear interpolation upsampling operation is performed on the high-frequency information enhancement image S3, and then a multi-scale high-pass filter is performed. The filtering result is added to the Laplacian image L2 to obtain the reconstructed high-frequency information enhancement image S2.

[0098] Finally, a bilinear interpolation upsampling operation is performed on the reconstructed high-frequency information enhancement image S2, and then a multi-scale high-pass filter is performed. The filtering result is added to the Laplacian image L1 to obtain the reconstructed high-frequency information enhancement image S1.

[0099] Further, the calculation expression is:

[0100] S k =L k +H(Up(L k+1 )),k=1,2,3,K;

[0101] Among them, k represents the number of layers of the Laplace pyramid, L k represents the Laplacian image of the kth layer, Up(·) represents the upsampling operation, H(·) represents the multi-scale high-pass filter, S kRepresents the reconstructed high-frequency information enhancement map of the kth layer.

[0102] S3: A latent encoder is used to encode the reconstructed high-frequency information enhancement map to obtain a low-dimensional latent space representation.

[0103] Given an input image, i.e., the reconstructed high-frequency information enhanced map S1, the encoder E maps it to the latent space z:

[0104] z=E(S1);

[0105] Among them, E(·) is the mapping function of the latent space z.

[0106] Further, the mapping z in the latent space is convolved to obtain m:

[0107] m = Conv(z);

[0108] Among them, Conv(·) is the convolution function.

[0109] Further, split m into mean and logarithmic variance logvar:

[0110] mean,logvar=chunk(m,2);

[0111] Among them, chunk(·) is the splitting function.

[0112] Further, the mean is scaled to obtain a low-dimensional latent space representation z r :

[0113] z r =mean×scale_factor;

[0114] Among them, scale_factor is the scaling factor.

[0115] S4: Taking the low-dimensional latent space representation as input, a diffusion model with gradually introduced residuals is adopted to generate deep features.

[0116] Furthermore, the method of generating deep features by gradually introducing a diffusion model with residuals is:

[0117] First, we represent z in a low-dimensional latent space r As input, the depth map is initialized by generating Gaussian noise, that is, a noise tensor is generated to represent the potential representation of the initial depth map

[0118] Furthermore, a zero tensor with the same shape as the initial depth map potential representation is initialized to store the depth features output at the previous time step, and the depth features output at the previous time step are used as the depth map potential representation of the current time step. This tensor is updated at each iteration.

[0119] Furthermore, at each time step t, the low-dimensional latent space representation z is concatenated r and the potential representation of the depth map at the current time step Get the splicing result x of the current time step t :

[0120]

[0121] Further, the concatenation result x of the current time step t Input to the pre-trained UNet neural network for learning noise prediction to predict the potential representation of the depth map at the current time step The noise residual ∈ t :

[0122] ∈ t =U(x t ,t);

[0123] Wherein, U(·) represents the pre-trained Unet neural network function, and Conditioned 2DUnet is used in the present invention.

[0124] Furthermore, using the predicted noise residual ∈ t Update and obtain the deep features of the current time step

[0125]

[0126] Among them, scheduler(·) represents the operation of simulating the reverse of the diffusion process.

[0127] Furthermore, the depth feature of the current time step Input into the stepwise residual introduction module to calculate the stepwise residual introduction. The steps are as follows:

[0128] First, the residual r of the previous time step is t-1 As the prior residual of the current time step

[0129] Furthermore, the prior residual of the current time step and the depth feature of the current time step Perform weighted averaging to obtain the smoothed residual of the current time step This smoothing operation can reduce the sharp fluctuations in the denoising process and make the update of each time step more stable:

[0130]

[0131] Further, using the deep features of the current time step Subtract the smoothed residual at the current time step Get the residual r of the current time step t :

[0132]

[0133] This residual represents the current denoised sample (i.e. the depth feature of the current time step ) and the smoothed denoised sample (i.e., the smoothed residual ) is used to dynamically adjust the deep features during the denoising process.

[0134] Further, calculate the weight factor α of the current time step t :

[0135] α t =compute(t,T,r t );

[0136] Where t is the current step number, T is the total number of steps, and compute(·) represents the calculation formula for the weights of residual adjustment in the generated diffusion model.

[0137] Further, the depth feature z is updated according to the calculated weight factor df , the update formula is:

[0138]

[0139] in, is the output of the current time step, which serves as the potential representation of the depth map of the next time step.

[0140] Furthermore, the relationship between the current denoising step number and half of the total denoising step number is determined. If the current denoising step number is less than half of the total denoising step number, the model is in the first half of the diffusion stage, and the residual introduction is strengthened according to the update formula to effectively recover the depth information from the noise; if the current denoising step number is greater than or equal to half of the total denoising step number, the model is in the second half of the diffusion stage, and the residual introduction is gradually weakened according to the update formula, focusing on detail recovery rather than over-adjustment to avoid excessive disturbance of the details.

[0141] Furthermore, the current number of steps is judged. If the current number of steps is less than the total number of steps, the diffusion process is not over and the next cycle continues; if the current number of steps is equal to the total number of steps, the diffusion process ends and the final depth feature is output.

[0142] S5: Use the potential decoder to decode the deep features to obtain the final depth map.

[0143] Given the deep features encoded in the latent space The decoder D decodes it back to the depth map space and generates a predicted depth map

[0144]

[0145] In order to further test the feasibility and effectiveness of the method of the present invention, experiments were conducted on the method of the present invention.

[0146] The error evaluation method and the accuracy evaluation method are used to evaluate the test results of the method of the present invention and the existing monocular depth estimation method on the data set used by the method of the present invention. The error evaluation methods include average relative error (Abs Rel), square relative error (Sq Rel), root mean square error (RMSE), and RMSE log. The average relative error calculates the average value of the relative error between the predicted depth and the true depth, which is used to measure the relative accuracy of the depth estimation. The smaller the value, the more accurate the prediction; the square relative error calculates the square relative error between the predicted depth and the true depth. It is more sensitive to larger errors and is used to penalize predictions with larger deviations. The smaller the value, the more stable the depth estimation; the root mean square error calculates the root mean square error between the predicted depth and the true depth, reflecting the average level of the overall error. The smaller the value, the smaller the deviation between the predicted value and the true value; RMSE log is the value obtained by taking the logarithm of the root mean square error. The smaller the value, the better the balance of prediction accuracy between long distance and short distance.

[0147] The accuracy evaluation methods include δ1, δ2, and δ3. These indicators represent the ratio of pixels whose predicted values ​​meet different threshold conditions to the true values. They are usually defined as follows: δ1 represents the ratio of pixels whose predicted values ​​meet the true values ​​in the interval [1 / 1.25, 1.25]; δ2 represents the ratio of pixels whose predicted values ​​meet the true values ​​in the interval [1 / 1.25, 1.25]. 2 ,1.25 2 ] interval; δ3 means the ratio of the predicted value to the true value is in [1 / 1.25 3 ,1.25 3 ] The larger the indicators are, the more pixel prediction values ​​are close to the true values, and the higher the overall accuracy of the model.

[0148] Tables 1 and 2 respectively give the evaluation values ​​of the average relative error (Abs Rel), square relative error (Sq Rel), root mean square error (RMSE), RMSE log, δ1, δ2 and δ3 obtained by experiments on the KITTI and DIODE datasets using the method of the present invention and the existing monocular depth estimation method.

[0149] Table 1 Experimental results evaluation values ​​obtained by the present invention on the KITTI dataset

[0150]

[0151] Table 2 Experimental results evaluation values ​​obtained by the present invention on the DIODE dataset

[0152]

[0153] Among them, Monodepth2 comes from the literature [Godard C, Mac Aodha O, Firman M, et al. Digging into self-supervised monocular depth estimation [C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2019: 3828-3838.].

[0154] Lite-mono comes from the literature [Zhang N, Nex F, Vosselman G, et al. Lite-mono: Alightweight cnn and transformer architecture for self-supervised monocular depth estimation [C] / / Proceedings of the IEEE / CVF Conference on ComputerVision and Pattern Recognition.2023:18537-18546.].

[0155] Manydepth comes from the literature [Watson J, Mac Aodha O, Prisacariu V, et al. Thetemporal opportunist: Self-supervised multi-frame monocular depth [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2021: 1164-1174.].

[0156] Marigold comes from the literature [Ke B, Obukhov A, Huang S, et al.Repurposingdiffusion-based image generators for monocular depth estimation[C] / / Proceedings of the IEEE / CVF Conferenc e on Computer Vision and PatternRecognition.2024:9492-9502.].

[0157] It can be seen from the data listed in Table 1 and Table 2 that the Abs Rel, Sq Rel, RMSE, RMSE log, δ1, δ2 and δ3 of the depth image obtained by the method of the present invention are better than those of the compared monocular depth estimation method, which shows that the depth image obtained by the present invention can more accurately predict the depth of the object in the image. The experimental results and data analysis fully demonstrate the advantages of the method of the present invention, effectively improving the accuracy of the monocular depth estimation method for detecting distant objects, and improving the robustness under different lighting environments.

[0158] like Figure 5 and Figure 6 As shown, in the present invention, experiments were conducted using the KITTI dataset. Figure 5 In the present invention, the method used in the present invention can estimate the depth of distant objects more accurately, and can more clearly see the outline of the object, and can better display the depth of distant objects. Figure 5 In the second picture in , Manydepth and Marigold did not estimate the depth of the distant electric pole, while the method used in the present invention clearly estimated the depth of the object. Figure 6 In the present invention, the method used can better estimate the depth of the object in the shadow part, and can better display the depth of the object in the shadow part. For example, in Figure 6 In the first picture in , ManyDepth and Marigold do not accurately estimate the depth of the pedestrian in the shadow, while the method used in the present invention more clearly estimates the depth of the pedestrian.

[0159] like Figure 7 As shown, in the present invention, the DIODE dataset is used for experiments. Compared with Marigold, the method used in the present invention is more accurate in estimating the depth of distant details and can better display the depth of distant objects. In the first and fifth images, the depth of the distant gate is estimated more carefully; in the sixth image, the depth estimation of the distant object is more accurate and is not affected by the light.

[0160] like Figure 8 As shown, in the present invention, the Cityscapes dataset is used for experiments. Compared with real pictures, the method used in the present invention can meet the depth estimation under different lighting conditions, and the depth estimation of objects in distant details and objects in shadows is relatively accurate.

[0161] like Fig. 9 As shown, in the present invention, the nuScenes dataset is used for experiments. The method used in the present invention is experimented in an environment with poor light (rainy day), and compared with real pictures, it satisfies the accurate depth estimation of distant objects.

[0162] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A long-distance monocular depth estimation method for driving scenes with multiple lighting conditions, characterized in that: Includes steps: S1: Perform multi-scale decomposition on the input driving scene image based on Gaussian smoothing and downsampling to obtain a multi-layer Gaussian image; S2: Perform high-frequency information enhancement on the multi-layer Gaussian image to obtain a multi-layer Laplacian image, and perform image reconstruction on the multi-layer Laplacian image to obtain a reconstructed high-frequency information enhancement image; S3: Use a latent encoder to encode the reconstructed high-frequency information enhancement map to obtain a low-dimensional latent space representation; S4: Taking the low-dimensional latent space representation as input, a diffusion model with gradually introduced residuals is used to generate deep features; S5: Use the potential decoder to decode the depth features and obtain the final depth map.

2. The long-distance monocular depth estimation method for multiple lighting driving scenes according to claim 1, characterized in that: The method for multi-scale decomposition of the input driving scene image based on Gaussian smoothing and downsampling is as follows: Gaussian smoothing and downsampling are used to perform recursive operations to generate multi-layer Gaussian images layer by layer, wherein, at each layer, the input image is Gaussian smoothed to obtain a blurred image B k : B k =G(σ)*G (k) ; Where * represents the convolution operation, k is the number of layers, σ is the standard deviation of the Gaussian filter, and G(·) represents the Gaussian function; For the blurred image B l Downsample to get the Gaussian image G of the next layer (k+1) : G (k+1) =downsample(B k ); Among them, downsample represents the downsampling operation.

3. The long-distance monocular depth estimation method for driving scenes with multiple lighting conditions according to claim 2, characterized in that: The method for enhancing high-frequency information of a multi-layer Gaussian image is as follows: starting from the highest layer, processing downward layer by layer, and performing high-frequency information enhancement on the Gaussian image G of the k+1th layer. k+1 After upsampling and adaptive blurring operations, the Gaussian image G of the k+1th layer is transformed into k+1 The upsampling and adaptive blurring results are compared with the k-th layer Gaussian image G k Subtract and get the Laplacian image L of the kth layer k , the Gaussian image of the highest level is directly used as the Laplacian image of the highest level: L k =G k -AGB(Up(G (k+1) )),k=1,2,k,…K; Among them, k represents the number of layers of the Laplace pyramid, K represents the highest layer, G k represents the Gaussian image of the kth layer, Up(·) represents the upsampling operation, AGB(·) represents the adaptive blur operation, and L k represents the Laplacian image of the kth layer.

4. The long-distance monocular depth estimation method for driving scenes with multiple lighting conditions according to claim 3, characterized in that: The method of the adaptive fuzzy operation is: Compute the Gaussian function at each position in the image: Where (x, y) is the current position, G(x) and G(y) are the one-dimensional Gaussian functions of x and y respectively, and σ is the standard deviation; Convert a one-dimensional Gaussian function to a two-dimensional Gaussian function by using the outer product: G2(x,y)=G(x)·G(y); Among them, G2(x,y) is a two-dimensional Gaussian function; Convolve the image with a two-dimensional Gaussian function: Among them, I(x,y) is the original image pixel, i is the offset, and I ′ (x, y) is the blurred image, G2(i, j) is the two-dimensional Gaussian function of the offset, and pad is the padding size.

5. The long-distance monocular depth estimation method for driving scenes with multiple lighting conditions according to claim 4, characterized in that: The method for reconstructing a multi-layer Laplacian image is as follows: starting from the highest layer, processing downward layer by layer, the Kth layer Laplacian image L K After bilinear interpolation upsampling and multi-scale high-pass filtering, the filtering result is consistent with the K-1 layer Laplacian image L K-1 Add together to get the reconstructed high-frequency information enhancement map S of the K-1th layer K-1 ; Then, the reconstructed high-frequency information enhancement map S of the K-1th layer K-1 After bilinear interpolation upsampling and multi-scale high-pass filtering, the filtering result is consistent with the K-2 layer Laplacian image L K-2 Add together to get the reconstructed high-frequency information enhancement map S of the K-2th layer K-2 , processing layer by layer until the first layer, the final reconstructed high-frequency information enhancement Figure S1 is obtained: S k =L k +H(Up(L k+1 )),k=1,2,3,K; Among them, k represents the number of layers of the Laplace pyramid, L k represents the Laplacian image of the kth layer, Up(·) represents the upsampling operation, H(·) represents the multi-scale high-pass filter, S k Represents the reconstructed high-frequency information enhancement map of the kth layer.

6. The long-distance monocular depth estimation method for driving scenes with multiple lighting conditions according to claim 5, characterized in that: The method of encoding the reconstructed high-frequency information enhancement image using a potential encoder is: The reconstructed high-frequency information enhancement map S1 is mapped to the latent space z, and the mapping z in the latent space is convolved and split into mean mean and logarithmic variance logvar. The mean mean is scaled to obtain the low-dimensional latent space representation z r , the calculation formula is: mean,logvar=chunk(Conv(E(S1)),2); Where E(·) is the mapping function of the latent space z, Conv(·) is the convolution function, and chunk(·) is the splitting function; z r =mean×scale_factor; Among them, scale_factor is the scaling factor.

7. The long-distance monocular depth estimation method for multiple lighting driving scenes according to claim 5 or 6, characterized in that: The method of generating deep features by gradually introducing a diffusion model with residuals is as follows: z is represented by a low-dimensional latent space r As input, the low-dimensional latent space represents z r It is spliced ​​with the output of the previous time step, and the noise residual of the splicing result is predicted using the pre-trained Unet neural network. The deep features of the current time step are obtained based on the predicted noise residual. According to the depth feature of the current time step Perform step-by-step residual introduction calculations to obtain the final deep features 8. The long-distance monocular depth estimation method for driving scenes with multiple lighting conditions according to claim 7, characterized in that: The calculation of gradual residual introduction is performed to obtain the final deep features The method is: take the depth feature z of the current time step df As input, the residual r of the previous time step is t-1 As the prior residual of the current time step The prior residual for the current time step and the depth feature of the current time step Perform weighted averaging to obtain the smoothed residual of the current time step Use the deep features of the current time step Subtract the smoothed residual at the current time step Get the residual r of the current time step t ; Calculate the weight factor α for the current time step t , based on the residual r of the current time step t and the weight factor α of the current time step t Update the depth feature and judge the current number of steps. If the current number of steps is less than the total number of steps, the diffusion process is not over and the next cycle continues; if the current number of steps is equal to the total number of steps, the diffusion process ends and the final depth feature is output.

9. The long-distance monocular depth estimation method for driving scenes with multiple lighting conditions according to claim 8, characterized in that: The deep features of the current time step are obtained The calculation formula is: Among them, ∈ t is the predicted noise residual, t is the current time step, represents the potential representation of the depth map at the current time step, scheduler(·) represents the operation that simulates the reverse of the diffusion process.

10. The long-distance monocular depth estimation method for driving scenes with multiple lighting conditions according to claim 8, characterized in that: The weight factor α for calculating the current time step is t The formula is: α t =compute(t,T,r t ); Where t is the current step number, T is the total number of steps; compute(·) represents the calculation formula for the weight of residual adjustment in the generated diffusion model; The residual r based on the current time step t and the weight factor α of the current time step t The method to update the deep features is: Determine the size relationship between the current number of steps and half of the total number of steps. If the current number of steps is less than half of the total number of steps, the residual introduction is strengthened by updating the formula; if the current number of steps is greater than or equal to half of the total number of steps, the residual introduction is gradually weakened by updating the formula; The update formula is: in, is the output of the current time step.

Citation Information

Patent Citations

  • Multi-scale enhanced monocular depth estimation method

    CN112785636A