A Video Dehazing and Depth Estimation Method in Real Moving Scenarios

By combining the atmospheric scattering model and brightness consistency constraint method, the learning framework for video defog removal and depth estimation is optimized, and the brightness inconsistency problem of video defog removal in real mobile scenes is solved, achieving high-quality defog removal and depth estimation effects.

CN119323533BActive Publication Date: 2025-06-10NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411322635.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2025-06-10
Estimated Expiration
2044-09-23

AI Technical Summary

Technical Problem

The existing video defogging technology has brightness inconsistency problems in real mobile scenes, resulting in flickering between adjacent frames and limited effectiveness in harsh outdoor weather conditions.

Method used

A video defog and depth estimation method combining atmospheric scattering model and brightness consistency constraints is proposed. By integrating ASM model and BCC constraints, the learning framework for video defog and depth estimation is optimized, and the discriminator network is used to improve the high-frequency details of the defog results and the accuracy of the depth estimation.

Benefits of technology

It effectively solves the brightness inconsistency of video defog removal in real mobile scenes, improves the accuracy of defog removal quality and depth estimation, especially in severe outdoor weather conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119323533B_ABST
    Figure CN119323533B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for video dehazing and depth estimation in a real moving scenario, providing a clearer vision and distance perception for a real fog scenario. This method combines luminance consistency constraints and an atmospheric scattering model to jointly optimize a depth estimation network; the luminance consistency constraint uses adjacent frames of dehazing as inputs to obtain a more accurate camera pose, uses the estimated depth map to complete reprojection, and constructs a self-supervised estimation method; with the self-supervised learning of the depth estimation network, each component of the entire atmospheric scattering model is effectively decoupled respectively, and at the same time, each component is used to reconstruct the foggy frame to form a reconstruction loss constraint to supervise the learning of the entire framework; the depth information shared by the luminance consistency constraint and the atmospheric scattering model is used to construct a jointly optimized learning framework. In the test usage stage, only one test image of real fog needs to be input, and the dehazing network and the depth estimation network can be used to quickly restore its clear scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video restoration and depth estimation, and particularly to a video dehazing and depth estimation algorithm in a real moving scene. Background Art

[0002] Early image dehazing methods mainly focused on combining the atmospheric scattering model (ASM) with different prior knowledge. However, recent progress in this field has shown that better performance can be achieved using deep learning techniques on large datasets of hazy / clear images. These methods use deep neural networks to learn the physical model parameters or directly map hazy images to clear images / videos. These methods mainly rely on aligned synthetic data for training, which, due to distribution differences, results in suboptimal performance in real-world scenarios. To reduce the domain gap, some methods adopt domain adaptation and unpaired dehazing models to handle real-world scenarios. Despite these efforts, when applying the image dehazing model to videos, the lack of luminance consistency constraints leads to inconsistent luminance between adjacent frames, resulting in flickering. Video dehazing techniques utilize the temporal information of adjacent frames to improve the recovery quality. Early methods mainly focused on post-processing, optimizing the transmission map and suppressing artifacts to ensure temporal consistency. Some methods solve multiple tasks simultaneously, including depth estimation and detection in hazy videos. Recently, the Real-World Video Dehazing Dataset (REVIDE) was introduced, and an improved deformable network based on confidence guidance was proposed. A phase-based memory network was designed for video dehazing based on REVIDE. Similarly, a memory-based physical prior guidance module was proposed to incorporate prior-related features into long-term memory for video dehazing. In addition, some image restoration methods have demonstrated excellent performance on the REVIDE dataset, especially under adverse weather conditions. However, these methods are mainly trained and evaluated in indoor smoke scenarios, which limits their effectiveness in real-world outdoor hazy conditions. Compared with dehazing and depth estimation works, the proposed DCL is trained on real hazy videos rather than synthetically blurred images and simultaneously optimizes video dehazing and depth estimation by integrating the ASM model and the brightness consistency constraint (BCC). Self-supervised monocular depth estimation (SMDE) has limited lidar perception capabilities in extreme weather conditions, leading to a growing interest in self-supervised methods. After the pioneering study by Zhou et al. demonstrated that superior performance could be achieved by exploiting only the geometric constraints between consecutive frames, researchers further explored the cues and methods for training self-supervised models using videos or stereo image pairs. Recently, some studies have utilized advanced networks to estimate depth in challenging scenarios (such as rain, snow, fog, and low-light environments) and explored SMDE. These degraded images, especially in regions with weak texture, significantly affect depth estimation. For weak texture, some methods mainly use image enhancement and domain adaptation learning to address this issue. However, these image enhancement methods rely on operator-based methods rather than learnable methods, which only enhance depth estimation. In addition, most domain adaptation learning based on synthetic data may not be effectively applied to real-world scenarios. Summary of the Invention

[0003] The purpose of this invention is to propose a video defogging and depth estimation method in a real mobile scene. On the one hand, the loss function is designed by using non-aligned clear images. The supervised training of the desmog network, on the other hand, redefines the mean-variance description of atmospheric light and the three-channel method of the transmission medium map, and proposes a corresponding neural network structure to better learn atmospheric light and transmission medium maps, making the reconstruction loss function Efficiently train a dehazing network.

[0004] A technical solution to achieve the purpose of the present invention: a video defogging and depth estimation method for a real moving scene, comprising:

[0005] Step 1: Preprocess the data set; use the non-aligned matching algorithm to obtain the matching information between all foggy video frames and clear reference frames, dedistort the data according to the calibrated camera internal parameters, and perform cropping;

[0006] Step 2: Transform the continuous fog video frames I [t-n:t+n] ∈R 3×h×w Input to the dehazing network Φ J The dehazing result is obtained in [J t ,J s ], J t represents the current defogging frame, J s Indicates J t Neighboring defogging frames, where n = 1; the current foggy video frame I t Input into the depth estimation network Φ respectively d Medium estimate d t ∈R 1×h×w , and the scattering coefficient estimation network Φ β Estimated β t ∈R 1×h×w , using the dark channel method to calculate the value of the infinite atmospheric light A ∞ ;

[0007] Step 3: The defogging result of step 2 t Then input into the camera pose prediction network Φ p To predict the camera’s pose information P x→y ∈R 4×4 , and then use Φ d The estimated d t and Φ p Get P x→y By reprojecting the adjacent dehazed frames onto the current frame, the depth d is estimated by self-supervised learning by calculating the photometric loss constrained depth estimation. t ;

[0008] Step 4: J generated from step 2 to step 3t , A ∞ , d t and β t Calculate and reconstruct the current fog video frame I′ using the atmospheric scattering model t , and then define I′ t and I t reconstruction loss function which consists of the sum of three loss functions, namely the first-order norm loss perceptual loss and structural similarity loss

[0009] Step 5: Construct a non-aligned reference loss function for J t and J′ t to regularize the feature distribution of color and texture of the defogging network; H t and H′ t non-aligned reference loss function to regularize the feature distribution of color and texture of the defogging network;

[0010] Step 6: Use the discriminator network to respectively regularize the defogging network Φ J and the depth estimation network Φ d ;

[0011] Step 7: According to the two loss functions in Steps 4, 5, and 6 and and the adversarial loss regularization of the two discriminator networks D MFIR and D MDR , optimize the network parameters of the entire DCL framework to obtain the defogging result; finally, input the test real RGB fog video frame I t , and input it into the defogging network and the depth estimation network respectively to directly generate the clear scene image result J t and the depth map d t .

[0012] An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned video defogging and depth estimation method in a real mobile scenario.

[0013] A computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the above-mentioned video defogging and depth estimation method in a real mobile scenario.

[0014] A computer program product, including a computer program. When the computer program is executed by a processor, it implements the above-mentioned video defogging and depth estimation method in a real mobile scenario.

[0015] Compared with the prior art, the present invention has the following remarkable advantages: (1) The present invention combines an atmospheric scattering model and a brightness consistency constraint to propose a depth-centered learning framework for real-scene video defogging and depth estimation. (2) The present invention proposes two discriminator networks D MFIR and D MDR to respectively improve the high-frequency details of the defogging result and the "black hole" problem of the depth map generated by weak textures.

[0016] The following further describes the present invention in detail with reference to the accompanying drawings. Description of the Drawings

[0017] Figure 1 It is a model network architecture diagram of the method of the present invention.

[0018] Figure 2 It is a comparison schematic diagram of the method of the present invention, showing that joint video defogging and depth estimation has better effects than separate foggy-day depth estimation and defogging first and then depth estimation.

[0019] Figure 3 It is an effect comparison diagram of the method of the present invention with other methods on three real foggy-day video datasets, namely GoProHazy, DrivingHazy, and InternetHazy, and our method simultaneously gives the effects of defogging and depth estimation.

[0020] Figure 4 It is an effect comparison diagram of the method of the present invention with other methods on three real foggy-day video datasets, namely GoProHazy, DENS-Fog (dense), and DENS-Fog (light), and our method simultaneously gives the effects of defogging and depth estimation.

[0021] Figure 5 It is an ablation experiment effect diagram of the proposed core module of the method of the present invention visualized on the DENS-Fog (light) dataset, and also verifies the effects of D MFIR and D MDR on the defogging result and the estimated depth on the GoProHazy dataset.

[0022] Figure 6 It is a diagram showing the influence of different β types on the proposed model of the method of the present invention visualized on the DENS-Fog (light) dataset, and only the influence on depth estimation is visualized here.

[0023] Figure 7 It is the effect of defogging and depth estimation of the method of the present invention on a continuous video frame. Detailed Embodiment

[0024] The present invention proposes a method for video dehazing and depth estimation in a real moving (driving) scenario, providing a clearer vision and distance perception for real fog scenarios, such as: driving scenarios and mobile phone photography in real foggy weather, etc. As Figure 2 shown, the combined video dehazing depth estimation has better effects compared to depth estimation alone in foggy weather and dehazing first and then depth estimation. This method is named DCL (Video Dehazing and Self-Supervised Depth Estimation with Depth as the Optimization Core), and it includes training and testing phases. In the training phase, three adjacent fog images I and three adjacent non-aligned and clear reference images J ref (corresponding to the fog images) are input. Here, the three adjacent non-aligned reference frames are obtained by non-aligned reference frame matching. This method uses three main loss functions and two regularization terms to train four sub-networks, namely the dehazing network Φ J , the depth estimation network Φ d , the scattering coefficient estimation network Φ β and the camera pose prediction network Φ p , the reconstruction loss function, the non-aligned reference loss function, and the brightness consistency constraint loss. Based on the above four sub-networks, the reconstruction loss function is composed of the perceptual, structural similarity, and L rec losses between I 1 and I, where I rec is the reconstructed fog image (calculated by the atmospheric scattering model). More importantly, this method combines the brightness consistency constraint (BCC) and the atmospheric scattering model (ASM) to jointly optimize the depth estimation network Φ d , and it is completed in a self-supervised manner. Here, the brightness consistency constraint uses adjacent frames after dehazing as input to obtain a more accurate camera pose, so as to more accurately complete reprojection using the estimated depth map and construct a self-supervised estimation method. With the self-supervised learning of the depth estimation network, we have respectively completed effective decoupling of each component of the entire ASM. At the same time, we use each component to reconstruct the fog frame to form a reconstruction loss constraint to supervise the learning of the entire framework. In this learning framework, our main innovation is to use the depth information shared by ASM and BCC to construct a jointly optimized learning framework. In the test and use phase, only a test image of real fog needs to be input, and the dehazing network Φ J and the depth estimation network Φ d can be used to quickly restore its clear scene. The video dehazing effect of DCL in real scenarios is shown in the attached drawings and reaches the current best effect.

[0025] The method of the present invention will be described in detail below with reference to the accompanying drawings.

[0026] Combined with Figure 1, A method for video dehazing and depth estimation in a real moving scenario. The real moving scenario of the present invention refers to the driving environment. Input sample pairs (real RGB hazy image I and non-aligned and clear reference image J′ t ), and the steps are as follows:

[0027] Step 1, Preprocess the dataset. First, use the non-aligned matching algorithm (NRFM) to obtain the matching information between all hazy video frames and clear reference frames, and then, according to the calibrated camera internal parameters, undistort the data and crop it from the size of 1920×1080 to 1600×512.

[0028] Step 2, Input the consecutive hazy video frames I [t-n:t+n] ∈R 3×h×w (n = 1) into the dehazing network Φ J to obtain the dehazing result [J t , J s (J t represents the current dehazed frame, and J s represents the neighboring dehazed frame of J t ). Input the current hazy video frame I t into the depth estimation network Φ d to estimate d t ∈R 1×h×w , and input it into the scattering coefficient estimation network Φ β to estimate β t ∈R 1×h×w , and then calculate the value of the infinite far atmospheric light A ∞ by using the dark channel method;

[0029] Step 3, Input the dehazing result J t of Step 2 into the camera pose prediction network Φ p to predict the pose information P x→y ∈R 4×4 , and then use the d d estimated by Φ t and Φ p to project the neighboring dehazed frames onto the current frame through reprojection, and then estimate the depth d x→y by self-supervised learning through calculating the photometric loss to constrain the depth estimation. t .

[0030] Step 4, According to J t , A ∞ , d t and β t generated in Steps 2-3, use the atmospheric scattering model to calculate and reconstruct the current hazy video frame I′ t , and then define I′ t and I tReconstruction loss function It consists of the sum of three loss functions, namely the first-order norm loss Perceptual loss and structural similarity loss Among them The perceptual content is calculated using VGG16

[0031] Step 5. In addition to the reconstruction loss in Step 4 This method also utilizes the non-alignment reference loss function regarding J t and J′ t

[0032] Step 6. In addition to the loss constraints in Steps 4 and 5, this method also proposes to use discriminator networks to respectively perform regularization constraints on the defogging network Φ J and the depth estimation network Φ d which are D MFIR and D MDR . They are respectively used to constrain Φ J to recover more high-frequency details and to constrain Φ d for the "black hole" problem that appears in weak texture regions

[0033] Step 7. According to the two loss functions in Steps 4, 5, and 6 and as well as the adversarial loss regularization of the two discriminator networks (D MFIR and D MDR ), the network parameters of the entire DCL framework are optimized (model training) to obtain the defogging result. Finally, the input test real RGB fog video frame image I t is respectively input into the defogging network and the depth estimation network to directly generate the clear scene image result J t and the depth map d t . Finally, we show the effects of continuous video defogging and depth estimation as Figure 7 shown

[0034] Preferably, in Step 1, the dataset is preprocessed. First, the matching information (the most similar clear reference frame) between all fog video frames and clear reference frames is obtained using the non-alignment matching algorithm, and then, according to the calibrated camera internal parameters, the data is de-distorted and cropped from the size of 1920×1080 to 1600×512. (Ensure that the calibration parameters of the camera are accurate)

[0035] Preferably, in Step 2, three networks are simultaneously used to respectively estimate the atmospheric scattering model parameters, namely, the depth estimation network Φ d is used to estimate d t , and the defogging network Φ J is used to predict J​t and the scattering coefficient estimation network Φ β to estimate β in t , where d t is the depth map of the current frame, and β t is the scattering coefficient of the current frame.

[0036] Input the consecutive fog video frames I [t-n:t+n] ∈ R 3×h×w (n = 1) into the defogging network Φ J to obtain the defogging result [J t , J s (J t represents the current defogged frame, and J s represents the neighboring defogged frame of J t ). The current fog video frame I t is respectively input into the depth estimation network Φ d to estimate d t ∈ R 1×h×w , and the scattering coefficient estimation network Φ β to estimate β t ∈ R 1×h×w , and then the value of the atmospheric light A ∞ is calculated by using the method of dark channel prior (DCP). (Φ d and Φ β are of the shared encoder network).

[0037] Preferably, in step 3, the defogged adjacent frames are used to perform the camera pose estimation, and the self-supervised learning of the depth estimation network Φ d is constrained by reprojection.

[0038] Input the defogging result J t of step 2 into the camera pose prediction network Φ p to predict the camera pose information P x→y ∈ R 4×4 , and then use the d d estimated by Φ t and Φ p to obtain P x→y by projecting the neighboring defogged frames onto the current frame through reprojection, and then the depth d t (x and y respectively refer to the pixel coordinates of the current frame and the neighboring frame) is estimated by calculating the photometric loss to constrain the depth estimation for self-supervised learning.

[0039] Preferably, in step 4, the J t , A ∞ , d t and β t generated in steps 1 - 3 are used to calculate and reconstruct the fog video frames by using the atmospheric scattering model Then, the real fog video frame I is utilized. t The reconstruction loss function is defined. And supervised learning is carried out on texture and color, where VGG16 is used to calculate the perceptual content (the content loss is calculated by selecting the features after the Relu layer).

[0040] Preferably, regarding J in step 5 t and J′ t the non - alignment reference loss function regularly constrains the feature distributions of color and texture of the defogging network, where VGG16 network is used to calculate the content loss between J t and J′ t .

[0041] Step 6: In addition to the loss constraints in steps 4 and 5, this method also proposes to use discriminator networks to respectively perform regular constraints on the defogging network Φ J and the depth estimation network Φ d , which are D MFIR and D MDR respectively. They are respectively used to constrain Φ J to recover more high - frequency details and to constrain the "black hole" problem that appears in the weak - texture area of Φ d .

[0042] Preferably, step 7 optimizes the network parameters (model training) of the entire DCL framework according to the two loss functions and in steps 4, 5 and 6, as well as the adversarial loss regularization of the two discriminator networks (D MFIR and D MDR ) to obtain the video defogging and depth estimation results. Finally, the input test real RGB fog video frame I t is input into the defogging network Φ J and the depth estimation network Φ d to directly generate the clear defogging J t and the depth estimation result d t .

[0043] The present invention establishes a video defogging and depth estimation framework (ASM - BCC) based on the atmospheric scattering model by sharing the same depth information between the atmospheric scattering model and self - supervised depth estimation, and optimizes the two tasks simultaneously. And two regularization discriminator networks D MFIR and D MDR are proposed to respectively enhance the high - frequency details of the defogging result and avoid inaccurate depth values caused by weak textures, such as the "black hole" problem of weak - texture road surfaces.

[0044] The present invention will be described in detail below in conjunction with embodiments.

[0045] Embodiment

[0046] As Figure 1 shown, a calculation process for video defogging and depth estimation in a real moving (driving) scenario is given. First, a continuous sequence of fog video frames I [t-n:t+n] ∈R 3×h×w (n = 1) is given, and then the non-aligned reference frame matching algorithm is used to match a clear reference frame with the highest similarity for each frame of this continuous video frame sequence. Then, we input the fog video frames into different atmospheric scattering model parameter prediction networks respectively. First, input into the defogging network Φ J to obtain the defogging result [j t , J s (J t represents the current defogged frame, and J s represents the neighboring defogged frame of J t ). Secondly, the current fog video frame I t is input into the depth estimation network Φ d to estimate d t ∈R 1×h×w , then the scattering coefficient estimation network Φ β is used to estimate β t ∈R 1×h×w . Finally, the value of the atmospheric light A ∞ is calculated by using the dark channel prior (DCP). Next, we will reconstruct the fog video frames by using the predicted atmospheric scattering model (ASM) parameters to construct a reconstruction loss function to constrain the learning of the entire framework. To ensure that each parameter variable can be fully disentangled, we establish a non-aligned reference loss function t about J t and J′ as a further supervised training signal for the defogging network, and jointly and are used to optimize the network parameters to obtain high-quality defogged images (in terms of color, texture, and brightness).

[0047] More importantly, we jointly use the atmospheric scattering model (ASM) and the brightness consistency constraint (BCC) to construct a jointly optimized paradigm for video defogging and depth estimation. We use the defogged video frames as the input of the camera pose network Φ p to predict more accurate camera pose information P x→y ∈R 4×4 . The more accurate camera pose provides a more effective supervision signal for constructing the reprojection self-supervised constraint. And the reconstruction of atmospheric scattering is for the depth estimation network Φd A reconstruction constraint based on the atmospheric scattering model is also provided, so the video dehazing and depth estimation are centered around depth and are a mutually compensatory learning process. We name the learning framework that simultaneously optimizes video dehazing and depth estimation centered around depth information on real fog videos as DCL. In addition, we also propose two discriminator networks D MFIR and D MDR respectively to enhance the high-frequency details in the dehazed video frames and to constrain the "black hole" problem that often occurs in predicting depths in weakly textured regions, such as a pure black road surface area.

[0048] The specific implementation steps of DCL are as follows: (Note that the size representation of the fog video frames is for a batch size of 1.)

[0049] Step 1: Preprocess the training dataset. First, after obtaining the camera internal parameters by photographing the data acquisition camera, use the camera internal parameters to perform a distortion correction operation on the video frames of the dataset, and then crop the resolution of the original video frames from 1920×1080 to 1600×512 that conforms to the normal driving perspective.

[0050] Step 2: For a given fog video frame segment I [t-n:t+n] ∈R 3×h×w (n = 1), the current frame is I t , and its neighboring frame is I s∈[t-n:t+n],s≠t ∈R 3×h×w , where h is the height and w is the width. In this step, we construct an ASM-BCC joint model by combining the atmospheric scattering model and luminance consistency. We define the ASM-BCC as follows:

[0051]

[0052] where J t is the dehazing result corresponding to the current fog video frame, that is, what the model needs to predict. J s is the neighboring frame of I s after dehazing. represents the bilinear interpolation sampling operation. x and y are the position coordinates of the corresponding pixel points on the current frame and the neighboring frame respectively. K represents the internal parameters of the camera.

[0053] Step 3: As Figure 1 shown, here we will define some different networks to separately predict the parameters J t , d t and β t of the atmospheric scattering model. We define three different networks Φ J , Φ d and Φ βEstimate these three key parameters separately, and a pose estimation network Φ p To predict the pose information P of the camera x→y ∈R 4×4 Here, the atmospheric light A ∞ is obtained through the calculation of DCP. We formulate this process as follows:

[0054]

[0055] where Φ J is the defogging network, Φ d is the depth estimation network, and Φ β is the scattering coefficient estimation network. Immediately afterwards, we use the estimated depth information d t and the pose P of the camera x→y to reconstruct the current defogged frame J t . The specific process is as follows:

[0056]

[0057] where SSIM represents the structural similarity loss, represents the bilinear interpolation sampling operation, and α = 0.85 which is a default setting. m a is a mask obtained by the automatic masking strategy. In addition, in this step, in order to obtain a smoother depth estimation result, we use the edge smoothing loss to constrain the depth estimated by Φ d . The specific constraint expression is as follows:

[0058]

[0059] Here, represents the average normalized inverse depth, and represent the gradients in the horizontal and vertical directions respectively.

[0060] Step 4: As Figure 1 shown, in Step 3, the J t , d t and β t predicted in the above steps are brought into the atmospheric scattering model to calculate the reconstructed fog video frame . Then, using the real fog video frame I t , three loss functions ( and ) are constructed to constrain the reconstructed I′ t in terms of texture and color to obtain a realistic reconstruction effect, It uses VGG16 to calculate the content loss (select the Relu layer to calculate the content loss, specific layers: 3, 8, 15, 22, 29). Note: The representation of the loss formula is for a pair of samples.

[0061]

[0062] Among them, the parameters θ, λ, and η are default set to 1. Φ l (I′ t ) and Φ l (I) respectively represent the feature maps corresponding to I rec and I at layer l, N represents the number of extracted feature layers, and represent the mean value of the image I′ t . and represent the variance of the image I t . represents the covariance between the image I′ t and I t .

[0063] Step 5, as Figure 1 shown, the present invention uses the non - aligned reference loss to constrain the color and texture of the generated J t to be closer to the reference image. Similarly, it uses the VGG16 network to calculate the content loss between the generated clear scene graph J t and the non - aligned reference frame J′ t .

[0064]

[0065] Among them, J′ t is the non - aligned clear reference frame image, and J t is the clear scene graph generated by the model of the present invention. is the content similarity between image features, Ω l (J t ) and Ω l (J′ t ) respectively represent the feature maps corresponding to the images J t and J′ t at layer l extracted by the VGG16 network.

[0066] Step 6, as Figure 1 shown, we respectively propose two discriminative networks D MFIR and D MDR to regularize the results of video de - fogging and depth estimation. D MFIR is used to enhance the high - frequency details of the de - fogging result, DMDR The "black hole" problem that occurs in the depth for regular estimation in weakly textured regions, such as: a pure black road surface. We respectively define D MFIR and D MDR as follows:

[0067]

[0068] Here is a normalization operation to remove the scale and make the training model more stable. cat(a,b) represents concatenating a and b, represents extracting the high-frequency features of the defogged video frame J t using the Haar wavelet and concatenating them. represents extracting the high-frequency features of the non-aligned clear reference frame J t using the Haar wavelet and concatenating them. D J is the defogging result discriminator network, D d is the depth estimation result discriminator network, is an expectation function, is the loss of the defogging result on the discriminator, is the loss of the defogging result on the generator, is the loss of the depth estimation on the discriminator, is the loss of the depth estimation on the generator.

[0069] Step 7, as Figure 1 shown, optimize the network parameters (model training) of the entire DCL framework according to the losses and regularization constraints in Steps 4, 5, and 6 to obtain the trained clear scene graph (J t ).

[0070] The present invention compares several existing SOTA defogging methods, namely DCP, RefineNet, CDD-GAN, D 4, PSD, RIDCP, PM-Net, MAP-Net, NSDNet, and DVD. We first evaluate the proposed DCL on three video dehazing benchmark datasets, namely GoProHazy, DrivingHazy, and InternetHazy. It is worth noting that during the testing process, our DCL model only requires a dehazing network. The results in unpaired, paired, and unaligned settings are summarized in Table 1. Overall, our DCL ranks first in terms of both FADE and NIQE among all methods. For example, the NIQE score of DCL is 22.04% higher than that of the second-best unaligned method DVD (Fan et al. 2024), and it also significantly outperforms other top methods in paired and unpaired methods. Figure 3 Shows a visual comparison with PM-Net (Liu et al. 2022b), RIDCP (Wu et al. 2023), NSDNet (Fan et al. 2023), MAPNet (Xu et al. 2023), and DVD (Fan et al. 2024). Although these methods generally produce visually satisfactory dehazing results, our DCL recovers clearer predictions with more accurate content and contours. In addition, our DCL is also able to estimate effective depth, which is not present in other dehazing methods.

[0071] Table 1 presents the quantitative results for three real-world blurred video datasets. ↓ indicates that lower is better. For dehazing methods that rely on ground truth for training, we use It is trained on the GoPro dataset. The DrivingHazy and InternetHazy datasets are tested using the dehazing model trained on GoProHazy. Note that all quantitative results are evaluated at an output resolution of 640×192.

[0072] Table 1: Quantitative Results for Three Real-World Blurred Video Datasets

[0073]

[0074]

[0075] In addition to the image and video dehazing methods compared in the above table, to verify the effectiveness of our proposed model in depth estimation, we also compared it with the current state-of-the-art self-supervised depth estimation methods and self-supervised depth estimation methods in extreme environments. To further verify the performance on the depth estimation task, we compared DCL with well-known self-supervised depth estimation methods, including MonoDepth2 (Godard et al. 2019), MonViT (Zhao et al. 2022), RobustDepth (Saunders, Vogiatzis, and Manso 2023), Mono-ViFI (Liu et al. 2024), and Lite-Mono (Zhang et al. 2023). Since the blurred data with depth annotations is very scarce, we trained these methods on the GoProHazy benchmark and then tested them on the DENSE-Fog dataset. It should be noted that during the test phase, our DCL only used the depth estimation subnetwork. Table 2 reports the numerical results. Overall, in both light fog and heavy fog scenarios, our DCL shows the best performance in almost all five metrics. In the heavy fog scenario, due to the blurred depth being closer to the mean of the true value, the accuracy and error metrics are inconsistent, resulting in smoother results. Figure 4 Visual results of MonoViT, Mono-ViFI, and DCL are shown. It can be found that the depths estimated by MonoViT and Mono-ViFI are blurred and unreliable. In contrast, our DCL can predict more reasonable and accurate depth results. In addition, DCL can also provide clean dehazed images.

[0076] Table 2 shows the comparison of quantitative results. We compared the framework of the present invention with the previous state-of-the-art methods on the DENSE-Fog dataset. All methods were trained on the GoProHazy dataset. It should be noted that RobustDepth was trained on the clear reference videos in GoProHazy because it used its own synthetic fog for training.

[0077] Table 2: Comparison of Quantitative Results

[0078]

[0079] In addition to the above comparisons, we also made a comparison in terms of the efficiency of the model. We compared the number of parameters, FLOPs, and inference time of state-of-the-art methods in image / video dehazing and self-supervised depth estimation tasks. The tests were conducted on an NVIDIA RTX4090 GPU. The running time was calculated when the input size was 640×192. As shown in Tables 1 and 2, our method achieved the shortest inference time in image / video dehazing and self-supervised estimation tasks, 0.075 seconds and 0.009 seconds respectively, indicating that DCL performs well in fast inference. All these facts prove that the proposed DCL performs very well based on our ASM-BCC model.

[0080] Table 3: Ablating the core module on the DENSE-Fog(light) dataset

[0081] Method BCC <![CDATA[D MFIR > <![CDATA[D MDR > Abs Real↓ RMSE log↓ <![CDATA[δ 1 ↑]]> DCL w / o BCC √ √ 0.636 0.569 0.439 <![CDATA[DCLw / oD MFIR > √ √ 0.320 0.366 0.621 <![CDATA[DCL w / oD MDR > √ √ 0.340 0.392 0.562 DCL(Ours) √ √ √ 0.311 0.364 0.623

[0082] To evaluate the effectiveness of our proposed BCC, D MFIR and D MDR we conducted experiments by excluding each component and training our model. The results in Table 3 and Figure 5 show that after integrating our BCC module, the video dehazing effect is significantly improved. This improvement stems from using dehazed images to enforce the brightness consistency constraint, which allows for better depth estimation (d) and more effective fog removal. In addition, D MFIR and D MDR also contribute to the improvement of dehazing and depth estimation.

[0083] Table 4: Ablating different loss functions on the GoProHazy dataset.

[0084]

[0085] To evaluate the impact of different loss functions on the proposed model, we conducted corresponding ablation experiments to verify the effectiveness of the loss functions and on GoProHazy. The results of FADE and NIQE are reported in Table 4. Obviously, plays a key role because the ASM model is an important physical mechanism for video dehazing in Equation (1), which ensures the independence between the dehazing result and the unaligned clear reference frame. In addition, the smoothness loss also makes important contributions to video dehazing and depth estimation.

[0086] Table 5: Comparing different β types on the DENSE-Fog(light) dataset

[0087] Shape ofβ Type Abs Real↓ RMSE log↓ <![CDATA[δ 1 ↑]]> (1,1,1) Constant 0.325 0.371 0.621 (1,192,640)(Ours) Non-uniform 0.311 0.364 0.623

[0088] In the method section description, we assume that β is a non-uniform variable, mainly because haze in real-world scenarios is usually non-uniform, such as local haze. Therefore, the scattering coefficient varies between different regions. Since existing studies mainly use a constant scattering coefficient, in this study, we conducted comparative experiments for different beta values (i.e., constant or non-uniform) in Table 5. The results show that non-uniform β can bring better depth estimation accuracy. We further demonstrate the visual comparison in Figure 6 with a visual comparison.

Claims

1. A video defogging and depth estimation method in a real mobile scene, characterized in that: include: Step 1: Preprocess the data set; The matching relationship between all foggy video frames and clear reference frames is obtained by using the non-aligned reference matching algorithm. The data is dedistorted and cropped according to the calibrated camera intrinsic parameters. Step 2: Transform the continuous fog video frames I [t-n:t+n] ∈R 3×h×w Input to the dehazing network Φ J The dehazing result is obtained in [J t , J s ], J t represents the current defogging frame, J s Indicates J t Neighboring defogging frames, n = 1; current foggy video frame I t Input into the depth estimation network Φ respectively d Medium estimate d t ∈R 1×h×w , and the scattering coefficient estimation network Φ β Estimated β t ∈R 1×h×w , using the dark channel method to calculate the value of the infinite atmospheric light A ∞ ; Step 3: The defogging result of step 2 t Then input into the camera pose prediction network Φ p To predict the camera’s position information p x→y ∈R 4×4 , and then use Φ d The estimated d t and Φ p Get p x→y By reprojecting the adjacent dehazed frames onto the current frame, the depth d is estimated by self-supervised learning by calculating the photometric loss constrained depth estimation. t ; Step 4: J generated from step 2 to step 3 t , A ∞ , d t and β t , using the atmospheric scattering model to calculate and reconstruct the current fog video frame I′ t , then define I′ t and I t The reconstruction loss function Contains the sum of three loss functions, namely the first-order norm loss Perceived loss and structural similarity loss Step 5: Build about J t and J′ t The non-aligned reference loss function J t and J′ t The non-aligned reference loss function Regularize the color and texture feature distribution of the dehazing network; Step 6: Use the discriminator network to analyze the dehazing network Φ J and the depth estimation network Φ d Perform regularization constraints; Step 7: Based on the two loss functions of step 4, step 5, and step 6 and And two discriminator networks D MFIR and D MDR The adversarial loss regularization optimizes the network parameters of the entire DCL framework to obtain the defogging result; finally, the real RGB fog video frame image I is input for testing t , respectively input into the dehazing network and the depth estimation network, and directly generate a clear scene image result J t and the depth map d t .

2. The video defogging and depth estimation method in a real mobile scene according to claim 1, characterized in that: In step 1, the calibrated camera intrinsic parameters are used to dedistort and crop the video frame to a size of 1600×512.

3. The video defogging and depth estimation method in a real mobile scene according to claim 2, characterized in that: Step 2 uses three networks to estimate the atmospheric scattering model parameters respectively, using the depth estimation network Φ d Estimated t , dehazing network Φ J Prediction J t and the scattering coefficient estimation network Φ β Estimated β t , where d t is the depth map of the current frame, β t is the scattering coefficient of the current frame.

4. The video defogging and depth estimation method in a real mobile scene according to claim 3, characterized in that: Step 3 uses the adjacent frames after defogging to estimate the camera posture and constrains the depth estimation network Φ by reprojection. d Self-supervised learning.

5. The video defogging and depth estimation method in a real mobile scene according to claim 4, characterized in that: In step 4, the predicted J t , A ∞ , d t and β t Bring it into the atmospheric scattering model to calculate and reconstruct the fog video frame Reusing real fog video frames I t and reconstructed fog video frame I′ t , construct the reconstruction loss function to I′ t Perform unsupervised learning of texture and color, where Use VGG16 to calculate the perceived content.

6. The method for video defogging and depth estimation in a real mobile scene according to claim 5, characterized in that: Multi-scale non-aligned reference loss in step 5 Supervised learning of color and texture for the dehazing network, where The VGG16 network is used to calculate J t and J′ t The feature distribution loss between .

7. The method for video defogging and depth estimation in a real mobile scene according to claim 6, characterized in that: Step 6 uses D MFIR and D MDR The dehazing results with richer high-frequency details are obtained by separately constraining the model, and accurate depth information is produced in the weak texture dehazing of the road surface.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the video defogging and depth estimation method in a real mobile scene as described in any one of claims 1-7 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the video defogging and depth estimation method in a real mobile scene as described in any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for video defogging and depth estimation in a real mobile scene described in any one of claims 1 to 7 is implemented.