Semi-supervised image defogging method based on Brownian motion bridge diffusion model
Through the combination of Brownian motion bridge diffusion model and differential convolution, the problem of unstable image defog training in the prior art is solved, effective defog removal and clarity recovery under unpaired data is achieved, and the visual quality of the image is improved.
Patent Information
- Application Number
- CN202510405997.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
Existing unsupervised or semi-supervised deep learning methods are unstable in training during image defog removal and are prone to pattern collapse. The existing diffusion-based defog removal model relies on a large amount of paired data, making it difficult to effectively remove haze and restore clarity and contrast.
The semi-supervised image defogging method based on the Brownian motion bridge diffusion model is adopted. By extracting the features of clear images and foggy images, the joint distribution is decoupled using the EM algorithm, and the image relationship is captured through the unified Brownian bridge diffusion model, and feature extraction is performed by combining differential convolutions (ConvDC, ConvAD, ConvHD, ConvVD and standard convolution). The KL divergence is designed as the objective function to perform unsupervised image defogging.
The image defogging training is achieved with stable unpaired data, which improves the model's characterization ability and generalization performance, restores the image's clarity and structural integrity, and the generated images are closer to the real fogging-free effect.
Smart Images

Figure CN120339128A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing technology, in particular to a semi-supervised image dehazing method based on a Brownian motion bridge diffusion model. Background Art
[0002] Haze is caused by the scattering effect of aerosol particles in the atmosphere, which results in photos taken in hazy weather usually showing low contrast and clarity. Dehazing methods aim to remove haze and enhance the contrast and color integrity of actual hazy images, which play a crucial role in the application of computer vision tasks such as image segmentation and object detection under hazy weather conditions.
[0003] In recent years, many unsupervised or semi-supervised deep learning methods have emerged for dehazing. However, directly inheriting the dehazing framework from unpaired image-to-image conversion methods to utilize unpaired hazy and haze-free images is not sufficient because the training process is usually unstable and prone to mode collapse problems. Diffusion models have shown powerful capabilities in modeling data distributions compared to unsupervised or semi-supervised deep learning methods, but existing diffusion-based dehazing models still rely heavily on a large amount of paired data for training. Summary of the Invention
[0004] The purpose of the present invention is to provide a semi-supervised image dehazing method based on a Brownian motion bridge diffusion model, including: step S100, extracting the features of clear images and hazy images; step S200, using the EM algorithm to decouple the joint distribution of hazy images and clear images into two conditional distributions, and capturing the relationship between hazy images and clear images through a unified Brownian bridge diffusion model; step S300, performing unsupervised image dehazing on unpaired clear images and hazy images.
[0005] Further, ConvDC, ConvAD, ConvHD, ConvVD, and a standard convolution are used for feature extraction.
[0006] Further, step S200 specifically includes: step S201, encoding the features of clear image x and hazy image y of the same scene to obtain a state distribution; step S202, using the EM algorithm to decouple the joint distribution of paired hazy images and clear images into two conditional distributions q(y|x) and p(x|y), where q(y|x) is the conditional distribution of hazy image y given clear image x; p(x|y) is the conditional distribution of clear image x given hazy image y; step S203, designing the KL divergence as the objective function and using a unified Brownian bridge diffusion model to learn the two conditional distributions.
[0007] Further, the specific process of step S202 is as follows: step S2021, let q(y|x) = qθ (y|x)q(x), p(x|y) = p θ (x|y)p(y), where q θ (y|x) and p θ (x|y) are Gaussian distributions modeled using a unified Brownian bridge diffusion model, q(x) is the prior distribution of the clear image, p(y) is the marginal distribution of the hazy image, and θ represents the relationship between the hazy image and the clear image; Step S2022, assume p θ (x|y) is a constant, q θ (y|x) is a latent variable, and calculate the expected value of the latent variable in the r-th iteration ; Step S2023, maximize the likelihood function to optimize the latent variable Step S2024, repeatedly execute Step S2023 and Step S2024 until the parameter θ converges.
[0008] Furthermore, the maximization of the likelihood function in Step S2023 is expressed as follows
[0009]
[0010] where
[0011] Furthermore, θ is optimized by minimizing the negative evidence lower bound. (r)
[0012] Furthermore, Step S203 specifically includes: Step S2031, construct a Markov chain to estimate the prior distribution of the Brownian bridge forward diffusion process; Step S2032, estimate the posterior distribution of the Brownian bridge reverse denoising process; Step S3203, design a joint KL divergence objective function to optimize the prior distribution and the posterior distribution.
[0013] Furthermore, in Step S2031, a Markov chain is constructed to estimate the prior distribution q BB (z t |z0, z T ), as shown in the following formula
[0014]
[0015] where T is the total number of steps of the diffusion process, I is the identity matrix, δ t is the variance, and the noise variance δ of the Brownian bridge process is used t = 2s(m t - m t 2 ), s is the time variable; the transition probability q BB (z t |zt-1 , z T )
[0016]
[0017] Among them,
[0018] Furthermore, in step S2032, the posterior distribution p BB (z0|z T ) is estimated
[0019]
[0020] Among them, is the reverse transition kernel from z t to z t-1 with learnable parameter θ, and the transition probability p BB (z t-1 |z t , z T )
[0021]
[0022] Among them, μ θ is the predicted mean of the noise, is the noise variance per step.
[0023] Furthermore, the joint KL divergence is
[0024]
[0025] Among them, q(x) and p(y) do not contain any parameters.
[0026] Compared with the prior art, the present invention has the following advantages: The present invention provides a semi-supervised image defogging framework, which uses a Brownian bridge diffusion model to enhance the performance of semi-supervised image defogging and achieve stable training; (2) Feature extraction is performed using ConvDC, ConvAD, ConvHD, ConvVD and a standard convolution, which can effectively capture gradient hierarchical information and greatly improve the representation ability and generalization performance of the model.
[0027] The present invention will be further described below with reference to the accompanying drawings of the specification. Description of the Drawings
[0028] Figure 1 is a schematic flow chart of the method of the present invention.
[0029] Figure 2 is a schematic diagram of the visualization results of various methods on the data set.
[0030] Figure 3Schematic diagram for RDC effectiveness evaluation. Detailed implementation manner
[0031] Combined with Figure 1 , a semi-supervised image dehazing method based on the Brownian motion bridge diffusion model, comprising:
[0032] Step S100, extracting the features of the clear image and the hazy image;
[0033] Step S200, using the EM algorithm to decouple the joint distribution of the hazy image and the clear image into two conditional distributions, and capturing the relationship between the hazy image and the clear image through a unified Brownian bridge diffusion model;
[0034] Step S300, performing unsupervised image dehazing on the unpaired clear image and the haze image.
[0035] In step S100, four differential convolutions (ConvDC, ConvAD, ConvHD, ConvVD) and a standard convolution are used for feature extraction. The standard convolution can extract local features in the image, such as edges, textures, etc. It slides the convolution kernel over the image and performs weighted summation on each local area to capture the basic features of the image. ConvDC (diagonal direction differential convolution) can capture the changes in the diagonal direction of the image and is very effective for feature extraction of some images with diagonal textures or structures; ConvAD (horizontal and vertical direction differential convolution) combines the differentials in the horizontal and vertical directions and can capture the changes in the horizontal and vertical directions of the image simultaneously, enhancing the perception of the overall structure of the image; ConvHD (horizontal direction differential convolution) focuses on feature extraction in the horizontal direction and can better capture the horizontal texture and edge information in the image; ConvVD (vertical direction differential convolution) focuses on feature extraction in the vertical direction and helps to extract the vertical texture and edge information in the image. Although the standard convolution can extract basic features, it may have limitations in retaining detailed information. The differential convolution can more sensitively capture the tiny changes and detailed information in the image by calculating the differences between adjacent pixels. For example, in a foggy image, the horizontal direction differential convolution can better capture the horizontal texture blurred by the fog, and the vertical direction differential convolution can capture the detailed changes in the vertical direction. These detailed information are crucial for restoring the clarity and structural integrity of the image and help to better restore the original details and structure of the image during the defogging process. Since the differential convolution can extract features from multiple directions, the model has stronger adaptability to image changes in different directions. In semi-supervised learning, the distributions of labeled data and unlabeled data may be different. The multi-directional feature extraction ability of the differential convolution can help the model better learn the common features under different data distributions, thereby improving the robustness and generalization ability of the model to different types of foggy images. In Brownian bridge diffusion, the multi-directional features extracted by the differential convolution can provide richer detailed information and structural information for the Brownian bridge diffusion model. These information can help the model more accurately simulate the diffusion process of the fog and its impact on the image, thereby more effectively removing the fog and restoring an effect closer to the real fog-free image.
[0036] The specific process of step S200 is as follows:
[0037] Step S201, encode the features of the clear image x and the hazy image y of the same scene to obtain the state distribution;
[0038] Step S202, use the EM algorithm to decouple the joint distribution of paired hazy images and clear images into two conditional distributions;
[0039] Step S203: Design the KL divergence as the objective function and use a unified Brownian bridge diffusion model to learn two conditional distributions.
[0040] The Brownian bridge diffusion is divided into a forward diffusion process and a reverse denoising process. In the forward diffusion process, within the time range from t = 0 to t = T, starting from the initial state distribution z0 of the clear image x, noise is gradually injected within the time steps, and finally the state distribution z of the hazy image y is reached. T For the prior distribution q BB (z t |z0,z T ) is estimated; in the reverse denoising process, within the time range from t = T to t = 0, starting from the initial state z T of the hazy image y, noise is gradually removed within the time steps, and finally the state distribution z0 of the clear image x is reached, and the posterior distribution p BB (z0|z T ) is estimated. In step S201, during the forward diffusion process of the Brownian bridge, the initial state z0 = x and the end state is z T = y; in the reverse denoising process of the Brownian bridge, the initial state z0 = y and the end state is z T = x.
[0041] In step S202, the EM algorithm is used to decouple the joint distribution q(x,y) into two conditional distributions q(y|x) and p(x|y); where q(y|x) is the conditional distribution of the hazy image y given the clear image x; p(x|y) is the conditional distribution of the clear image x given the hazy image y. The joint distribution q(x,y) describes the probability that the clear image x and the hazy image y appear simultaneously. The specific process of step S202 is as follows:
[0042] Step S2021: Let q(y|x) = q θ (y|x)q(x), p(x|y) = p θ (x|y)p(y), where q θ (y|x) and p θ (x|y) are Gaussian distributions modeled using a unified Brownian bridge diffusion model, q(x) is the prior distribution of the clear image, p(y) is the marginal distribution of the hazy image; θ is the optimization parameter representing the relationship between the hazy image and the clear image;
[0043] Step S2022: Assume that p θ (x|y) is a constant and q θ (y|x) is a latent variable, and calculate the expected value of the latent variable in the r - th iteration;
[0044] Step S2023: Maximize the likelihood function to optimize the latent variable
[0045] In step S2024, steps S2023 and S2024 are repeatedly executed until the parameter θ converges.
[0046] The likelihood function in step S2023 is expressed as shown in formula (1)
[0047]
[0048] Since p θ (x|y) is a constant, formula (1) is transformed into formula (2) to optimize the latent variable for optimization
[0049]
[0050] where
[0051] The parameter θ is optimized by minimizing the negative evidence lower bound (ELOB) function L of formula (3) (r) for optimization
[0052]
[0053] where the target distribution q BB (z t-1 |z t ,z0,z T )
[0054]
[0055] is a deep neural network with parameter θ (r) , aiming to predict z0;
[0056]
[0057] In the forward diffusion process of step S203, the prior distribution q BB (z t |z0,z T ) is estimated by constructing a Markov chain. The constructed Markov chain is shown in equation (5)
[0058]
[0059] where T is the total number of steps of the diffusion process, I is the identity matrix, δ t is the variance, and the noise variance δ t = 2s(m t -m t 2 ) of the Brownian bridge process is adopted, and s is the time variable.
[0060] The transition probability q is obtained according to formula (5). BB (z t |z t-1 ,z T )
[0061]
[0062] Among them,
[0063] In step S203, the posterior distribution p BB (z0|z T ) is estimated through formula (7).
[0064]
[0065] Among them is the reverse transition kernel from z t to z t-1 with learnable parameters θ; the transition probability p BB (z t-1 |z t ,z T ) is further obtained through the formula
[0066]
[0067] Among them, μ θ is the predicted mean of the noise, is the noise variance per step.
[0068] By designing the joint KL divergence as the objective function, the prior distribution q BB (z t |z0,z T ) and the posterior distribution p BB (z0|z T ) are learned and optimized, where formula (9) represents the KL divergence
[0069]
[0070] Among them, q(x) and p(y) do not contain any parameters and can be regarded as constants.
[0071] Embodiment
[0072] In this embodiment, the proposed method is trained and evaluated on the publicly available datasets SOTS and NH-HAZE 2. The ratio of paired data to unpaired data is 1:1. The SOTS synthetic dataset includes indoor and outdoor subsets. In the indoor dataset, there are 500 pairs of hazy and haze-free images. This embodiment randomly divides the dataset into 450 pairs of training data and 50 pairs of test data. Among the 450 pairs of training data, 225 pairs are randomly selected as the paired dataset, and the remaining 225 hazy images are used as the unpaired dataset. The outdoor dataset is processed in the same way as the indoor dataset. For NH-HAZE2, this embodiment also divides the training set into paired data and unpaired data at a ratio of 1:1.
[0073] This embodiment compares and evaluates the performance of the SID-BBDM of the present invention with a variety of the latest dehazing methods, including the prior-based method (DCP ] ), the supervised method (DehazeNet ] , AOD-Net, MSCNN ] , GDN) trained on the entire training set, the unsupervised methods (CycleDehaze, YOLY, USID-Net, D4, ODCR) trained on the completely unpaired training set, and the semi-supervised methods (DA, PSD, DTS) using the same training set as ours. For fair comparison, this embodiment uses the official codes provided by their respective authors to retrain these methods under the same experimental settings. For the methods that cannot be retrained, this embodiment directly uses the provided models for testing.
[0074] 1. Training details
[0075] The proposed method is implemented using PyTorch 1.12.1 and trained on a computer equipped with an Intel(R) Core(TM) i5-13600K CPU@5.10GHz and an NVIDIA GeForce RTX 4090 GPU. The framework consists of two components: a pre-trained VQGAN model and the proposed Brownian bridge diffusion model. This embodiment adopts the same pre-trained VQGAN model as in the Latent Diffusion Model. The number of time steps in the training phase is set to 1000, while 200 sampling steps are used in the inference phase to balance sample quality and efficiency. The optimizer Adam is selected, using the default settings of PyTorch and a fixed learning rate of 5e-5. All training samples are resized to 256×256.
[0076] The training strategy of this embodiment is divided into two stages: in the first stage, training is carried out using limited paired data, and in the second stage, only a large amount of unpaired data is used for training. It should be noted that in the second stage, the model of the first stage is not fine-tuned, but is independently trained completely based on unpaired data.
[0077] 2. Experimental Evaluation
[0078] (1) Quantitative Analysis
[0079] Quantitative comparison on the dataset in Table 1. The best results are in bold, and the second-best results are underlined.
[0080] Table 1
[0081]
[0082] As shown in Table 1, this embodiment conducts a detailed quantitative comparison of various state-of-the-art dehazing methods on multiple datasets (including SOTS-indoor, SOTS-outdoor, and NH-HAZE2). SID-BBDM achieves the highest PSNR (28.96 dB) on the SOTS-Outdoor dataset and also reaches the highest PSNR (19.37 dB) on the NH-HAZE 2 dataset, and performs well in terms of SSIM value. The above results indicate that SID-BBDM inherits the advantages of diffusion models in generating realistic images. In addition, this embodiment also evaluates different methods using non-reference evaluation metrics NIQE and MUSIQ, and the results are shown in Table 2.
[0083] (2) Qualitative Analysis
[0084] Table 2 Evaluates NIQE and MUSIQ, the best results are in bold, and the second-best results are underlined.
[0085] Table 2
[0086]
[0087] As Figure 2 shown, this embodiment compares the visualization effects of the SID-BBDM method with previous state-of-the-art methods on the SOTS-indoor and SOTS-outdoor datasets. Methods such as YOLY, GDN, and USID-Net are prone to causing darkened areas and blurred contours, while PSD introduces color distortion. In contrast, SID-BBDM restores clearer contours and edges, and there is less residual haze in the output. DCP and YOLY occasionally fail when processing the sky region, resulting in obvious color shifts and artifacts in the processed images. However, the results of SID-BBDM are closer to the true values, presenting a more visually appealing dehazing effect.
[0088] (3) Ablation Experiments
[0089] SID-BBDM introduces DDIM to accelerate inference by leveraging non-Markov processes. By evaluating on the SOTS-indoor dataset with different sampling steps, the impact of the number of sampling steps in the reverse diffusion process on the performance of SID-BBDM is studied. The quantitative scores for the defogging task are reported in Table 3. The results show that when the number of sampling steps is small (less than 200), the image quality improves rapidly as the number of steps increases. When the number of steps exceeds 200, the PSNR and SSIM metrics only show slight improvements.
[0090] Table 3
[0091]
[0092] The variance of the Brownian Bridge characterizes its temporal uncertainty and volatility. The dynamic evolution of this variance is used to control the inherent randomness in the generation process. The Brownian Bridge diffusion process tends to stabilize at the two endpoints while maintaining high randomness in the intermediate stages, thus ensuring sufficient diversity and flexibility in the generation process. This variance regulation is crucial for achieving a smooth transition between the initial and target distributions. By scaling the maximum variance of the Brownian Bridge by a factor s at t = T / 2, the quality of the generated images can be adjusted. In this example, s ∈ {0.5, 1, 2, 4} is experimented with to evaluate its impact on the performance of SID-BBDM. The quantitative results are shown in Table 3. As s increases, the quality and realism of the generated images decrease. If the original variance formula of the unscaled Brownian Bridge is used, SID-BBDM cannot generate reasonable samples due to the excessive maximum variance.
[0093] Table 4 RDC Effectiveness Evaluation
[0094]
[0095] To verify the effectiveness of the proposed RDC module, an ablation study is conducted in this example to evaluate its contribution through 500k training iterations. Specifically, the ResBlock in UNet is used as the baseline in this example, and five parallel convolutional layers are gradually removed, and the results are summarized in Table 4. As shown in the results, the PSNR performance steadily improves from 27.60 dB to 28.85 dB, and the SSIM shows a similar trend. As Figure 3As shown, ConvAD focuses on regions with significant angular changes, while ConvHD and ConvVD mainly capture gradient changes in the horizontal and vertical directions, highlighting edge or texture changes. ConvCD enhances the changes in details and textures. The RDC of SID-BBDM effectively provides a comprehensive representation ability to capture diverse textures, edges, and fine details in images.
Claims
1. A semi-supervised image defogging method based on the Brownian motion bridge diffusion model, characterized in that Including: Step S100, extracting the features of the clear image and the hazy image; Step S200, using the EM algorithm to decouple the joint distribution of the hazy image and the clear image into two conditional distributions, and capturing the relationship between the hazy image and the clear image through a unified Brownian bridge diffusion model; Step S300, performing unsupervised image dehazing on unpaired clear images and haze images.
2. The method according to claim 1, wherein In step S100, ConvDC, ConvAD, ConvHD, ConvVD and a standard convolution are used for feature extraction.
3. The method according to claim 1, characterized in that Step S200 specifically includes: Step S201, encoding the features of the clear image x and the haze image y of the same scene to obtain the state distribution; Step S202, using the EM algorithm to decouple the joint distribution of the paired haze image and the clear image into two conditional distributions q(y|x) and p(x|y), where q(y|x) is the conditional distribution of the haze image y when the clear image x is given; p(x|y) is the conditional distribution of the clear image x when the haze image y is given; Step S203, designing the KL divergence as the objective function and using a unified Brownian bridge diffusion model to learn the two conditional distributions.
4. The method according to claim 3, wherein The specific process of step S202 is as follows: Step S2021, let q(y|x) = q θ (y|x)q(x), p(x|y) = p θ (x|y)p(y), where q θ (y|x) and p θ (x|y) are Gaussian distributions modeled using a unified Brownian bridge diffusion model, q(x) is the prior distribution of the clear image, p(y) is the marginal distribution of the hazy image, and θ represents the relationship between the hazy image and the clear image; Step S2022, assume p θ (x|y) is a constant, q θ (y|x) is a latent variable, calculate the expected value of the latent variable in the r-th iteration; Step S2023, maximizing the likelihood function to optimize the latent variable Step S2024, repeatedly execute step S2023 and step S2024 until the parameter θ converges.
5. The method according to claim 4, characterized in that The maximum likelihood function in step S2023 is expressed as follows Among them, 6. The method according to claim 5, wherein Optimize θ by minimizing the negative evidence lower bound (r) Do the optimization 7. The method according to claim 3, characterized in that, Step S203 specifically includes: Step S2031, constructing a Markov chain to estimate the prior distribution of the Brownian bridge forward diffusion process; Step S2032, estimating the posterior distribution of the Brownian bridge reverse denoising process; Step S3203, designing a joint KL divergence objective function to optimize the prior distribution and the posterior distribution.
8. The method according to claim 7, wherein In step S2031, a Markov chain is constructed to estimate the prior distribution q BB (z t |z0,z T ), as shown in the following formula where T is the total number of steps of the diffusion process, and I is the identity matrix, δ t is the variance, and the noise variance δ of the Brownian bridge process is used t = 2s(m t - m t 2 ), where s is the time variable; Excessive probability q BB (z t |z t-1 ,z T ) Among them, 9. The method according to claim 8, characterized in that In step S2032, the posterior distribution p BB (z0|z T ) is estimated according to the following formula Among them, is the reverse transition kernel from z t to z t-1 with learnable parameters θ and transition probability p BB (z t-1 |z t ,z T ) where μ θ is the predicted mean of the noise, and is the noise variance per step.
10. The method according to claim 9, wherein The joint KL divergence is where q(x) and p(y) do not contain any parameters.