Method and device for restoring bad weather degraded image based on dynamic degradation generation

By building a diffusion model and a dynamic degradation generator, combining the expected maximization algorithm and semi-supervised learning, the problem of insufficient generalization ability of image restoration models under severe weather conditions is solved, and high-quality and robust image restoration effect is achieved.

CN120259142AActive Publication Date: 2025-07-04HUAQIAO UNIVERSITY

Patent Information

Application Number
CN202510748261.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The prior art image restoration model has limited generalization capabilities under severe weather conditions, and its recovery effect is unstable, making it difficult to adapt to a variety of severe weather scenarios.

Method used

A conditional diffusion model and dynamic degradation generator are built, and the hidden variable optimization is optimized using the expected maximization algorithm, and semi-supervised training is performed in combination with labeled and unlabeled data sets. Denoising is performed through the diffusion time step of the Markov chain, degradation layers are generated and semi-supervised learning is performed.

Benefits of technology

It realizes high-quality image restoration under different bad weather conditions, improves the generalization ability and robustness of the model, and is highly adaptable, and is suitable for image restoration and enhancement in various visual tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259142A_ABST
    Figure CN120259142A_ABST
Patent Text Reader

Abstract

The invention discloses a bad weather degraded image restoration method and device based on dynamic degradation generation, and relates to the field of image processing, and the method comprises the steps: employing an E step of an expectation maximization algorithm to optimize a hidden variable in a training process of a noise predictor and a dynamic degradation generator, and obtaining an optimized hidden variable; in M steps of an expectation maximization algorithm, inputting the optimized hidden variables and state variables into a dynamic degradation generator, generating a degradation layer to perform semi-supervised training on a noise predictor and the dynamic degradation generator, and obtaining a trained noise predictor; in an image restoration stage, an image degraded in severe weather and random Gaussian noise are input into a trained noise predictor, an iterative denoising step obeying a Markov chain is executed, and a clean restored image is obtained. The problems that an existing severe weather image restoration model is limited in generalization ability and unstable in restoration effect are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a method and device for restoring degraded images in bad weather generated based on dynamic degradation. Background Art

[0002] In the field of computer vision, bad weather conditions (such as fog, rain, snow, and dust) will seriously affect the image quality, thereby reducing the accuracy of tasks such as autonomous driving, video surveillance, and target detection. Therefore, it is of great practical significance to study efficient image restoration methods for image degradation problems under different weather conditions.

[0003] Traditional image restoration methods are mainly based on physical models, such as the atmospheric scattering model, the rain streak removal model, etc. Such methods require accurate estimation of physical parameters, such as transmittance, aerosol concentration, etc., but it is often difficult to obtain accurate physical information in complex environments, resulting in unstable restoration effects. In recent years, data-driven deep learning methods have made remarkable progress in the field of image restoration. Among them, supervised learning methods rely on a large amount of paired data for training, but it is difficult to obtain high-quality real paired data, and the model generalization ability is limited. In addition, a single model can usually only handle specific types of weather degradation and is difficult to adapt to multiple bad weather scenarios. Summary of the Invention

[0004] The purpose of this application is to propose a method and device for restoring degraded images in bad weather generated based on dynamic degradation for the above-mentioned technical problems.

[0005] In the first aspect, the present invention provides a method for restoring degraded images in bad weather generated based on dynamic degradation, including the following steps:

[0006] Construct a conditional diffusion model and a dynamic degradation generator. The diffusion model includes a forward noise addition process and a reverse denoising process. In the reverse denoising process, a noise predictor is used for iterative denoising. The dynamic degradation generator includes a transition model, a fully connected layer, and an emission model connected in sequence;

[0007] Construct a labeled dataset and an unlabeled dataset. During the training process of the noise predictor and the dynamic degradation generator, the E-step of the expectation maximization algorithm is used to optimize the introduced latent variables to obtain optimized latent variables. In the M-step of the expectation maximization algorithm, the optimized latent variables and the introduced state variables are input into the dynamic degradation generator to generate a degradation layer. The labeled dataset and the unlabeled dataset are used in combination with the degradation layer to perform semi-supervised training on the noise predictor and the dynamic degradation generator to obtain a trained noise predictor and a trained dynamic degradation generator; the reverse denoising process of the diffusion model including the trained noise predictor is used as the bad weather degraded image restoration model;

[0008] The degraded image of bad weather to be restored and random noise conforming to the standard Gaussian distribution are simultaneously input into the bad weather degraded image restoration model. Through the trained noise predictor, a denoising process of T diffusion time steps obeying the Markov chain is performed to obtain a clean restored image.

[0009] Preferably, the E-step of the expectation-maximization algorithm is used to optimize the introduced latent variable, and the optimized latent variable is obtained, specifically including:

[0010] In the E-step of the expectation-maximization algorithm, the latent variable is iteratively optimized using Langevin dynamics to obtain the optimized latent variable , and the iterative optimization process is represented by the following formula:

[0011] ;

[0012] where τ represents the τ-th iteration of Langevin dynamics, δ represents the step size factor, ξ (τ) represents the Gaussian noise used in the τ-th iteration of Langevin dynamics, represents the objective function with respect to the latent variable , represents the latent variable optimized after the τ-th iteration of Langevin dynamics, represents the latent variable optimized after the (τ + 1)-th iteration of Langevin dynamics, represents the gradient of the objective function with respect to the latent variable at ;

[0013] Repeat the above iterative optimization process to obtain the latent variable optimized after the last iteration of Langevin dynamics and use it as the optimized latent variable .

[0014] Preferably, the optimized latent variable and the introduced state variable are input into the dynamic degradation generator to generate a degradation layer, specifically including:

[0015] In the i-th dynamic degradation generation process corresponding to the dynamic degradation generator, the (i - 1)-th state variable and the i-th optimized latent variable are introduced, and , represents the normal distribution, represents the identity matrix, represents subject to;

[0016] The (i - 1)-th state variable and the i-th optimized latent variable Input into the transition model to obtain the \(i\)-th state variable , as shown in the following formula:

[0017] ;

[0018] where \(T( )\) represents the function corresponding to the transition model, represents the parameters of the transition model;

[0019] Input the \(i\)-th state variable into the fully connected layer and reshape to obtain the \(i\)-th tensor. Input the \(i\)-th tensor into the emission model to obtain the \(i\)-th degraded layer , as shown in the following formula:

[0020] =E( ; );

[0021] where, represents the \(i\)-th tensor, \(E( )\) represents the function corresponding to the emission model, represents the parameters of the emission model.

[0022] Preferably, a semi-supervised training is performed on the noise predictor and the dynamic degradation generator by using a labeled dataset and an unlabeled dataset and combining the degraded layer to obtain a trained noise predictor and a trained dynamic degradation generator, specifically including:

[0023] During the forward noise addition process of the diffusion model, a fixed Markov chain is adopted. Within \(T\) diffusion time steps, the original clean image is gradually perturbed into a noisy image. The perturbation process follows a preset variance scheduling parameter. At the -th diffusion time step, the state transition is defined by the following formula:

[0024] ;

[0025] where, represents the conditional transition probability distribution at the -th diffusion time step during the forward noise addition process, represents the noisy image at the -th diffusion time step, represents the noisy image at the -th diffusion time step, represents the variance scheduling parameter at the -th diffusion time step, , represents the identity matrix, represents the probability distribution of the original clean image, denote being subject to;

[0026] Using the properties of the Markov chain, the noise image at the -th diffusion time step can be directly sampled from the original clean image as follows: as shown in the following equation:

[0027] ;

[0028] where denotes directly sampling the noise image at the -th diffusion time step under the condition of the original clean image is the forward diffusion conditional probability distribution of the noise image at the -th diffusion time step, , and , denotes the cumulative signal retention rate after diffusion time steps, denotes the signal retention coefficient at the -th diffusion time step, denotes the signal retention coefficient at the -th diffusion time step, , denotes the noise term;

[0029] In the reverse denoising process of the diffusion model, sampling starts from and, conditional on the degraded image y, gradually performs denoising sampling steps to recover the clean image. The reverse denoising process forms a parameterized Markov chain, and its joint probability distribution can be expressed as:

[0030] ;

[0031] where denotes the state sequence from the 0-th diffusion time step to the -th diffusion time step, denotes the conditional probability distribution of obtaining under the degraded image y, denotes the parameters of the noise predictor, denotes the -th diffusion time step of the noise image is the probability distribution, , denotes the normal distribution, denote being subject to, I denotes the identity matrix, denotes obtaining the -th diffusion time step of the noise image under the degraded image y and the The noisy image at the th diffusion time step;

[0032] During the reverse denoising process, the following formula is used to obtain the estimated clean image of the skip-step reconstruction from the noisy image at the th diffusion time step : :

[0033] ;

[0034] where represents the predicted noise obtained by inputting the noisy image at the th diffusion time step into the noise predictor;

[0035] The estimated clean image of the skip-step reconstruction is linearly superimposed with each degradation layer to obtain the estimated degraded image, as shown in the following formula:

[0036] ;

[0037] where represents the estimated degraded image;

[0038] The loss function used during the semi-supervised training of the noise predictor and the dynamic degradation generator is:

[0039] ;

[0040] where represents the L2 norm, represents the parameters of the noise predictor, represents the parameters of the dynamic degradation generator, represents the degraded image, represents the estimated degraded image. If the degraded image comes from a labeled dataset, then ; if the degraded image comes from an unlabeled dataset, then ; represents the function corresponding to the adverse weather degradation image restoration model; represents the function corresponding to the dynamic degradation generator, represents the optimized latent variable, represents the initial state variable, represents the variance of the residual term, represents the Kullback-Leibler divergence, represents obtaining the noisy image at the th diffusion time step under the condition of the given original clean image ​ The posterior probability distribution represents the noisy image at the first diffusion time step under the condition of obtaining the original clean image The conditional probability distribution represents at the th diffusion time step of the noisy image and the original clean image under the condition of obtaining the th diffusion time step of the noisy image The true posterior probability distribution represents at the th diffusion time step of the noisy image under the condition of obtaining the th diffusion time step of the noisy image The conditional probability distribution represents the true clean image in the labeled dataset represents the hyperparameter represents taking the expectation of the probability distribution q( )

[0041] In each round of the training process, the E-step and M-step of the expectation-maximization algorithm are alternately executed once. By minimizing the loss function, the optimized noise predictor and the optimized dynamic degradation generator for each round are obtained.

[0042] Using the optimized noise predictor and the optimized dynamic degradation generator of the current round, the E-step and M-step of the expectation-maximization algorithm are executed in the next round of the training process. Repeat the above training process until convergence to obtain the trained noise predictor and the trained dynamic degradation generator.

[0043] Preferably, the transition model includes a first fully connected layer, a second fully connected layer, and a third fully connected layer connected in sequence, where the first fully connected layer and the second fully connected layer are equipped with ReLU activation functions, and the third fully connected layer is equipped with a Tanh activation function.

[0044] Preferably, the emission model includes a first convolutional layer, a second convolutional layer, an upsampling layer, and a third convolutional layer connected in sequence, where the first convolutional layer, the second convolutional layer, and the third convolutional layer are all equipped with ReLU activation functions.

[0045] In a second aspect, the present invention provides a device for restoring degraded images in bad weather based on dynamic degradation generation, including:

[0046] A model construction module, configured to construct a conditional diffusion model and a dynamic degradation generator, where the diffusion model includes a forward noise addition process and a reverse denoising process, and a noise predictor is used for iterative denoising in the reverse denoising process, and the dynamic degradation generator includes a transition model, a fully connected layer, and an emission model connected in sequence;

[0047] A model training module, configured to construct a labeled dataset and an unlabeled dataset. During the training processes of the noise predictor and the dynamic degradation generator, the E-step of the expectation maximization algorithm is used to optimize the introduced latent variables to obtain optimized latent variables; in the M-step of the expectation maximization algorithm, the optimized latent variables and the introduced state variables are input into the dynamic degradation generator to generate a degradation layer, and the labeled dataset, the unlabeled dataset, and the degradation layer are combined to perform semi-supervised training on the noise predictor and the dynamic degradation generator to obtain a trained noise predictor and a trained dynamic degradation generator; the reverse denoising process of the diffusion model including the trained noise predictor is used as a degraded image restoration model for bad weather;

[0048] A restoration module, configured to obtain a degraded image for bad weather to be restored and random noise conforming to a standard Gaussian distribution and input them into the degraded image restoration model for bad weather at the same time. Through the trained noise predictor, a denoising process of T diffusion time steps obeying a Markov chain is performed to obtain a clean restored image.

[0049] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0050] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.

[0051] In a fifth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] (1) The method for restoring degraded images of bad weather based on dynamic degradation generation proposed by the present invention realizes the efficient optimization of latent variables and the joint estimation of model parameters by designing the degradation layer generation process, constructing a conditional denoising process based on degradation generation, and combining the Monte Carlo-based expectation maximization algorithm, thereby achieving high-quality restoration of degraded images under different bad weather conditions in a semi-supervised learning framework.

[0054] (2) The method for restoring degraded images of bad weather based on dynamic degradation generation proposed by the present invention enables the diffusion model to perform semi-supervised learning through the Monte Carlo-based Expectation-Maximization algorithm, making the model generalize better in real scenarios; in addition, by optimizing the latent variables in the E-step of the Expectation-Maximization algorithm and combining the emission model and the transition model, the characteristics of degraded layers of bad weather such as rain and snow layers can be realistically simulated, improving the quality of image restoration.

[0055] (3) The method for restoring degraded images of bad weather based on dynamic degradation generation proposed by the present invention improves the restoration performance by using the reverse denoising process of the diffusion model, and realizes the semi-supervised learning paradigm of the diffusion model through the Monte Carlo Expectation-Maximization algorithm, thus effectively improving the robustness and generalization ability of visual tasks. It has the characteristics of high robustness, strong adaptability and excellent generalization ability, and is applicable to image restoration and enhancement in various visual tasks. Brief Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0057] Figure 1 It is a schematic flowchart of the method for restoring degraded images of bad weather based on dynamic degradation generation according to the embodiment of the present application;

[0058] Figure 2 It is a schematic diagram of the training process of the noise predictor and the dynamic degradation generator of the method for restoring degraded images of bad weather based on dynamic degradation generation according to the embodiment of the present application;

[0059] Figure 3 It is a schematic diagram of the transition model of the method for restoring degraded images of bad weather based on dynamic degradation generation according to the embodiment of the present application;

[0060] Figure 4 It is a schematic diagram of the emission model of the method for restoring degraded images of bad weather based on dynamic degradation generation according to the embodiment of the present application;

[0061] Figure 5 It is a schematic diagram of the device for restoring degraded images of bad weather based on dynamic degradation generation according to the embodiment of the present application;

[0062] Figure 6 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Detailed Embodiments

[0063] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] Figure 1 A method for restoring a degraded image of bad weather generated based on dynamic degradation provided by an embodiment of the present application is shown, including the following steps:

[0065] S1. Construct a conditional diffusion model and a dynamic degradation generator. The diffusion model includes a forward noise addition process and a reverse denoising process. In the reverse denoising process, a noise predictor is used for iterative denoising. The dynamic degradation generator includes a transition model, a fully connected layer, and an emission model connected in sequence.

[0066] In a specific embodiment, the transition model includes a first fully connected layer, a second fully connected layer, and a third fully connected layer connected in sequence, where the first fully connected layer and the second fully connected layer are provided with ReLU activation functions, and the third fully connected layer is provided with a Tanh activation function.

[0067] In a specific embodiment, the emission model includes a first convolutional layer, a second convolutional layer, an upsampling layer, and a third convolutional layer connected in sequence, where the first convolutional layer, the second convolutional layer, and the third convolutional layer are all provided with ReLU activation functions.

[0068] Specifically, referring to Figure 2 , an embodiment of the present application adopts a conditional diffusion model. The diffusion model includes a forward noise addition process and a reverse denoising process. In the forward noise addition process, Gaussian noise is gradually added to a randomly generated original clean image to obtain a noise image at each diffusion time step; in the reverse denoising process, taking the degraded image as a condition, that is, the degraded image and the noise image at the th diffusion time step are input into the noise predictor to obtain the predicted noise at the th diffusion time step. The noise image at the th diffusion time step is denoised by the predicted noise at the th diffusion time step to restore the noise image at the th diffusion time step. The above steps are repeated to gradually remove the noise and obtain the original clean image.

[0069] The dynamic degradation generator is used in the training process of the noise predictor in the diffusion model, and the parameters of the dynamic degradation generator are also optimized synchronously during the training process of the noise predictor in the diffusion model. The dynamic degradation generator is composed of a transition model, a fully connected layer, and an emission model. Referring toFigure 3 and Figure 4 For the transfer model, a multi-layer perceptron (MLP) is used for modeling. In one embodiment, the transfer model consists of a first fully-connected layer, a second fully-connected layer, and a third fully-connected layer connected in sequence. Moreover, the first fully-connected layer and the second fully-connected layer are equipped with ReLU activation functions, and the third fully-connected layer is equipped with a Tanh activation function. The emission model consists of a first convolutional layer, a second convolutional layer, an upsampling layer, and a third convolutional layer connected in sequence, and each convolutional layer is equipped with a ReLU activation function.

[0070] S2. Construct a labeled dataset and an unlabeled dataset. During the training process of the noise predictor and the dynamic degradation generator, the E-step of the expectation-maximization algorithm is used to optimize the introduced latent variables to obtain optimized latent variables. In the M-step of the expectation-maximization algorithm, the optimized latent variables and the introduced state variables are input into the dynamic degradation generator to generate a degradation layer. The labeled dataset, the unlabeled dataset, and the degradation layer are combined to perform semi-supervised training on the noise predictor and the dynamic degradation generator to obtain a trained noise predictor and a trained dynamic degradation generator. The reverse denoising process of the diffusion model is used as a model for restoring images degraded by bad weather.

[0071] In a specific embodiment, using the E-step of the expectation-maximization algorithm to optimize the introduced latent variables to obtain optimized latent variables specifically includes:

[0072] In the E-step of the expectation-maximization algorithm, the latent variables are iteratively optimized using Langevin dynamics to obtain optimized latent variables , and the iterative optimization process is represented by the following formula:

[0073] ;

[0074] where τ represents the τ-th iteration of Langevin dynamics, δ represents the step size factor, ξ (τ) represents the Gaussian noise used in the τ-th iteration of Langevin dynamics, represents the gradient of the objective function with respect to the latent variables , represents the latent variables optimized after the τ-th iteration of Langevin dynamics, represents the latent variables optimized after the (τ + 1)-th iteration of Langevin dynamics, represents the gradient of the objective function with respect to the latent variables at ;

[0075] Repeat the above iterative optimization process to obtain the latent variables optimized after the last iteration of Langevin dynamics and use them as the optimized latent variables 。

[0076] In a specific embodiment, the optimized latent variable and the introduced state variable are input into the dynamic degradation generator to generate a degradation layer, specifically including:

[0077] Introduce the (i - 1)-th state variable in the i-th dynamic degradation generation process corresponding to the dynamic degradation generator and the i-th optimized latent variable , where and , represents a normal distribution, represents an identity matrix, represents subject to;

[0078] Input the (i - 1)-th state variable and the i-th optimized latent variable into the transition model to obtain the i-th state variable , as shown in the following formula:

[0079] ;

[0080] where T( ) represents the function corresponding to the transition model, represents the parameter of the transition model;

[0081] Input the i-th state variable into the fully connected layer and reshape it to obtain the i-th tensor, and input the i-th tensor into the emission model to obtain the i-th degradation layer , as shown in the following formula:

[0082] = E( ; );

[0083] where represents the i-th tensor, E( ) represents the function corresponding to the emission model, represents the parameter of the emission model.

[0084] In a specific embodiment, a labeled dataset and an unlabeled dataset are used in combination with the degradation layer to perform semi-supervised training on the noise predictor and the dynamic degradation generator, and a trained noise predictor and a trained dynamic degradation generator are obtained, specifically including:

[0085] In the forward noise addition process of the diffusion model, a fixed Markov chain is used. Within T diffusion time steps, the original clean image is gradually perturbed into a noisy image, and the perturbation process follows a preset variance scheduling parameter. At the In a diffusion time step, the state transition is defined by the following formula:

[0086] ;

[0087] where represents the conditional transition probability distribution at the -th diffusion time step in the forward noise addition process, represents the noisy image at the -th diffusion time step, represents the noisy image at the -th diffusion time step, represents the variance scheduling parameter at the -th diffusion time step, , represents the identity matrix, represents the probability distribution of the original clean image, represents being subject to;

[0088] Using the properties of the Markov chain, the noisy image at the -th diffusion time step can be directly sampled from the original clean image , as shown in the following formula:

[0089] ;

[0090] where represents the forward diffusion conditional probability distribution of directly sampling the noisy image at the -th diffusion time step under the condition of the original clean image , , and , represents the cumulative signal retention rate after diffusion time steps, represents the signal retention coefficient at the -th diffusion time step, represents the signal retention coefficient at the -th diffusion time step, , represents the noise term;

[0091] In the reverse denoising process of the diffusion model, sampling starts from , and with the degraded image y as the condition, the denoising sampling steps are gradually executed to recover the clean image. The reverse denoising process constitutes a parameterized Markov chain, and its joint probability distribution can be expressed as:

[0092] ;

[0093] Among them, represents the state sequence from the 0th diffusion time step to the th diffusion time step, represents the conditional probability distribution obtained under the degraded image y . represents the parameters of the noise predictor, represents the th diffusion time step of the noise image 's probability distribution, , represents the normal distribution, represents being subject to, I represents the identity matrix, represents under the degraded image y and the th diffusion time step of the noise image to obtain the th diffusion time step of the noise image 's conditional probability distribution;

[0094] During the reverse denoising process, the following formula is used to obtain the estimated clean image of the skip reconstruction from the th diffusion time step of the noise image :

[0095] ;

[0096] Among them, represents the predicted noise obtained by inputting the th diffusion time step of the noise image into the noise predictor;

[0097] The estimated clean image of the skip reconstruction is linearly superimposed with each degraded layer to obtain the estimated degraded image, as shown in the following formula:

[0098] ;

[0099] Among them, represents the estimated degraded image;

[0100] The loss function used in the semi-supervised training process of the noise predictor and the dynamic degradation generator

[0101] ;

[0102] Among them, represents the L2 norm, represents the parameters of the noise predictor, Denote the parameters of the dynamic degradation generator, Denote the degraded image, Denote the estimated degraded image. If the degraded image comes from a labeled dataset, then ; If the degraded image comes from an unlabeled dataset, then ; Denote the function corresponding to the adverse weather degraded image restoration model; Denote the function corresponding to the dynamic degradation generator, Denote the optimized latent variable, Denote the initial state variable, Denote the variance of the residual term, Denote the Kullback-Leibler divergence, Denote obtaining the noise image at the th diffusion time step under the condition of the given original clean image posterior probability distribution, Denote the conditional probability distribution of obtaining the original clean image under the condition of the noise image at the first diffusion time step, Denote the true posterior probability distribution of obtaining the noise image at the th diffusion time step under the conditions of the noise image at the th diffusion time step and the original clean image ; Denote the conditional probability distribution of obtaining the noise image at the th diffusion time step under the condition of the noise image at the th diffusion time step, Denote the true clean image in the labeled dataset, Denote the hyperparameter, Denote taking the expectation of the probability distribution q( );

[0103] In each round of training process, alternately execute one E-step and one M-step of the expectation-maximization algorithm. By minimizing the loss function, obtain the optimized noise predictor and the optimized dynamic degradation generator for each round;

[0104] Use the optimized noise predictor and the optimized dynamic degradation generator in the current round to execute one E-step and one M-step of the expectation-maximization algorithm in the next round of training process. Repeat the above training process until convergence to obtain the trained noise predictor and the trained dynamic degradation generator.

[0105] Specifically, during the training process of the noise predictor and the dynamic degradation generator in the diffusion model, the E-step of the Expectation-Maximization (EM) algorithm is first adopted. The latent variable z is iteratively optimized through Langevin Dynamics so that the latent variable gradually approaches the true data distribution during the update process. Gaussian noise is introduced into the optimization iteration formula of the latent variable z to prevent the optimization iteration process from falling into a local optimal solution. After obtaining the optimized latent variable, the introduced state variable and the optimized latent variable are used in each dynamic degradation generation process corresponding to the dynamic degradation generator. After multiple dynamic degradation generation processes, the degraded layers generated in each dynamic degradation generation process are obtained. By linearly superimposing the degraded layers generated in each dynamic degradation generation process and the estimated clean image reconstructed by skipping steps a corresponding estimated degraded image can be obtained.

[0106] Furthermore, in order to reflect the matching between the estimated degraded image and the true degraded image y, the Mean Squared Error (MSE) loss function is introduced:[[]]

[0107] 2;

[0108] where represents the Mean Squared Error loss function;

[0109] In the M-step of the Expectation-Maximization algorithm, the maximum likelihood estimation of the model parameters is performed from the probability model containing the latent variable to further optimize the overall performance of the model. Specifically, by maximizing the following log-likelihood function:[[]]

[0110] ;

[0111] where represents the joint posterior probability distribution of the parameters of the noise predictor and the parameters of the dynamic degradation generator under the degraded image y, represents the likelihood function of generating the degraded image y under the parameters of the noise predictor and the parameters of the dynamic degradation generator, represents the prior probability distribution of the parameters of the noise predictor, represents a constant term, represents that the left expression is equivalently defined as the right symbol, represents the objective function to be optimized, represents the parameters of the dynamic degradation generator, ~ , denotes being subject to, and I denotes the identity matrix.

[0112] For the parameters of the noise predictor included in the diffusion model , embodiments of the present application use the variational lower bound of the negative log-likelihood for training, and its corresponding KL divergence loss function is defined as follows:

[0113] ;

[0114] where denotes the KL divergence loss function. It uses the KL divergence to measure the distance between two probability distributions.

[0115] Furthermore, a labeled dataset and an unlabeled dataset are constructed. In the labeled dataset, each sample has a corresponding real clear image for the degraded image ; while in the unlabeled dataset, each sample does not have a corresponding real clear image for the degraded image ; therefore, a semi-supervised loss function is introduced for the cases of the labeled dataset and the unlabeled dataset; in each training epoch, the model alternately processes the labeled dataset and the unlabeled dataset. The training of the labeled dataset and the unlabeled dataset is carried out simultaneously and alternately.

[0116] Combining the above four loss functions or objective functions, the loss function is derived. By minimizing this loss function , for the parameters of the noise predictor and the parameters of the dynamic degradation generator, optimization and adjustment are performed. In the above training process, the E-step and M-step of the expectation-maximization algorithm are carried out synchronously, and repeated iteration is performed until convergence to obtain a trained noise predictor and a trained dynamic degradation generator. By minimizing this loss function, the model can gradually approximate the distribution of the real data during the optimization process and ensure higher stability and accuracy in the complex reverse denoising process. The pseudo-code of its specific training process is shown in Table 1.

[0117] Table 1:

[0118]

[0119] In the training process of the noise predictor in the diffusion model of the embodiments of the present application, not only is a dynamic degradation generator introduced to guide and synchronously optimize the training process, but also the expectation-maximization algorithm is nested in the training process, so that the restored model of the bad weather degraded image obtained by training (i.e., the reverse denoising process of the diffusion model including the trained noise predictor) can restore bad weather degraded images in various different situations.

[0120] S3. Obtain the degraded image in bad weather to be restored and random noise conforming to the standard Gaussian distribution, and input them into the degraded image restoration model in bad weather at the same time. Through the trained noise predictor, perform the denoising process for T diffusion time steps that follow the Markov chain to obtain a clean restored image.

[0121] Specifically, regard the reverse denoising process containing the trained noise predictor in the diffusion model as the degraded image restoration model in bad weather. Input the degraded image in bad weather to be restored and the obtained random noise vector conforming to the standard Gaussian distribution into the degraded image restoration model in bad weather. Use the degraded image in bad weather to be restored as a condition, and use the trained noise predictor to predict the predicted noise at each diffusion time step, and gradually denoise the degraded image in bad weather to be restored to obtain the corresponding restored clean image.

[0122] The method for restoring the degraded image in bad weather based on dynamic degradation proposed in the embodiments of this application introduces the semi-supervised learning method into the image restoration task. Semi-supervised learning can use a small amount of labeled data and a large amount of unlabeled data for training, reduce the dependence on large-scale labeled data, and improve the generalization ability of the model. In addition, by designing a unified framework, the image restoration tasks under different weather conditions are integrated into one model, so as to realize the integrated restoration of various degraded images in different bad weathers, and can better meet the actual application requirements.

[0123] The technical effects of the embodiments of this application will be described below through specific experiments.

[0124] The situation of the experimental platform used in the embodiments of this application is as follows:

[0125] 1) Hardware:

[0126] GPU: NVIDIA GeForce RTX 3090;

[0127] CPU: Intel Xeon(R)CPU E5-2697A v4@ 2.60 GHz×64;

[0128] Memory: 256 GB RAM;

[0129] 2) Software:

[0130] Operating system: Ubuntu 20.04.6 LTS;

[0131] Framework: PyTorch 2.1.0 + CUDA 12.1, Python 3.8;

[0132] Dependencies: NumPy, OpenCV, SciPy, torchvision;

[0133] The specific parameter settings are shown in Table 2.

[0134] Table 2:

[0135]

[0136] The performance of the degraded image restoration model for bad weather proposed in the embodiments of this application is compared with the image restoration models of other existing fully supervised methods (including Transweather, AirNet, Gridformer*, and PromptIR*). The comparison results are shown in Table 3. According to the comparison results, it can be seen that:

[0137] 1. The degraded image restoration model for bad weather proposed in the embodiments of this application is significantly better than Transweather in terms of PSNR, but there is still a gap compared with fully supervised methods (such as AirNet and Gridformer*). This gap is in line with the expected performance of semi-supervised learning. Especially in pixel-level reconstruction tasks, fully supervised methods have an advantage due to directly using label information.

[0138] 2. The difference in PSNR between the degraded image restoration model (30.95) proposed in the embodiments of this application and the optimal fully supervised method (Gridformer*, 36.61) is about 5.66 dB. The difference in structural similarity SSIM is 0.069 (0.902 vs. 0.971). This gap is within the typical range of semi-supervised learning and is an acceptable result, especially when data annotation is limited.

[0139] 3. Semi-supervised methods are weaker than fully supervised methods in terms of PSNR / SSIM, mainly due to insufficient supervision signals, but they have significant advantages in annotation cost and generalization ability. If the test data contains unknown degradation types or low-quality inputs, the performance degradation of semi-supervised methods may be less than that of fully supervised methods.

[0140] Table 3:

[0141]

[0142] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, this application provides an embodiment of a degraded image restoration device for bad weather based on dynamic degradation generation. This device embodiment corresponds to the method embodiment shown in Figure 1 and can be specifically applied to various electronic devices.

[0143] An embodiment of the present application provides a device for restoring degraded images in bad weather generated based on dynamic degradation, including:

[0144] A model construction module 1, configured to construct a conditional diffusion model and a dynamic degradation generator. The diffusion model includes a forward noise addition process and a reverse denoising process. In the reverse denoising process, a noise predictor is used for iterative denoising. The dynamic degradation generator includes a transition model, a fully connected layer, and an emission model connected in sequence;

[0145] A model training module 2, configured to construct a labeled dataset and an unlabeled dataset. During the training process of the noise predictor and the dynamic degradation generator, the E-step of the expectation maximization algorithm is used to optimize the introduced latent variables to obtain optimized latent variables; in the M-step of the expectation maximization algorithm, the optimized latent variables and the introduced state variables are input into the dynamic degradation generator to generate a degradation layer. The labeled dataset, the unlabeled dataset, and the degradation layer are combined to perform semi-supervised training on the noise predictor and the dynamic degradation generator to obtain a trained noise predictor and a trained dynamic degradation generator; the reverse denoising process of the diffusion model including the trained noise predictor is used as the bad weather degraded image restoration model;

[0146] A restoration module 3, configured to obtain a bad weather degraded image to be restored and random noise conforming to a standard Gaussian distribution and input them into the bad weather degraded image restoration model at the same time. Through the trained noise predictor, a denoising process of T diffusion time steps obeying a Markov chain is performed to obtain a clean restored image.

[0147] Figure 6 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. As Figure 6 shown, the electronic device of this embodiment includes: a processor 601 and a memory 602; wherein the memory 602 is used to store computer execution instructions; the processor 601 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. For specific reference, see the relevant descriptions in the foregoing method embodiments.

[0148] Optionally, the memory 602 can be either independent or integrated with the processor 601.

[0149] When the memory 602 is independently set, the electronic device further includes a bus 603 for connecting the memory 602 and the processor 601.

[0150] An embodiment of the present invention also provides a computer storage medium, in which computer execution instructions are stored. When the processor 601 executes the computer execution instructions, the above method is implemented.

[0151] An embodiment of the present invention also provides a computer program product, including a computer program, which when executed by a processor 601, implements the above method.

[0152] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.

[0153] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0154] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a hardware plus software functional unit.

[0155] The above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 601 to execute some steps of the methods in various embodiments of the present application.

[0156] It should be understood that the above processor 601 can be a central processing unit (Central Processing Unit, abbreviated as CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor 601 can also be any conventional processor 601, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by the hardware processor 601, or executed by a combination of hardware and software modules in the processor 601.

[0157] The memory 602 may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk or an optical disc, etc.

[0158] The bus 603 may be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus 603 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus 603 in the attached drawings of this application is not limited to only one bus 603 or one type of bus 603.

[0159] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0160] An exemplary storage medium is coupled to the processor 601, so that the processor 601 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 601. The processor 601 and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor 601 and the storage medium can also exist as discrete components in an electronic device or a main control device.

[0161] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disk or optical disc and other media that can store program codes.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for restoring degraded images of severe weather generated based on dynamic degradation, characterized in that, The steps include: Construct a conditional diffusion model and a dynamic degradation generator. The diffusion model includes a forward noise-adding process and a reverse denoising process. In the reverse denoising process, a noise predictor is used for iterative denoising. The dynamic degradation generator includes a transition model, a fully connected layer, and an emission model connected in sequence; Construct a labeled dataset and an unlabeled dataset. During the training process of the noise predictor and the dynamic degradation generator, the E-step of the expectation-maximization algorithm is used to optimize the introduced latent variables to obtain optimized latent variables. In the M-step of the expectation-maximization algorithm, the optimized latent variables and the introduced state variables are input into the dynamic degradation generator to generate a degraded layer. The labeled dataset and the unlabeled dataset are used in combination with the degraded layer to perform semi-supervised training on the noise predictor and the dynamic degradation generator to obtain a trained noise predictor and a trained dynamic degradation generator; Use the reverse denoising process of the diffusion model as a model for restoring images degraded by bad weather; Obtain a bad-weather degraded image to be restored and random noise conforming to a standard Gaussian distribution and input them simultaneously into the model for restoring images degraded by bad weather. Through the trained noise predictor, perform a denoising process with T diffusion time steps following a Markov chain to obtain a clean restored image.

2. The method for restoring a degraded image of bad weather generated based on dynamic degradation according to claim 1, wherein Use the E-step of the expectation-maximization algorithm to optimize the introduced latent variables to obtain optimized latent variables, specifically including: In the E-step of the expectation-maximization algorithm, the Langevin dynamics is used to iteratively optimize the latent variable to obtain the optimized latent variable . The iterative optimization process is represented by the following formula: ; where τ represents the τ-th iteration of Langevin dynamics, δ represents the step size factor, and ξ (τ) represents the Gaussian noise used in the τ-th iteration of Langevin dynamics, represents the objective function with respect to the latent variable ; represents the latent variable optimized after the τ-th iteration of Langevin dynamics, represents the latent variable optimized after the (τ + 1)-th iteration of Langevin dynamics, represents the gradient of the objective function with respect to the latent variable at ; Repeat the above iterative optimization process, and use the latent variable obtained in the last iteration as the optimized latent variable .

3. The method for restoring a degraded image of bad weather generated based on dynamic degradation according to claim 1, wherein, Input the optimized latent variables and the introduced state variables into the dynamic degradation generator to generate a degraded layer, specifically including: Introduce the (i-1)th state variable and the ith optimized latent variable during the ith dynamic degradation generation process corresponding to the dynamic degradation generator and , and , denotes a normal distribution, denotes an identity matrix, denotes being subject to; Input the (i - 1)-th state variable and the optimized i-th latent variable into the transition model to obtain the i-th state variable , as shown in the following formula: ; Among them, T( ) represents the function corresponding to the transfer model, represents the parameters of the transfer model; Input the i-th state variable into the fully connected layer and reshape it to obtain the i-th tensor, and input the i-th tensor into the emission model to obtain the i-th degraded layer , as shown in the following formula: =E( ; ); Among them, represents the i-th tensor, and E( ) represents the function corresponding to the emission model, represents the parameters of the emission model.

4. The method for restoring a degraded image of bad weather generated based on dynamic degradation according to claim 1, characterized in that Use the labeled dataset and the unlabeled dataset in combination with the degraded layer to perform semi-supervised training on the noise predictor and the dynamic degradation generator to obtain a trained noise predictor and a trained dynamic degradation generator, specifically including: During the forward noise addition process of the diffusion model, a fixed Markov chain is adopted. Within T diffusion time steps, the original clean image is gradually perturbed into a noisy image. The perturbation process follows a preset variance scheduling parameter. At the th diffusion time step, the state transition is defined by the following formula: ; Among them, represents the conditional transition probability distribution at the -th diffusion time step during the forward noise addition process, represents the noisy image at the -th diffusion time step, represents the noisy image at the -th diffusion time step, represents the variance scheduling parameter at the -th diffusion time step, , represents the identity matrix, represents the probability distribution of the original clean image, means subject to; Using the properties of Markov chains, the noise image at the th diffusion time step can be directly sampled from the original clean image , as shown in the following equation: ; Among them, represents the noise image obtained by directly sampling under the condition of the original clean image at the -th diffusion time step, which is the forward diffusion conditional probability distribution, , and represents the cumulative signal retention rate after diffusion time steps, represents the signal retention coefficient at the -th diffusion time step, represents the signal retention coefficient at the -th diffusion time step, represents the noise term; During the reverse denoising process of the diffusion model, starting from sampling is performed, and conditional on the degraded image y, the denoising sampling steps are gradually executed to recover the clean image. The reverse denoising process forms a parameterized Markov chain, and its joint probability distribution can be expressed as: ; Among them, represents the state sequence from the 0th diffusion time step to the th diffusion time step, represents the conditional probability distribution obtained under the degraded image y ; represents the parameters of the noise predictor, represents the th diffusion time step of the noise image 's probability distribution, , represents the normal distribution, represents being subject to, I represents the identity matrix, represents under the degraded image y and the th diffusion time step of the noise image to obtain the th diffusion time step of the noise image 's conditional probability distribution; During the reverse denoising process, the following formula is used to obtain the estimated clean image for skip-step reconstruction from the noisy image at the -th diffusion time step : ​ ; Among them, represents the predicted noise obtained by inputting the noise image at the th diffusion time step into the noise predictor; ​ The estimated clean image after skip-step reconstruction is linearly superimposed with each degradation layer to obtain an estimated degraded image, as shown in the following formula: ; Among them, represents the estimated degraded image; The loss function used in the semi-supervised training process of the noise predictor and the dynamic degradation generator is as follows: ; Among them, represents the L2 norm, represents the parameters of the noise predictor, represents the parameters of the dynamic degradation generator, represents the degraded image, represents the estimated degraded image. If the degraded image comes from the labeled dataset, then ; if the degraded image comes from the unlabeled dataset, then ; represents the function corresponding to the adverse weather degradation image restoration model; represents the function corresponding to the dynamic degradation generator, represents the optimized latent variable, represents the initial state variable, represents the variance of the residual term, represents the Kullback-Leibler divergence, represents obtaining the noise image at the th diffusion time step under the condition of the given original clean image posterior probability distribution, represents the conditional probability distribution of obtaining the original clean image at the first diffusion time step given the noise image , represents the true posterior probability distribution of obtaining the noise image at the th diffusion time step given the noise image at the th diffusion time step and the original clean image , represents the conditional probability distribution of obtaining the noise image at the th diffusion time step given the noise image at the th diffusion time step, represents the true clean image in the labeled dataset, represents the hyperparameter, represents taking the expectation of the probability distribution q( ); In each round of training process, alternately execute the E-step and the M-step of the expectation-maximization algorithm once. By minimizing the loss function, obtain the optimized noise predictor and the optimized dynamic degradation generator for each round; Use the optimized noise predictor and the optimized dynamic degradation generator in the current round to execute the E-step and the M-step of the expectation-maximization algorithm in the next round of training process. Repeat the above training process until convergence to obtain a trained noise predictor and a trained dynamic degradation generator.

5. The method for restoring a degraded image of bad weather generated based on dynamic degradation according to claim 1, wherein, The transition model includes a first fully connected layer, a second fully connected layer, and a third fully connected layer connected in sequence, where the first fully connected layer and the second fully connected layer have ReLU activation functions, and the third fully connected layer has a Tanh activation function.

6. The method for restoring a deteriorated weather degraded image generated based on dynamic degradation according to claim 1, wherein The emission model includes a first convolutional layer, a second convolutional layer, an upsampling layer, and a third convolutional layer connected in sequence, where the first convolutional layer, the second convolutional layer, and the third convolutional layer all have ReLU activation functions.

7. An apparatus for restoring degraded images of severe weather generated based on dynamic degradation, characterized in that, It includes: A model construction module, configured to construct a conditional diffusion model and a dynamic degradation generator, where the diffusion model includes a forward noise addition process and a reverse denoising process, and a noise predictor is used for iterative denoising in the reverse denoising process, and the dynamic degradation generator includes a transition model, a fully connected layer, and an emission model connected in sequence; A model training module, configured to construct a labeled data set and an unlabeled data set. During the training process of the noise predictor and the dynamic degradation generator, the E-step of the expectation maximization algorithm is used to optimize the introduced latent variables to obtain optimized latent variables; in the M-step of the expectation maximization algorithm, the optimized latent variables and the introduced state variables are input into the dynamic degradation generator to generate a degraded layer, and the labeled data set and the unlabeled data set are used in combination with the degraded layer to perform semi-supervised training on the noise predictor and the dynamic degradation generator to obtain a trained noise predictor and a trained dynamic degradation generator; The reverse denoising process of the diffusion model including the trained noise predictor is used as a bad weather degraded image restoration model; A restoration module, which obtains a bad weather degraded image to be restored and a random noise conforming to a standard Gaussian distribution and inputs them into the bad weather degraded image restoration model at the same time. Through the trained noise predictor, a denoising process of T diffusion time steps obeying a Markov chain is performed to obtain a clean restored image.

8. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-6 is implemented.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Degraded image restoration method and system

    CN102646267A

  • Tone modification voice restoration method and system based on deep learning model

    CN117612544A

  • General image restoration method for severe weather based on two-stage distillation learning

    CN118521511A

  • Mining area multi-severe weather image enhancement method and system

    CN118537251A

  • Image restoration method and system suitable for various severe weathers

    CN119180766A

Cited By

  • Energy spectrum CT image joint denoising and reconstruction method based on degradation perception diffusion model

    CN121391657A

  • A spectral CT image joint denoising and reconstruction method based on a degenerate perception diffusion model

    CN121391657B