Image defogging method and system based on physical guide diffusion model
By decomposing the fog map using a physical perception dehazing model and combining it with the Markov chain mechanism of the diffusion model, and using pseudo-sharp images and transmittance maps as conditional guidance, the problem of unclear dehazing process in existing technologies is solved, thereby improving image clarity and visual quality.
Patent Information
- Application Number
- CN202511143552.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-28
AI Technical Summary
Existing diffusion models lack modeling of the physical mechanisms of haze formation in real-world scenes during image dehazing, resulting in an unclear denoising process that affects image clarity and quality.
A physical perception dehazing model is used to decompose fog images into pseudo-sharp images, transmittance images, and atmospheric light images. Combined with a diffusion model, noise is gradually reduced through a Markov chain mechanism. The pseudo-sharp images and transmittance images are used as conditions to guide the optimization of model parameters through a self-circulating consistency loss.
It improves the clarity and visual quality of the dehazed image, enhances the model's generalization ability in real and complex scenes, and ensures the fidelity of image details and structure.
Smart Images

Figure CN121032845A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an image dehazing method and system based on a physically guided diffusion model. Background Technology
[0002] In outdoor vision systems, haze is a common atmospheric phenomenon that significantly reduces image visibility and contrast, leading to a severe deterioration in image quality. These images typically appear dim and blurry, accompanied by loss of detail and increased noise, severely impacting the performance of advanced vision tasks such as autonomous driving, surveillance, and remote sensing. Therefore, effectively enhancing hazy images to improve their visual quality is a crucial step in ensuring the accurate and efficient execution of advanced vision tasks.
[0003] In existing technologies, diffusion models have gained significant attention in image generation due to their stable training process and ability to avoid mode collapse. First, the diffusion model analyzes the noise distribution characteristics in foggy images, accurately modeling the noise to identify its components. Then, using a back-diffusion process, it gradually guides the image from a noisy distribution to a clearer one, predicting and removing noise at each step through a neural network to achieve progressive dehazing. During this transformation, a conditional guidance mechanism is incorporated, utilizing local texture details and global structural information to obtain the dehazed image.
[0004] The drawback of the aforementioned existing technology is that the image dehazing algorithm based on the diffusion model lacks a clear model of the physical mechanism of haze formation in real scenes when adding and removing noise from hazy images. This results in an unclear sampling direction during the denoising process, which in turn affects the clarity and quality of the enhanced image. Summary of the Invention
[0005] Therefore, it is necessary to provide an image dehazing method based on a physically guided diffusion model to address the aforementioned technical problems.
[0006] The technical solution adopted in this invention is as follows:
[0007] In a first aspect, this invention proposes an image dehazing method based on a physically guided diffusion model, comprising the following steps:
[0008] (1) Obtain real-world fog and clear images;
[0009] (2) Input the fog image into the trained physical perception defogging model and output the pseudo-sharp image J* and the transmittance image T, wherein the pseudo-sharp image J* is the preliminary estimated fog-free image and the transmittance image represents the degree of attenuation of scene light by the medium.
[0010] The physical perception dehazing model includes a pseudo-sharp image estimation sub-network, a transmittance image estimation sub-network, and an atmospheric light image estimation sub-network, which are used to generate the pseudo-sharp image J*, the transmittance image T, and the atmospheric light image A, respectively; among them, the atmospheric light image estimation sub-network is only used in the training process of the physical perception dehazing model.
[0011] (3) Input the clear image into the diffusion model, gradually add noise to generate a noise image through the Markov chain mechanism, and use the conditional guidance mechanism to inject the pseudo clear image J* and the transmittance image T into the diffusion model by feature stitching to guide the denoising direction and gradually predict and remove noise components to generate an intermediate image. The intermediate image at the final time step is output as the clear image after dehazing. At the same time, minimize the error between the predicted noise and the real noise, as well as the error between the clear image and the clear image after dehazing, to train the diffusion model.
[0012] (4) Image dehazing is performed using a physical-guided diffusion model consisting of a trained physical perception dehazing model and a diffusion model.
[0013] Furthermore, the sharp image estimation subnetwork adopts a symmetric encoder-decoder structure and integrates a multi-scale feature fusion module between the encoder and the decoder. The encoder extracts hierarchical features through downsampling, the decoder restores resolution through upsampling, and the decoder ends by converting the feature map output by the last layer of the decoder into a three-channel image through a 1×1 convolutional layer, and outputs a normalized pseudo-sharp image after passing through the Tanh activation function.
[0014] Furthermore, the computation process of the multi-scale feature fusion module includes:
[0015] First, multi-scale features from the last m+1 layer of the encoder are extracted using parallel dilated convolution kernels and then initially fused.
[0016]
[0017] Among them, Conv dil=s This represents a convolution operation with an inflation rate of s. Let || denote the feature representation of the i-th layer at scale s, and || denote the channel concatenation operation. Conv 1×1 F represents a 1×1 convolution. fuse This indicates the initial fusion characteristics, where m ≤ n-1;
[0018] Then, the initial fusion features are dynamically adjusted using a channel attention mechanism:
[0019] w = σ(MLP(GAP(F) fuse )))
[0020]
[0021] Where GAP represents global average pooling, MLP represents a fully connected layer, σ represents the sigmoid activation function, and w represents the generated channel attention weight vector. F represents channel-by-channel multiplication. att This represents the dynamically adjusted fusion features, which serve as the output of the multi-scale feature fusion module.
[0022] Furthermore, the output of the multi-scale feature fusion module, after being compressed by a 3×3 convolution, is added to the features of the same layer of the decoder via skip connections and then used as the input to the next layer of the decoder, as follows:
[0023] F out,i =Deconv(Conv 3×3 (F att ))+F dec,i
[0024] Among them, F att Conv represents the output of the multi-scale feature fusion module. 3×3 Indicates a 3×3 convolutional layer, F dec,i F represents the feature map output of the i-th level of the decoder. out,i This represents the output feature map after resolution restoration and feature fusion, which serves as the input to the (i+1)th level of the decoder.
[0025] Furthermore, the transmittance map estimation subnetwork adopts a multi-layer non-degenerate structure. Each non-degenerate structure contains a convolutional layer, a batch normalization layer, and a LeakyReLU activation function. The output of the last layer generates a transmittance map T through a convolutional layer and a sigmoid function.
[0026] Furthermore, the atmospheric light map estimation subnetwork adopts a symmetric encoder-decoder structure. The encoder extracts hierarchical features through downsampling, the decoder restores the resolution through upsampling, and the decoder generates the atmospheric light map at the end through transposed convolutional layers and average pooling.
[0027] Furthermore, the process of training the physical perception defogging model includes:
[0028] Using real-world fog images as input to the physical perception dehazing model, estimated pseudo-clear images J*, transmittance maps T, and atmospheric light maps A are generated. The physical perception fog image I is then reconstructed using the following formula:
[0029] I(x)=J * (x)t(x)+A(1-t(x))
[0030] Where I(x) represents the value of pixel x in the reconstructed fog image, J *(x) represents the value of pixel x in the estimated pseudo-clear image, A represents the value of pixel x in the estimated atmospheric light image, and t(x) represents the value of pixel x in the estimated transmittance image.
[0031] We employ a self-circulating consistency loss method and optimize model parameters by comparing real-world fog maps with reconstructed fog maps.
[0032] Furthermore, the denoising process in step (3) is expressed as follows:
[0033]
[0034] Where, x t The fog map represents the current time step t, x t-1 α represents the image output at the previous time step t-1 after denoising. t and β t These are noise scheduling parameters, ∈ θ (·) represents the noise prediction network, t is the current denoising time step, and σ t This represents the standard deviation of noise associated with the time step, where z is the standard Gaussian noise, and Concat[J] * [,T] is the feature stitching operation, which stitches the pseudo-sharp image J... * The transmittance map T is combined with the transmittance map T as a conditional input for noise prediction.
[0035] Secondly, this invention proposes an image dehazing system based on a physically guided diffusion model, which is used to implement the aforementioned image dehazing method based on a physically guided diffusion model.
[0036] Thirdly, the present invention provides a computer electronic device, including a memory and a processor;
[0037] The memory is used to store computer programs;
[0038] The processor is configured to implement the above-described image dehazing method based on the physical guided diffusion model when executing the computer program.
[0039] The image dehazing method based on a physically guided diffusion model provided in this invention has the following advantages compared with the prior art:
[0040] This invention utilizes a physical perception dehazing model to decompose fog images, generating pseudo-sharp images (J). * J and the transmission map (T) serve as physical guiding conditions for the diffusion process, where J *The model provides scene structure constraints to ensure detail fidelity, while T models the spatial distribution characteristics of haze. The two are injected into the diffusion model through feature concatenation. Their synergistic effect provides clear physical guidance for the denoising process, improving the clarity and visual quality of the dehazed image. At the same time, a self-circulating consistency loss is used to replace the traditional adversarial loss. The physical perception haze map is reconstructed through the atmospheric scattering model and compared with the input haze map for consistency, solving the problem of unsupervised training and thus enhancing the model's generalization ability in real complex scenes. Attached Figure Description
[0041] Figure 1 This is a schematic diagram of a physical perception dehazing model for an image dehazing method based on a physical guided diffusion model, provided in one embodiment.
[0042] Figure 2 This is a flowchart illustrating an image dehazing method based on a physically guided diffusion model, as provided in one embodiment.
[0043] Figure 3 The image dehazing method based on a physical guided diffusion model is shown as a peak signal-to-noise ratio curve in one embodiment. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0045] In one embodiment, an image dehazing method based on a physically guided diffusion model is provided, the method comprising:
[0046] Obtain real-world fog and clear images;
[0047] A pseudo-sharp image estimation subnetwork (J-Net) is designed using a symmetric encoder-decoder architecture. After receiving the fog image, the encoder extracts hierarchical features through four downsampling layers, and the decoder restores the resolution through corresponding upsampling. A multi-scale feature fusion module (MSF) is integrated, which fuses features from layers n, n-1, and n-2 of the encoder through cross-scale dilated convolutions (dilation rates 1 / 3 / 5). After compression by a 3×3 convolution, these features are added to the features from the same layer of the decoder via skip connections. At the end of the decoder, a 1×1 convolutional layer converts the final feature map into a three-channel image, and after passing through a Tanh activation function, a normalized pseudo-sharp image is output.
[0048] A four-layer non-degenerate structure is used to design a transmittance map estimation subnetwork (T-Net). Its input end receives a fog map, and its output end outputs a transmittance map in the range [0,1] through a Sigmoid function. Each layer contains a convolutional layer, a batch normalization layer, and a LeakyReLU activation function.
[0049] An atmospheric light map estimation sub-network (A-Net) based on the U-Net architecture is designed. Its input end receives the fog map, and its output end outputs the atmospheric light map through transposed convolutional layer, LeakyReLU activation and average pooling. The encoder and decoder are each composed of two convolutional blocks. The encoder block performs convolution and ReLU activation, and the decoder block performs upsampling, convolution and ReLU activation. The above three sub-networks are integrated to obtain the physical perception defogging model.
[0050] Inputting a real-world fog image into a physics-based defogging model yields a pseudo-sharp image J. * (i.e., the preliminary estimated fog-free image), transmittance map T (characterizing the degree of attenuation of scene light by the medium), and atmospheric light map A (characterizing the global atmospheric illumination intensity), based on the atmospheric scattering model and the estimated J * The physical perception fog map I is reconstructed using T and A. A self-cyclic consistency loss is used, and the physical perception defogging model is trained by comparing the real-world fog map with the physical perception fog map.
[0051] A Markov chain mechanism is added before the Unet network to form a diffusion generation module. The input of the diffusion generation module is connected to the output of the physical perception dehazing model to obtain a physical-guided dehazing diffusion model.
[0052] The fog image to be processed is input into the trained physical perception dehazing model to obtain the output pseudo-clear image J. * The physically guided diffusion model is input with a transmittance map T and a real clear image. At each time step, the model gradually adds noise to the real clear image to generate noisy image samples. The noise removal process is then dynamically adjusted based on the conditions provided by the pseudo-clear image and the transmittance map to generate a clear dehazed image. The physically guided diffusion model is trained by simultaneously minimizing the error between predicted noise and real noise, as well as the error between the real clear image and the clear dehazed image.
[0053] The fog image to be processed is input into a trained physical perception dehazing model to obtain a pseudo-sharp image and a transmittance image. A noise model combining the current fog image is initialized and input into a trained physical guidance dehazing diffusion model. The pseudo-sharp image and the transmittance image are injected into the UNet network through a feature concatenation method using a conditional guidance mechanism to guide the direction of noise removal. Then, the UNet network is used to capture the features of the noise distribution, predict and remove the noise components in the current time step to obtain an intermediate image. The intermediate image corresponding to the final time step is used as the clear image after dehazing.
[0054] A specific embodiment of the present invention is provided:
[0055] S1. Construct a physical perception defogging model.
[0056] A pseudo-sharp image estimation subnetwork (J-Net) is designed using a symmetric encoder-decoder architecture. The encoder receives a fogged image and extracts hierarchical features through four levels of downsampling, while the decoder recovers the resolution through corresponding upsampling. Furthermore, a multi-scale feature fusion module (MSF) is integrated, which fuses features from layers n, n-1, and n-2 of the encoder through cross-scale dilated convolutions (dilation rates 1 / 3 / 5). After compression by a 3×3 convolution, these features are added to the features from the same layer of the decoder via skip connections. At the end of the decoder, a 1×1 convolutional layer converts the final feature map into a three-channel image, which is then processed by a Tanh activation function to output a normalized pseudo-sharp image.
[0057] A four-layer non-degenerate structure is used to design a transmittance map estimation subnetwork (T-Net). Its input end receives a fog map, and its output end outputs a transmittance map in the range [0,1] through a convolutional layer and a sigmoid function. Each non-degenerate structure contains a convolutional layer, a batch normalization layer and a LeakyReLU activation function.
[0058] An atmospheric light map estimation sub-network (A-Net) based on the U-Net architecture is designed. The encoder input receives the fog map and extracts hierarchical features through four levels of downsampling. The decoder restores the resolution through corresponding upsampling. The decoder outputs the atmospheric light estimate through transposed convolutional layers, LeakyReLU activation, and average pooling. In this embodiment, the encoder and decoder are each composed of two convolutional blocks. The encoding block performs convolution and ReLU activation, and the decoding block performs upsampling, convolution, and ReLU activation. The above three sub-networks are integrated to obtain the physical perception dehazing model.
[0059] In the J-Net network, the multi-scale feature fusion module specifically includes: the input of the MSF module is connected to the feature maps {F} output from the last three layers of the encoder. n-2 ,F n-1 ,F n Multi-scale features are extracted using parallel k×k dilated convolution kernels:
[0060]
[0061] Among them, Conv dil=s This represents a convolution operation with an inflation rate of s. Let i represent the feature representation of the i-th layer at scale s, where i = n, n-1, n-2;
[0062] Cross-scale feature fusion is achieved through feature concatenation and 1×1 convolution:
[0063]
[0064] Where || represents the channel splicing operation, Conv 1×1 This indicates that a 1×1 convolution achieves channel dimension transformation, F fuse It is the output feature after cross-scale feature fusion;
[0065] Dynamically adjust the fused features using a channel attention mechanism:
[0066] w = σ(MLP(GAP(F) fuse )))
[0067]
[0068] Where GAP represents global average pooling, σ represents the sigmoid activation function, and w represents the generated channel attention weight vector. F represents channel-by-channel multiplication. att Indicates the adjusted features;
[0069] Feature maps are compressed using 3×3 convolution:
[0070] F compress =Conv 3×3 (F att )
[0071] Among them, Conv 3×3 It is a 3×3 ordinary convolutional layer used for spatial local feature compression and dimensionality reduction, F compress It is the compressed output feature map.
[0072] Feature map resolution restoration is achieved through transposed convolution.
[0073] F out,i =Deconv(F compress )+F dec,i
[0074] Where Deconv represents the transpose convolution operation, F dec,i F represents the feature map output of the i-th level of the decoder. out,i This represents the output feature map after resolution restoration and feature fusion, which serves as the input to the (i+1)th level of the decoder; the F value output from the last level... dec,n At the end of the decoder, the image is converted into a three-channel image through a 1×1 convolutional layer and then normalized to a pseudo-sharp image after passing through the Tanh activation function.
[0075] S2. Preprocess the fog map in the dataset, including resizing and normalization, to ensure the quality and consistency of the input data. For example... Figure 1 As shown, the preprocessed images are fed into a physically-aware dehazing model to obtain a pseudo-sharp image, transmittance map, and atmospheric illumination map for each image. Then, an atmospheric scattering model is used to reconstruct a physically-aware fog image from the pseudo-sharp image, transmittance map, and atmospheric illumination map, and the self-cyclic consistency loss between the real-world fog image and the physically-aware fog image is calculated. By minimizing the difference between the reconstructed fog image and the input fog image, the parameters of the physically-aware dehazing model are optimized inversely, improving the accuracy of the pseudo-sharp image and transmittance map estimation. This provides more accurate physical guidance conditions for the subsequent physically-guided dehazing diffusion model, ensuring that the generated dehazed sharp image conforms to physical laws while maintaining visual realism.
[0076] Specifically, the physical perception dehazing model decomposes the fog map into a pseudo-sharp map, a transmittance map, and an atmospheric light map through three sub-networks. Then, an atmospheric scattering model is used to reconstruct the pseudo-sharp map, transmittance map, and atmospheric light map into a physical perception fog map. In this way, by minimizing the difference between the reconstructed fog map and the input fog map, the parameters of the physical perception dehazing model are optimized in reverse, and the accuracy of the pseudo-sharp map and transmittance map estimation can be improved.
[0077] In this embodiment, the atmospheric scattering model is used to reconstruct the formula as follows:
[0078] I(x)=J * (x)t(x)+A(1-t(x))
[0079] Where I(x) is the reconstructed fog map (pixel position x), J * (x) is the estimated pseudo-sharp image (i.e., the preliminary estimated haze-free image), A is the estimated atmospheric light map (characterizing the global atmospheric illumination intensity), and t(x) is the estimated transmittance map (characterizing the degree of attenuation of scene light by the medium).
[0080] It is important to note that the pseudo-sharp image J generated by the physical perception dehazing model * (x) is not a truly dehazed image. This is because traditional atmospheric scattering models assume uniform fog and single scattering, neglecting the spatial heterogeneity and multiple scattering effects in real-world scenes. Furthermore, they fail to consider depth-dependent scattering and strong scattering in localized dense fog regions, resulting in a falsely clear image J. * (x) may contain residual haze and color distortion, resulting in a significant deviation from the true clear image J(x) after dehazing.
[0081] S3. A Markov chain mechanism is added before the diffusion model (Unet network) to form a diffusion generation module. The input of this diffusion generation module is connected to the output of the trained physical awareness dehazing model obtained in S2, thereby constructing a physical-guided dehazing diffusion model.
[0082] The working process of the physics-guided defogging diffusion model is as follows: Figure 2 As shown, the fog image to be processed is input into a trained physically-aware dehazing model, which outputs a pseudo-sharp image (i.e., a preliminary estimate of the fog-free image) and a transmittance map (characterizing the degree of light attenuation by the medium). Subsequently, the sharp image obtained from the real world is input into the diffusion generation module. In the diffusion generation module, a Markov chain mechanism is used for noise addition. At each time step, the model gradually adds noise to the sharp image until it approximates Gaussian noise. In this way, the model can progressively simulate the noise accumulation process.
[0083] Specifically, this involves using a Markov chain mechanism to add noise to the clear image obtained from the real world through the physical perception dehazing model. At each time step, the model gradually adds noise to the clear image, thereby generating noisy image samples. The specific formula is as follows:
[0084]
[0085] Where t represents a time step, T is the last time step, and x0, x1, ..., x T It is a sequence of images generated at each step, where x0 is the original image, i.e., the true, clear image; x t The image at time step t represents the noisy image, where ∈ is Gaussian noise distributed in (0,1). It is the product of variables that gradually decrease with increasing time steps, x T It is the noisy image at the last time step, which theoretically conforms to a Gaussian noise distribution.
[0086] The relationship between variables and hyperparameters is as follows:
[0087] α t =1-β t
[0088]
[0089] Where, α t β represents the variable at time step t. t It is a hyperparameter that gradually increases with the time step.
[0090] S4. For a noisy image x that conforms to a Gaussian noise distribution TThe Unet network is used to predict and remove noise components in the current step. This denoising process is gradual; the model removes noise step by step over a series of time steps, generating a clearer intermediate image at each step until a noise-free dehazed image that closely approximates the real data is recovered. Simultaneously, during the denoising process, pseudo-sharp images and transmittance maps serve as conditional inputs, effectively guiding the diffusion model's generation direction. The pseudo-sharp image provides the basic structure of the scene and the initial dehazing effect, helping the diffusion model determine the overall layout and detailed features of the clear image in the initial stage; the transmittance map characterizes the degree of light attenuation caused by fog, providing the model with spatial information on noise intensity and distribution.
[0091] At each time step, the model dynamically adjusts the noise removal process by combining details from the pseudo-sharp image and spatial information from the transmittance map. This allows the Unet network to more accurately simulate the noise characteristics of real data at each time step, thereby optimizing the dehazing effect and generating an output that more closely resembles the dehazed sharp image. The physically guided diffusion model is trained by simultaneously minimizing the error between predicted noise and real noise, as well as the error between the sharp image and the dehazed sharp image.
[0092] Specifically, this includes: removing noise components from the current time step based on condition-guided noise prediction; the formula for the removal process is as follows:
[0093]
[0094] Where, x t The fog map represents the current time step t, x t-1 α represents the image output at the previous time step t-1 after denoising. t and β t These are noise scheduling parameters that control the proportion of noise added and removed during the diffusion process. θ (·) is a noise prediction network used to estimate the noise components in the current image, t is the current denoising time step, which affects the noise intensity and denoising strategy, and σ t The standard deviation of noise related to the time step is represented to control the randomness in the denoising process. z is standard Gaussian noise, used to maintain diversity during sampling. Concat[J] * [,T] is the feature stitching operation, which combines the pseudo-sharp image and the transmittance image as the conditional input for noise prediction. The specific formula is as follows:
[0095] Concat[J * [T] = 0.6 * J * +0.4*T
[0096] Among them, J *T is a pseudo-clear image estimated by the physical perception defogging model; T is a transmittance map estimated by the physical perception defogging model, which describes the fog concentration distribution at each point in the scene. The two together serve as defogging guidance conditions to assist in defogging.
[0097] S5. Input the fog image to be processed into the trained physics-guided dehazing diffusion model. The fog image to be processed first outputs a pseudo-sharp image and a transmittance image through the trained physics-aware dehazing model. Then, initialize a noise based on the current fog image and input it into the trained diffusion generation module. The pseudo-sharp image and the transmittance image are concatenated as conditional input to guide the direction of noise reduction. Subsequently, the diffusion generation module is used to capture the characteristics of the noise distribution, predict and remove the noise components in the current time step to obtain an intermediate image. The intermediate image corresponding to the final time step is used as the dehazed clear image.
[0098] like Figure 3 As shown, the accuracy of each method was tested on datasets with different haze concentrations, and the peak signal-to-noise ratio (PSNR) was used to characterize the accuracy results. By comparing the various methods, it can be seen that the overall success rate of our proposed method is higher than that of methods based on Generative Adversarial Networks (GANs) and Convolutional Neural Networks (CNNs), and the advantage of our method gradually increases with the increase of haze concentration. This indicates that the physical-guided dehazing diffusion model proposed in our method can more accurately simulate the physical properties of fog and its impact on images, thereby better restoring image details and structure during the dehazing process. This advantage of physical modeling is particularly evident when haze concentrations are high.
[0099] This embodiment also provides an image dehazing system based on a physically guided diffusion model that implements the above method, comprising:
[0100] The fog image input module is used to obtain a real-world fog image input physical perception defogging model, and to obtain a corresponding clear image input diffusion model training module;
[0101] The physical perception dehazing model module includes a pseudo-sharp image estimation sub-network, a transmittance image estimation sub-network, and an atmospheric light image estimation sub-network, which are used to generate a pseudo-sharp image J*, a transmittance image T, and an atmospheric light image A, respectively. The atmospheric light image estimation sub-network is only used in the training process of the physical perception dehazing model. During the inference process, its output pseudo-sharp image J* and transmittance image T are used as the conditional inputs of the diffusion model.
[0102] The diffusion model training module is used to input a clear image into the diffusion model, gradually add noise to generate a noisy image through a Markov chain mechanism, and inject the pseudo-clear image J* and the transmittance image T into the diffusion model using a feature concatenation method using a conditional guidance mechanism to guide the denoising direction and gradually predict and remove noise components to generate intermediate images. The intermediate image at the final time step is output as the dehazed clear image. At the same time, the module minimizes the error between the predicted noise and the actual noise, as well as the error between the clear image and the dehazed clear image, to train the diffusion model.
[0103] The image dehazing module based on the physical guided diffusion model is used to perform image dehazing processing using a physical guided diffusion model composed of a trained physical perception dehazing model and a diffusion model.
[0104] For the system embodiments, since they basically correspond to the method embodiments, relevant details can be found in the descriptions of the method embodiments; the implementation methods of the remaining modules will not be repeated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0105] The system embodiments of the present invention can be applied to any device with data processing capabilities, such as a computer or other similar device. The system embodiments can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution.
[0106] It should also be noted that the image dehazing method based on a physically guided diffusion model in the above embodiments can essentially be executed by a computer program. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the method provided in the above embodiments, which includes a memory and a processor;
[0107] The memory is used to store computer programs;
[0108] The processor is configured to implement the image dehazing method based on the physical guided diffusion model in the above embodiments when executing the computer program.
[0109] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0110] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.
Claims
1. An image dehazing method based on a physically guided diffusion model, characterized in that, Includes the following steps: (1) Obtain real-world fog and clear images; (2) Input the fog image into the trained physical perception defogging model and output the pseudo-sharp image J* and the transmittance image T, wherein the pseudo-sharp image J* is the preliminary estimated fog-free image and the transmittance image represents the degree of attenuation of scene light by the medium. The physical perception dehazing model includes a pseudo-sharp image estimation sub-network, a transmittance image estimation sub-network, and an atmospheric light image estimation sub-network, which are used to generate the pseudo-sharp image J*, the transmittance image T, and the atmospheric light image A, respectively; among them, the atmospheric light image estimation sub-network is only used in the training process of the physical perception dehazing model. (3) Input the clear image into the diffusion model, gradually add noise to generate a noise image through the Markov chain mechanism, and use the conditional guidance mechanism to inject the pseudo clear image J* and the transmittance image T into the diffusion model by feature stitching to guide the denoising direction and gradually predict and remove noise components to generate an intermediate image. The intermediate image at the final time step is output as the clear image after dehazing. At the same time, minimize the error between the predicted noise and the real noise, as well as the error between the clear image and the clear image after dehazing, to train the diffusion model. (4) Image dehazing is performed using a physical-guided diffusion model consisting of a trained physical perception dehazing model and a diffusion model.
2. The image dehazing method based on a physically guided diffusion model according to claim 1, characterized in that, The sharp image estimation subnetwork adopts a symmetric encoder-decoder structure and integrates a multi-scale feature fusion module between the encoder and the decoder. The encoder extracts hierarchical features through downsampling, and the decoder restores the resolution through upsampling. At the end of the decoder, a 1×1 convolutional layer is used to convert the feature map output by the last layer of the decoder into a three-channel image, and then the normalized pseudo-sharp image is output after passing through the Tanh activation function.
3. The image dehazing method based on a physically guided diffusion model according to claim 2, characterized in that, The computation process of the multi-scale feature fusion module includes: First, multi-scale features from the last m+1 layer of the encoder are extracted using parallel dilated convolution kernels and then initially fused. Among them, Conv dil=s This represents a convolution operation with an inflation rate of s. Let || denote the feature representation of the i-th layer at scale s, and || denote the channel concatenation operation. Conv 1×1 F represents a 1×1 convolution. fuse This indicates the initial fusion characteristics, where m ≤ n-1; Then, the initial fusion features are dynamically adjusted using a channel attention mechanism: w=σ(MLP(GAP(F fuse ))) Where GAP represents global average pooling, MLP represents a fully connected layer, σ represents the sigmoid activation function, and w represents the generated channel attention weight vector. F represents channel-by-channel multiplication. att This represents the dynamically adjusted fusion features, which serve as the output of the multi-scale feature fusion module.
4. The image dehazing method based on a physically guided diffusion model according to claim 2, characterized in that, The output of the multi-scale feature fusion module is compressed by a 3×3 convolution and then added to the features of the same layer of the decoder via skip connections. This result is then used as the input to the next layer of the decoder, as shown below: F out,i =Deconv(Conv 3×3 (F att ))+F dec,i Among them, F att Conv represents the output of the multi-scale feature fusion module. 3×3 F represents a 3×3 convolutional layer. dec,i F represents the feature map output of the i-th level of the decoder. out,i This represents the output feature map after resolution restoration and feature fusion, which serves as the input to the (i+1)th level of the decoder.
5. The image dehazing method based on a physically guided diffusion model according to claim 1, characterized in that, The transmittance map estimation subnetwork adopts a multi-layer non-degenerate structure. Each non-degenerate layer contains a convolutional layer, a batch normalization layer, and a LeakyReLU activation function. The output of the last layer generates a transmittance map T through a convolutional layer and a sigmoid function.
6. The image dehazing method based on a physically guided diffusion model according to claim 1, characterized in that, The atmospheric light map estimation subnetwork adopts a symmetric encoder-decoder structure. The encoder extracts hierarchical features through downsampling, the decoder restores the resolution through upsampling, and the decoder generates the atmospheric light map at the end through transposed convolutional layers and average pooling.
7. The image dehazing method based on a physically guided diffusion model according to claim 1, characterized in that, The process of training a physics-based dehazing model includes: Using real-world fog images as input to the physical perception dehazing model, estimated pseudo-clear images J*, transmittance maps T, and atmospheric light maps A are generated. The physical perception fog image I is then reconstructed using the following formula: I(x)=J * (x)t(x)+A(1-t(x)) Where I(x) represents the value of pixel x in the reconstructed fog image, J * (x) represents the value of pixel x in the estimated pseudo-clear image, A represents the value of pixel x in the estimated atmospheric light image, and t(x) represents the value of pixel x in the estimated transmittance image. We employ a self-circulating consistency loss method and optimize model parameters by comparing real-world fog maps with reconstructed fog maps.
8. The image dehazing method based on a physically guided diffusion model according to claim 1, characterized in that, The denoising process in step (3) is represented as follows: Where, x t The fog map represents the current time step t, x t-1 α represents the image output at the previous time step t-1 after denoising. t and β t These are noise scheduling parameters, ∈ θ (·) represents the noise prediction network, t is the current denoising time step, and σ t This represents the standard deviation of noise associated with the time step, where z is the standard Gaussian noise, and Concat[J] * [,T] is the feature stitching operation, which stitches the pseudo-sharp image J... * The transmittance map T is combined with the transmittance map T as a conditional input for noise prediction.
9. An image dehazing system based on a physically guided diffusion model, characterized in that, include: The fog image input module is used to obtain a real-world fog image input physical perception defogging model, and to obtain a corresponding clear image input diffusion model training module; The physical perception dehazing model module includes a pseudo-sharp image estimation sub-network, a transmittance image estimation sub-network, and an atmospheric light image estimation sub-network, which are used to generate a pseudo-sharp image J*, a transmittance image T, and an atmospheric light image A, respectively. The atmospheric light image estimation sub-network is only used in the training process of the physical perception dehazing model. During the inference process, its output pseudo-sharp image J* and transmittance image T are used as the conditional inputs of the diffusion model. The diffusion model training module is used to input a clear image into the diffusion model, gradually add noise to generate a noisy image through a Markov chain mechanism, and inject the pseudo-clear image J* and the transmittance image T into the diffusion model using a feature concatenation method using a conditional guidance mechanism to guide the denoising direction and gradually predict and remove noise components to generate intermediate images. The intermediate image at the final time step is output as the dehazed clear image. At the same time, the module minimizes the error between the predicted noise and the actual noise, as well as the error between the clear image and the dehazed clear image, to train the diffusion model. The image dehazing module based on the physical guided diffusion model is used to perform image dehazing processing using a physical guided diffusion model composed of a trained physical perception dehazing model and a diffusion model.
10. A computer electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the image dehazing method based on the physical guided diffusion model as described in any one of claims 1 to 8.
Citation Information
Cited By
Image restoration method and device based on residual denoising diffusion model
CN121563847A