Image defogging method and device based on wavelet transform, equipment and storage medium
Through the image defogging method based on wavelet transformation, the image is decomposed into low-frequency and high-frequency components, and the diffusion model and high-frequency enhancement processing are performed respectively, which solves the problem of poor image defogging effect in the prior art, and achieves clear recovery of image details and global features.
Patent Information
- Application Number
- CN202510083743.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The existing image defog removal method is difficult to clearly restore the details and global features of the image, resulting in unsatisfactory defog removal.
The image defogging method based on wavelet transformation is adopted, and the image is decomposed into low-frequency components and high-frequency components through wavelet transformation, and the low-frequency conditional diffusion model and high-frequency enhancement module are respectively processed, and finally the defogging image is reconstructed through inverse wavelet transformation and multi-scale pooling.
Effectively remove the impact of haze on the global features of the image, restore the details and global information of the image, and improve the image defog removal effect and quality.
Smart Images

Figure CN119991485A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing and application technology, and in particular to an image defogging method, device, equipment and storage medium based on wavelet transform. Background Art
[0002] With the rapid development of information technology, computer vision and image processing technologies have been widely used in many fields such as autonomous driving, intelligent monitoring, and remote sensing image processing. However, haze weather is one of the key factors affecting image quality and subsequent analysis. Haze not only causes image blur and reduced contrast, but also may hide important details in the image, seriously affecting the accuracy of image understanding and subsequent image analysis tasks. Therefore, how to effectively remove haze and restore image clarity and details during image processing has always been a major problem that needs to be solved in the field of computer vision.
[0003] At present, image defogging methods can be mainly divided into physical model-based defogging methods and end-to-end defogging methods. Traditional physical model-based methods mainly simulate the impact of haze on light propagation, combine atmospheric scattering models or optical transmission models, reverse the concentration and scattering effect of haze, and then restore the clarity of the image. Although this type of method can intuitively explain the physical mechanism of the defogging process through a strong theoretical basis, this method is usually suitable for scenes with relatively uniform haze, that is, it can only effectively restore the detailed information in the image under certain conditions. In comparison, end-to-end defogging methods use deep learning technology, especially models such as convolutional neural networks (CNNs), to learn the mapping relationship from foggy images to defogging images through a large amount of training data. This type of method does not rely on any physical model or manually designed features, but instead shows strong robustness and adaptability in complex haze environments through data-driven methods, and can handle images in dynamic environments and non-uniform haze.
[0004] With the development of generative models, especially the widespread application of generative adversarial networks (GAN) and variational autoencoders (VAE) in the field of image restoration, image dehazing technology has entered a new research stage. Although GAN and VAE show strong generative capabilities in image restoration, they are unstable during the training process, and the generated images are prone to artifacts or blurred details. Therefore, it is difficult for existing solutions to clearly restore the details and global features of the image, resulting in poor image dehazing effect. Summary of the invention
[0005] In view of this, in order to address the above shortcomings, it is necessary to propose an image defogging method, device, equipment and storage medium based on wavelet transform to clearly restore the details and global information of the image and improve the image defogging effect and image quality.
[0006] In a first aspect, the present invention provides an image defogging method based on wavelet transform, comprising:
[0007] Obtain a foggy image to be defogged;
[0008] Performing wavelet transform on the foggy image to decompose it into low-frequency components and high-frequency components; wherein the low-frequency components contain global structural information of the image, and the high-frequency components contain details and texture information of the image;
[0009] Inputting the low-frequency component into a pre-constructed low-frequency conditional diffusion model, and outputting the restored low-frequency component; wherein the low-frequency conditional diffusion model is used to restore the low-frequency structure of the image from the noise based on the diffusion and denoising process;
[0010] Restoring and enhancing the edge and texture of the high-frequency component using a pre-constructed high-frequency enhancement module to obtain a restored high-frequency component; wherein the high-frequency enhancement module is constructed based on Gabor convolutions in multiple different directions;
[0011] Reconstructing the restored low-frequency component and the restored high-frequency component into a spatial domain image through inverse wavelet transformation;
[0012] Based on multi-scale pooling, local information and global information in the spatial domain image are aggregated to obtain a final defogging image of the foggy image.
[0013] Preferably, performing wavelet transform on the foggy image includes:
[0014] For the foggy image J∈R H×W×C , the foggy image J is decomposed using two-dimensional discrete wavelet transform to obtain four components: Among them, X A is the low frequency component, X H , X V , X D They are horizontal high-frequency component, vertical high-frequency component and diagonal high-frequency component respectively; the basis function of the two-dimensional discrete wavelet transform is the Haar wavelet basis function.
[0015] Preferably, the pre-trained low-frequency conditional diffusion model is expressed as:
[0016]
[0017] in, Used to characterize the image obtained in step t, Used to characterize low-frequency components, α t As a hyperparameter to control the variance of the noise added at each step, represents the noise in the model predictions.
[0018] Preferably, the restoring and enhancing the edge and texture of the high-frequency component by using a pre-built high-frequency enhancement module comprises:
[0019] For the horizontal high frequency component X H , vertical high frequency component X V and the diagonal high frequency component X D In any of the above, execute:
[0020] Based on the following calculation formula, the current high-frequency component is mapped to a high-dimensional feature space using a depth-separable convolution, and four 7×7 Gabor convolutions are used to obtain high-dimensional features in four directions:
[0021]
[0022] Among them, x0, They are respectively the current high frequency components X α Corresponding to the high-dimensional features in four directions, Conv Gabor Characterizing Gabor convolution, Conv Depth Used to characterize depthwise separable convolutions;
[0023] Based on the following calculation formula, two cross attention layers are used to convert the current high-frequency component X α The high-frequency features in four directions are fused and spliced to obtain the restored high-frequency component of the current high-frequency component:
[0024]
[0025] in, It is used to characterize the restored high-frequency component of the current high-frequency component, ψ(·) represents batch normalization, Attn(·) represents the cross self-attention mechanism, Conv Depth Used to characterize depthwise separable convolution, Cat(·) is an abbreviation of the Concat(·) operation.
[0026] Preferably, aggregating local information and global information in the spatial domain image based on multi-scale pooling includes:
[0027] This is achieved using the following calculation formula group:
[0028]
[0029] Among them, MP n×nIndicates maximum pooling with n×n kernel, represents element-by-element addition, x n×n Represents the feature map after maximum pooling with n×n kernel, Representing spatial domain images Feature map after convolution and batch normalization, ψ(·) represents batch normalization.
[0030] Preferably, when constructing a complete model for image defogging, the restored low-frequency component output by the low-frequency conditional diffusion model is defined as The restored high-frequency component obtained by the high-frequency enhancement module is right The spatial domain image after inverse wavelet transform is is the reference image X GT Wavelet decomposition image of;
[0031] The low-frequency loss function is expressed as follows:
[0032]
[0033] Where, L low It is used to represent the low-frequency loss function, ε is the real noise added, and ε θ is the noise predicted by the model;
[0034] High-frequency loss function L high It is expressed as follows:
[0035]
[0036] The overall image loss function is expressed as follows:
[0037]
[0038] Where, L image It is used to characterize the overall image loss function, and SSIM(·) is used to characterize the SSIM loss function;
[0039] The total loss function of the constructed complete model for image dehazing is expressed as:
[0040] L total =αL low +βL high +γL image
[0041] Among them, L total Used to characterize the total loss function, α, β, and γ are weight parameters of low-frequency loss, high-frequency loss, and image loss, respectively.
[0042] In a second aspect, the present invention provides an image defogging device based on wavelet transform, comprising: an image acquisition module, a wavelet transform module, a low-frequency recovery module, a high-frequency recovery module, an inverse wavelet transform module and an image aggregation module;
[0043] The image acquisition module is configured to acquire a foggy image to be defogged;
[0044] The wavelet transform module is configured to perform wavelet transform on the foggy image to decompose it into a low-frequency component and a high-frequency component; wherein the low-frequency component contains global structural information of the image, and the high-frequency component contains details and texture information of the image;
[0045] The low-frequency recovery module is configured to input the low-frequency component into a pre-constructed low-frequency conditional diffusion model and output the restored low-frequency component; wherein the low-frequency conditional diffusion model is used to recover the low-frequency structure of the image from the noise based on the diffusion and denoising process;
[0046] The high-frequency restoration module is configured to restore and enhance the edge and texture of the high-frequency component using a pre-constructed high-frequency enhancement module to obtain a restored high-frequency component; wherein the high-frequency enhancement module is constructed based on Gabor convolutions in multiple different directions;
[0047] The inverse wavelet transform module is configured to reconstruct the restored low-frequency component and the restored high-frequency component into a spatial domain image through inverse wavelet transformation;
[0048] The image aggregation module is configured to aggregate local information and global information in the spatial domain image based on multi-scale pooling to obtain a final defogging image of the foggy image.
[0049] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute any of the methods described in the first aspect.
[0050] In a fourth aspect, the present invention provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements any method described in the first aspect.
[0051] It can be seen from the above technical scheme that in the image defogging method based on wavelet transform provided by the embodiment of the present invention, firstly, a defogging foggy image is obtained, and then the foggy image is subjected to wavelet transform to decompose into low-frequency components and high-frequency components; further, the low-frequency components are input into a pre-constructed low-frequency conditional diffusion model to restore the low-frequency structure of the image from the noise based on the diffusion and denoising process, and at the same time, the edges and textures of the high-frequency components are restored and enhanced using a pre-constructed high-frequency enhancement module; then the obtained restored low-frequency components and restored high-frequency components are reconstructed into a spatial domain image through inverse wavelet transform, and local information and global information in the spatial domain image are aggregated based on multi-scale pooling, so as to obtain the final defogging image. It can be seen that this scheme can decompose the input image into a low-frequency component containing the global structural information of the image and a high-frequency component containing the details and texture information of the image through wavelet transform, and then the low-frequency components are processed through the low-frequency conditional diffusion model, so that the low-frequency structure of the image can be gradually restored from the noise, thereby removing the influence of haze on the global characteristics of the image. At the same time, the high-frequency enhancement module is used to process the high-frequency components, which can improve the clarity of the high-frequency components through Gabor convolution in different directions, ensuring that the edges and textures are fully restored and enhanced. Finally, multi-scale pooling is performed on the reconstructed spatial domain image to ensure the balance between global consistency and local details of the output dehazed image, greatly improving the quality of the dehazed image. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A flow chart of an image defogging method based on wavelet transform provided by an embodiment of the present invention.
[0053] Figure 2 A schematic diagram of the overall architecture of image defogging provided by the present invention.
[0054] Figure 3 Schematic diagram of Haar wavelet transform for dense fog and clear images and reorganization of low-frequency and high-frequency information.
[0055] Figure 4 A visualization of the four directional features extracted using Gabor convolution.
[0056] Figure 5 Visual comparison of the present invention with different image dehazing methods on the Dense-Haze dataset.
[0057] Figure 6 Visual comparison of the proposed method with different image dehazing methods on the NH-Haze dataset.
[0058] Figure 7 Visual comparison of the proposed method with different image dehazing methods on the SOTS dataset. DETAILED DESCRIPTION
[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0060] like Figure 1 As shown, the present invention provides an image defogging method based on wavelet transform, which comprises the following steps:
[0061] Step 101: Obtain a foggy image to be defogged;
[0062] Step 102: performing wavelet transform on the foggy image to decompose it into low-frequency components and high-frequency components; wherein the low-frequency components contain global structural information of the image, and the high-frequency components contain details and texture information of the image;
[0063] Step 103: inputting the low-frequency component into a pre-constructed low-frequency conditional diffusion model, and outputting the restored low-frequency component; wherein the low-frequency conditional diffusion model is used to restore the low-frequency structure of the image from the noise based on the diffusion and denoising process;
[0064] Step 104: using a pre-built high-frequency enhancement module to restore and enhance the edge and texture of the high-frequency component to obtain a restored high-frequency component; wherein the high-frequency enhancement module is constructed based on Gabor convolutions in multiple different directions;
[0065] Step 105: Reconstructing the restored low-frequency component and the restored high-frequency component into a spatial domain image through inverse wavelet transformation;
[0066] Step 106: Based on multi-scale pooling, local information and global information in the spatial domain image are aggregated to obtain the final defogging image of the foggy image.
[0067] like Figure 2As shown in the overall architecture of image defogging, when this embodiment performs defogging on an image, the input image can be decomposed into a low-frequency component containing the global structural information of the image and a high-frequency component containing the details and texture information of the image through wavelet transform, and then the low-frequency component is processed through the low-frequency conditional diffusion model, which can gradually restore the low-frequency structure of the image from the noise, thereby removing the influence of haze on the global characteristics of the image. At the same time, the high-frequency component is processed by the high-frequency enhancement module, and the clarity of the high-frequency component can be improved through Gabor convolution in different directions, ensuring that the edges and textures are fully restored and enhanced. Finally, multi-scale pooling is performed on the reconstructed spatial domain image through multi-scale pooling, which ensures the balance between global consistency and local details in the output defogging image, greatly improving the quality of the defogging image.
[0068] For step 102, wavelet transform is performed on the foggy image to decompose it into low-frequency components and high-frequency components; wherein the low-frequency components contain global structural information of the image, and the high-frequency components contain details and texture information of the image;
[0069] In this embodiment, when performing wavelet transform on the foggy image, the most basic wavelet basis function in discrete wavelet transform, Haar wavelet, is selected to decompose the input image. Specifically, for the foggy image J∈R H×W×C , the foggy image J is decomposed using two-dimensional discrete wavelet transform 2D-DWT to obtain four components: If {X A , X H , X V , X D}=2D-DWT(J), where X A is the low frequency component, X H , X V , X D They are horizontal high frequency component, vertical high frequency component and diagonal high frequency component respectively;
[0070] The low-frequency components contain the overall structure of the image and the most important visual information, while the high-frequency components mainly represent the details and edges of the image. They are sensitive to subtle features in the image. The presence of fog will reduce the difference between the brightness of distant objects and the background brightness, which is reflected in the reduction of image contrast. The fog will also affect the overall brightness of the image, making the image look brighter and uneven. Since the effect of fog is uniform and smooth throughout the image, it will not cause rapid brightness changes like edges or details. Therefore, the impact of fog is more related to the low-frequency information of the image rather than the high-frequency information. Figure 3The figure shows the Haar small transform and low-frequency and high-frequency information reorganization of dense fog and clear images, and it is found that the haze information mainly exists in the low-frequency information. The low-frequency components of the dense fog image and the high-frequency components of the clear image are reorganized, and the resulting image still contains a lot of haze information, while the exchange of the high-frequency information of the two images does not change significantly from the original image. The quantitative results also show that the haze information affects the low-frequency components of the image more and has little effect on the high-frequency components.
[0071] For step 103, the low-frequency component is input into a pre-constructed low-frequency conditional diffusion model, and the low-frequency component is output and restored; wherein the low-frequency conditional diffusion model is used to restore the low-frequency structure of the image from the noise based on the diffusion and denoising process;
[0072] Since haze information affects the low-frequency components of the image more, and the low-frequency components contain the global structural information of the image. Therefore, we consider using the low-frequency conditional diffusion model to process the low-frequency components. For the decision of the model training space, we choose to diffuse in the original pixel space instead of the latent space. Although diffusion in the latent space can reduce computational costs and improve training efficiency, diffusion in the pixel space can avoid reconstruction errors that may be introduced during decoding and decoding, and better maintain the authenticity and details of the image, thereby achieving a higher quality generation effect. Specifically, the low-frequency conditional diffusion model is used to generate clear low-frequency components in the original space. The single-step denoising process of the low-frequency conditional diffusion model can be expressed as:
[0073]
[0074] in, Used to characterize the image obtained in step t, the low-frequency component X A Recorded as α t As a hyperparameter to control the variance of the noise added at each step, represents the noise in the model predictions.
[0075] Specifically, when constructing the low-frequency conditional diffusion model, the clean data is gradually converted into complete Gaussian noise, and then denoised step by step through the reverse process to generate new data. The whole process consists of two stages, the forward diffusion process and the reverse denoising process.
[0076] Forward diffusion process: given a data sample x0 (such as an image), gradually add a small amount of noise to form a series of intermediate states x1, x2, ..., x T , and finally get x which is close to the standard Gaussian noise T . Its formula is:
[0077]
[0078] where α t As a hyperparameter or reparameter to control the variance of the noise added at each step, after T noise additions, the original data will gradually become blurred and eventually become pure noise. I is the unit matrix. In the forward process, the data x of the original sample x0 at any time step t can be derived. t , use q(x t |x0) indicates that the reverse diffusion step q(x t-1 |x t , x0). Specifically:
[0079]
[0080] in, α t =1-β t and
[0081] Reverse denoising process: Given the current state x t , the goal is to gradually denoise and obtain the state x of the previous time step t-1 Specifically, each step in the reverse process is assumed to be a Gaussian distribution, and the parameters of the Gaussian distribution are learned through the model. Its conditional probability is expressed as:
[0082] p θ (x t-1 |x t )=N(x t-1 ;μ θ (x t ,t),∑ θ (x t ,t))
[0083] in, For approximation Use the network to learn the noise ε in the mean θ (x t ,t), which is used to predict the noise component that needs to be removed during denoising at each time step, or to directly learn the mean μ θ (x t ,t), used to predict the image x at the previous time step t-1 .
[0084] Training objectives and sampling process: In order to enable the model to effectively denoise, the loss function of the training process is used to minimize the noise ε predicted by the model. θ (x t ,t) and the actual noise ε. The mean square error MSE is usually used to measure the accuracy of noise prediction. Therefore, the training objective is:
[0085]
[0086] In the sampling process of the low-frequency conditional diffusion model, starting from standard Gaussian noise, the model gradually removes the noise through the reverse process until the final image is generated. In each step, the model predicts the mean and variance of the noise and samples from the conditional Gaussian distribution to generate the image of the previous time step. After multiple iterations of denoising, the final generated sample will be similar to the original image distribution, resulting in a high-quality generated image.
[0087] For step 104, the edge and texture of the high-frequency component are restored and enhanced using a pre-constructed high-frequency enhancement module to obtain a restored high-frequency component; wherein the high-frequency enhancement module is constructed based on Gabor convolutions in multiple different directions;
[0088] In this step, a high-frequency enhancement module is proposed for foggy images, the purpose of which is to transform the three high-frequency components X H , X V , X D To restore and enhance the details. Specifically, the horizontal high-frequency component X H , vertical high frequency component X V and the diagonal high frequency component X D In any of the above, execute:
[0089] Based on the following calculation formula, the current high-frequency component is mapped to a high-dimensional feature space using a depth-separable convolution, and four 7×7 Gabor convolutions are used to obtain high-dimensional features in four directions:
[0090]
[0091] Among them, x0, are the current high frequency components X α Corresponding to the high-dimensional features in four directions, Conv Gabor Characterizing Gabor convolution, Conv Depth Used to characterize depthwise separable convolutions;
[0092] Based on the following calculation formula, two cross attention layers are used to convert the current high-frequency component X α The high-frequency features in four directions are fused and spliced to obtain the restored high-frequency component of the current high-frequency component:
[0093]
[0094] in, It is used to characterize the restored high-frequency component of the current high-frequency component, ψ(·) represents batch normalization, Attn(·) represents the cross self-attention mechanism, ConvDepth It is used to represent the depth-wise separable convolution. Cat(·) is the abbreviation of Concat(·) operation. With X H Having the same dimensions and sizes.
[0095] In this embodiment, a core advantage of considering Gabor convolution is its directional selectivity, which can extract image features in a specific direction. In addition, the convolution kernel parameters used in Gabor convolution can be set to a Gabor filter with a specific direction before training begins, rather than randomly initialized. This means that the convolution kernel in the initial stage already has the characteristics of direction and frequency sensitivity, and there is no need to learn from scratch, thus reducing the complexity of parameter training. During the convolution process, the features of high-frequency information during training are enhanced using Gabor convolutions in four directions, where each Gabor convolution kernel decomposes the information into a specific direction, so that the edge and texture information in the image is more clearly expressed in the feature map.
[0096] like Figure 4 As shown in the figure, the process of extracting four directional features by Gabor convolution is demonstrated. The high-frequency components are processed by Gabor convolution kernels in different directions to generate multi-angle output feature maps. This method uses multi-directional convolution operations to make the edge and texture information in the image more clearly expressed in the feature map. The convolution kernel parameters used in Gabor convolution can be set to a Gabor filter with a specific direction before training begins, rather than randomly initialized. This shows that the convolution kernel in the initial stage already has the characteristics of direction and frequency sensitivity, and there is no need to learn from scratch, thus reducing the complexity of parameter training. In the convolution process, Gabor convolution in four directions is used to enhance the features of high-frequency information during training, where each Gabor convolution kernel decomposes the information into a specific direction, thereby generating response features in the corresponding direction.
[0097] For step 105, the restored low-frequency component and the restored high-frequency component are reconstructed into a spatial domain image through inverse wavelet transformation;
[0098] In this embodiment, the restored low-frequency components are obtained respectively. and restore high frequency components After that, consider using the inverse Haar wavelet transform to reconstruct the spatial domain image
[0099] For step 106, local information and global information in the spatial domain image are aggregated based on multi-scale pooling to obtain a final defogging image of the foggy image.
[0100] In this step, in order to take into account both local and global features, multi-scale maximum pooling (x 5×5、x 9×9 、x 13×13 ) is used to perform feature downsampling to capture feature information of different scales, thereby improving the model’s ability to recognize multi-scale features.
[0101] The low-frequency conditional diffusion module generates and high frequency enhancement module After inverse Haar wavelet transform, the image in the spatial domain is obtained Then it is processed by the multi-scale pooling module to finally obtain the dehazed image X, which is implemented using the following calculation formula group:
[0102]
[0103] Among them, MP n×n Indicates maximum pooling with n×n kernel, represents element-by-element addition, x n×n Represents the feature map after maximum pooling with n×n kernel, Representing spatial domain images Feature map after convolution and batch normalization, ψ(·) represents batch normalization.
[0104] In addition, when constructing a complete model for image dehazing, the restored low-frequency component of the output of the low-frequency conditional diffusion model is defined as The restored high-frequency component obtained by the high-frequency enhancement module is right The spatial domain image after inverse wavelet transform is is the reference image X GT Wavelet decomposition image of;
[0105] The low-frequency loss function is expressed as follows:
[0106]
[0107] Where, L low It is used to represent the low-frequency loss function, ε is the real noise added, and ε θ is the noise predicted by the model;
[0108] For the low-frequency loss function, only the noise and low-frequency components of the diffusion model are constrained, so consider designing a high-frequency loss function L high , which is expressed as follows:
[0109]
[0110] For high-frequency information, the L1 norm constraint can maintain sparsity, so that there will not be too much smoothing when extracting detail information. Therefore, consider combining the L2 norm with the SSIM loss to design the overall image loss function, which is expressed as follows:
[0111]
[0112] Where, L image It is used to characterize the overall image loss function, and SSIM(·) is used to characterize the SSIM loss function;
[0113] Therefore, the total loss function of the constructed complete model for image dehazing is expressed as:
[0114] L total =αL low +βL high +γL image
[0115] Among them, L total It is used to characterize the total loss function. α, β, and γ are the weight parameters of low-frequency loss, high-frequency loss, and image loss, which can be set to 2, 1, and 1 respectively. Therefore, during the model training process, by increasing the low-frequency loss term L low The weight of can effectively suppress the impact of haze on the overall brightness and contrast of the image, thereby better restoring the background information of the image.
[0116] The following uses Python software and uses Pytorch to implement the model framework on a single NVIDIA GeForce RTX 4090 GPU. Images from the natural haze dataset and the synthetic haze dataset are used as input data, and the LPIPS, PSNR, and SSIM indicators are used for evaluation. The present invention is qualitatively and quantitatively compared with the DCP model based on traditional priors and six deep learning-based methods: DehazeNet, FFA-Net, Dehamer, FSDGN, WeatherDiff, and FDCM, where WeatherDiff and FDCM are both dehazing models based on diffusion models.
[0117] Figure 5 The comparative experiment of the present invention and different image dehazing models on the Dense-Haze dataset is demonstrated. When removing dense haze, existing methods often cannot avoid significant loss of content and color, resulting in unsatisfactory dehazing effect. In contrast, the proposed method can better preserve the details and color information of the image, and the generated dehazing result is visually close to the reference clear image.
[0118] Figure 6The comparative experiment of the present invention and different image dehazing models on the NH-Haze dataset is demonstrated. Compared with the FSDGN and Dehamer methods, the present invention can more accurately remove uneven haze and generate a high-fidelity dehazing effect. In particular, in terms of the accuracy of restoring details and colors, the present invention shows obvious advantages, can effectively restore the details of blurred areas, and avoids the common artifacts and blurs in existing methods. The generated images are more realistic and natural in visual effects.
[0119] Figure 7 The comparative experiment of the present invention and different image dehazing models on the SOTS dataset is demonstrated. The present invention and other methods show comparable visual effects on the dataset, and can also effectively restore image details and color information, providing better perceptual quality.
[0120] The quantitative comparison results of different methods on three datasets are shown in Table 1. From the perspective of various evaluation indicators, the present invention shows significant advantages on the Dense-Haze and NH-Haze datasets.
[0121] Table 1
[0122]
[0123] In addition, the present invention also provides an image defogging device based on wavelet transform, comprising: an image acquisition module, a wavelet transform module, a low-frequency recovery module, a high-frequency recovery module, an inverse wavelet transform module and an image aggregation module;
[0124] An image acquisition module, configured to acquire a foggy image to be defogged;
[0125] A wavelet transform module is configured to perform wavelet transform on the foggy image to decompose it into a low-frequency component and a high-frequency component; wherein the low-frequency component contains global structural information of the image, and the high-frequency component contains details and texture information of the image;
[0126] A low-frequency recovery module is configured to input the low-frequency component into a pre-built low-frequency conditional diffusion model and output the restored low-frequency component; wherein the low-frequency conditional diffusion model is used to recover the low-frequency structure of the image from the noise based on the diffusion and denoising process;
[0127] A high-frequency restoration module is configured to restore and enhance the edge and texture of the high-frequency component using a pre-built high-frequency enhancement module to obtain a restored high-frequency component; wherein the high-frequency enhancement module is constructed based on Gabor convolutions in multiple different directions;
[0128] An inverse wavelet transform module is configured to reconstruct the restored low-frequency component and the restored high-frequency component into a spatial domain image through inverse wavelet transform;
[0129] The image aggregation module is configured to aggregate local information and global information in the spatial domain image based on multi-scale pooling to obtain a final defogging image of the foggy image.
[0130] The present specification also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute a method in any one of the embodiments in the specification.
[0131] The present specification also provides a computing device, including a memory and a processor, wherein executable codes are stored in the memory, and when the processor executes the executable codes, a method in any embodiment of the present specification is implemented.
[0132] Since the device embodiments provided by the present invention are based on the same inventive concept as the method embodiments of this specification, the specific contents can be found in the description of the method embodiments of this specification and will not be repeated here.
[0133] The modules or units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. The above disclosure is only the preferred embodiment of the present invention, and of course it cannot be used to limit the scope of the rights of the present invention. Those skilled in the art can understand that all or part of the processes of the above embodiment are implemented, and the equivalent changes made according to the claims of the present invention still fall within the scope of the invention.
Claims
1. An image defogging method based on wavelet transform, characterized in that: include: Obtain a foggy image to be defogged; Performing wavelet transform on the foggy image to decompose it into low-frequency components and high-frequency components; wherein the low-frequency components contain global structural information of the image, and the high-frequency components contain details and texture information of the image; Inputting the low-frequency component into a pre-constructed low-frequency conditional diffusion model, and outputting the restored low-frequency component; wherein the low-frequency conditional diffusion model is used to restore the low-frequency structure of the image from the noise based on the diffusion and denoising process; Restoring and enhancing the edge and texture of the high-frequency component using a pre-constructed high-frequency enhancement module to obtain a restored high-frequency component; wherein the high-frequency enhancement module is constructed based on Gabor convolutions in multiple different directions; Reconstructing the restored low-frequency component and the restored high-frequency component into a spatial domain image through inverse wavelet transformation; Based on multi-scale pooling, local information and global information in the spatial domain image are aggregated to obtain a final defogging image of the foggy image.
2. The image defogging method based on wavelet transform according to claim 1, characterized in that: The performing wavelet transform on the foggy image comprises: For the foggy image J∈R H×W×C , the foggy image J is decomposed using two-dimensional discrete wavelet transform to obtain four components: Among them, X A is the low frequency component, X H , X V , X D They are horizontal high-frequency component, vertical high-frequency component and diagonal high-frequency component respectively; the basis function of the two-dimensional discrete wavelet transform is the Haar wavelet basis function.
3. The image defogging method based on wavelet transform according to claim 1, characterized in that: The pre-trained low-frequency conditional diffusion model is expressed as: in, Used to characterize the image obtained in step t, Used to characterize low-frequency components, α t As a hyperparameter to control the variance of the noise added at each step, represents the noise in the model predictions.
4. The image defogging method based on wavelet transform according to claim 2 is characterized in that: The method of restoring and enhancing the edge and texture of the high-frequency component by using a pre-built high-frequency enhancement module includes: For the horizontal high frequency component X H , vertical high frequency component X V and the diagonal high frequency component X D In any of the above, execute: Based on the following calculation formula, the current high-frequency component is mapped to a high-dimensional feature space using a depth-separable convolution, and four 7×7 Gabor convolutions are used to obtain high-dimensional features in four directions: Among them, x0, are the current high frequency components X α Corresponding to the high-dimensional features in four directions, Conv Gabor Characterizing Gabor convolution, Conv Depth Used to characterize depthwise separable convolutions; Based on the following calculation formula, two cross attention layers are used to convert the current high-frequency component X α The high-frequency features in four directions are fused and spliced to obtain the restored high-frequency component of the current high-frequency component: in, It is used to characterize the restored high-frequency component of the current high-frequency component, ψ(·) represents batch normalization, Attn(·) represents the cross self-attention mechanism, Conv Depth Used to characterize depthwise separable convolution, Cat(·) is an abbreviation of the Concat(·) operation.
5. The image defogging method based on wavelet transform according to claim 1, characterized in that: The aggregating local information and global information in the spatial domain image based on multi-scale pooling includes: This is achieved using the following calculation formula group: Among them, MP n×n Indicates maximum pooling with n×n kernel, represents element-by-element addition, x n×n Represents the feature map after maximum pooling with n×n kernel, Representing spatial domain images Feature map after convolution and batch normalization, ψ(·) represents batch normalization.
6. The image defogging method based on wavelet transform according to claim 1, characterized in that: When constructing a complete model for image dehazing, the restored low-frequency component of the output of the low-frequency conditional diffusion model is defined as The restored high-frequency component obtained by the high-frequency enhancement module is right The spatial domain image after inverse wavelet transform is is the reference image X GT Wavelet decomposition image of; The low-frequency loss function is expressed as follows: Where, L low It is used to represent the low-frequency loss function, ε is the real noise added, and ε θ is the noise predicted by the model; High-frequency loss function L high It is expressed as follows: The overall image loss function is expressed as follows: Where, L image It is used to characterize the overall image loss function, and SSIM(·) is used to characterize the SSIM loss function; The total loss function of the constructed complete model for image dehazing is expressed as: L total =αL low +βL high +γL image Among them, L total Used to characterize the total loss function, α, β, and γ are weight parameters of low-frequency loss, high-frequency loss, and image loss, respectively.
7. An image defogging device based on wavelet transform, characterized in that: include: Image acquisition module, wavelet transform module, low frequency recovery module, high frequency recovery module, inverse wavelet transform module and image aggregation module; The image acquisition module is configured to acquire a foggy image to be defogged; The wavelet transform module is configured to perform wavelet transform on the foggy image to decompose it into a low-frequency component and a high-frequency component; wherein the low-frequency component contains global structural information of the image, and the high-frequency component contains details and texture information of the image; The low-frequency recovery module is configured to input the low-frequency component into a pre-constructed low-frequency conditional diffusion model and output the restored low-frequency component; wherein the low-frequency conditional diffusion model is used to recover the low-frequency structure of the image from the noise based on the diffusion and denoising process; The high-frequency restoration module is configured to restore and enhance the edge and texture of the high-frequency component using a pre-constructed high-frequency enhancement module to obtain a restored high-frequency component; wherein the high-frequency enhancement module is constructed based on Gabor convolutions in multiple different directions; The inverse wavelet transform module is configured to reconstruct the restored low-frequency component and the restored high-frequency component into a spatial domain image through inverse wavelet transformation; The image aggregation module is configured to aggregate local information and global information in the spatial domain image based on multi-scale pooling to obtain a final defogging image of the foggy image.
8. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 6.
9. A computing device, comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Frequency domain decomposition based single image defogging acceleration method
CN107274369A
Method for removing haze in urban remote sensing image
CN111275652A
Cited By
Image defogging system and method for low-altitude scene
CN120598820A
Remote sensing image defogging method and device based on wavelet multi-scale decomposition
CN120876239A