Emergency image defogging method based on wavelet transform frequency domain characteristics
Through the image defog removal method based on the frequency domain characteristics of wavelet transform, the problem of poor image defog removal in emergencies is solved, and the rapid and accurate defog removal effect is achieved, and the calculation cost is reduced, which is suitable for real-time processing needs.
Patent Information
- Application Number
- CN202510623107.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The prior art has poor image defog removal effect when handling emergencies, especially when thick fog or mist is unevenly distributed, resulting in poor reconstruction effect, and the high computing cost of the Transformer structure is not suitable for real-time processing needs.
The image defogging method based on the frequency domain characteristics of wavelet transform is adopted. By constructing the emergency defogging data set, the image defogging model is trained. The high-frequency features of the fogging image are extracted using the wavelet transform, and frequency domain inverse fusion with the low-frequency features of the fogging image is carried out to achieve the recovery of texture information.
It realizes the rapid and accurate removal of fog occlusion in emergencies, restores high-quality foggy-free images, reduces calculation costs, and is suitable for real-time processing needs.
Smart Images

Figure CN120125469A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for dehazing emergency images based on frequency domain features of wavelet transform. Background Art
[0002] The influence of fog on the visual system cannot be ignored in some specific environments. Especially in downstream tasks such as emergency detection, autonomous driving, and target recognition, the presence of fog will cause serious loss of image information, thereby affecting the performance and reliability of the system. In recent years, image dehazing has received extensive attention as an important and challenging image inversion problem in the field of computer vision. Its core task is to restore a clear fog-free image from a blurred image. Traditional image dehazing methods usually perform dehazing processing based on prior knowledge through an atmospheric scattering model. However, when dealing with images that do not fully match the prior conditions, the dehazing effect of these methods is poor. The Transformer structure has achieved good application results due to its excellent performance in processing sequence data and has gradually been introduced into the image dehazing task. The image dehazing method using the Transformer structure can make full use of its multi-head self-attention mechanism to effectively model the global information in the image, thereby achieving a better dehazing effect. However, using the Transformer for dehazing also comes with certain challenges. Especially when the image has thick fog and the fog distribution is uneven, the original semantic information of the image is severely blocked, and the edge and texture information of the image are often severely lost, resulting in poor reconstruction effects. The high-frequency features of the fog-free image can be used first and then fused with its reconstructed image in the frequency domain to obtain the edge information and texture information that are difficult to reconstruct by the image itself, making the dehazed image closer to the fog-free image. Moreover, the advantages of the Transformer are accompanied by an increase in its computational cost. Due to its large model parameter scale, it leads to high computational energy consumption, especially in emergency scenarios that require efficient and real-time processing, this problem is particularly prominent. Therefore, it is particularly important to find a low-computation and high-precision image dehazing method suitable for emergencies. Summary of the Invention
[0003] Aiming at the above deficiencies in the prior art, the present invention provides a method for dehazing emergency images based on frequency domain features of wavelet transform, which solves the problem of occlusion of the original natural environment caused by fog during emergencies. The dehazing network eliminates fog to accurately, quickly, and clearly obtain the state of the environment in the current emergency.
[0004] In order to achieve the above objectives, the technical solution adopted by the present invention is: A method for dehazing emergency images based on frequency domain features of wavelet transform, comprising the following steps:
[0005] S1. Obtain virtual defogging data and real defogging data respectively, and construct an emergency defogging data set;
[0006] S2. Use the emergency defogging data set to train an image defogging model guided by high-frequency feature data;
[0007] S3. Use the trained image defogging model to perform defogging processing on emergency images.
[0008] Further, the S1 includes:
[0009] Obtain fog-free images in multiple scenarios, and generate corresponding foggy images by performing non-uniform depth fogging processing on the fog-free images, and form an image pair of virtual foggy images and fog-free images with the original fog-free images, that is, construct virtual defogging data;
[0010] Screen out outdoor images from the real image defogging data set, and construct an image pair of real foggy images and fog-free images, that is, construct real defogging data;
[0011] Mix the virtual defogging data and the real defogging data in proportion and cross them to form an emergency defogging data set.
[0012] Still further, the image defogging model includes:
[0013] An input layer, which is used to input the fog-free image and the foggy image into a single-branch image defogging model in sequence based on the emergency defogging data set. Among them, through a training strategy, the fog-free image and the foggy image input in each batch are constrained by a loss function to obtain the mapping relationship between the fog-free image and the foggy image;
[0014] A frequency domain feature extraction module, which is used to obtain the low-frequency feature LL of the fog-free image and the high-frequency features LH, HL, and HH of the fog-free image by using wavelet transform. Among them, the high-frequency features LH, HL, and HH of the fog-free image are saved to a feature list, and the low-frequency feature LL of the fog-free image is used to reconstruct the fog-free image; the high-frequency features of the foggy image are not saved, and the low-frequency feature of the foggy image is used to reconstruct the defogged image;
[0015] A defogging module, which is used to downsample the low-frequency feature LL of the fog-free image, and at the same time implement self-attention calculation within a local window based on the multi-head self-attention mechanism of Taylor expansion;
[0016] A downsampling module, which is used to adjust the resolution of the downsampled fog-free image by using wavelet convolution to make it reach the lowest resolution;
[0017] An upsampling module, which is used to upsample the fog-free image and fuse the upsampled features with the downsampled features;
[0018] The frequency-domain information reconstruction module is used to perform an inverse wavelet transform on the high-frequency features LH, HL, and HH of the haze-free image as prior data and the low-frequency features of the reconstructed haze-removed image based on the haze-free image processed by the upsampling module, so as to restore the texture information;
[0019] The physical information reconstruction module is used to cut the reconstructed haze-removed image processed by the frequency-domain information reconstruction module in the channel dimension and simulate the global atmospheric light using the cut one-dimensional tensor Simulate the transfer function with a three-dimensional tensor to calculate the reconstructed haze-removed image through the physical model prior;
[0020] The image restoration module is used to obtain the multi-level image features of each branch through the encoder for the reconstructed haze-removed image output by the frequency-domain information reconstruction module and the reconstructed haze-removed image output by the physical information reconstruction module, fuse the multi-level image features of each branch with a learnable weight, and then use the decoder to restore the final haze-removed image to achieve a reconstructed haze-removed image with higher perceptual quality.
[0021] Furthermore, the expression of the high-frequency features of the haze-free image is as follows:
[0022]
[0023]
[0024] where represents the high-frequency features of the haze-free image, represents the convolution operation, represents the low-pass filter, represents the high-pass filter, represents the input haze-free image;
[0025] The expression of the inverse wavelet transform is as follows:
[0026]
[0027] where represents the transposed convolution operation.
[0028] Furthermore, the expression of the loss function of the image dehazing model is as follows:
[0029]
[0030]
[0031]
[0032]
[0033]
[0034] Among them, represents the combined loss function of the image defogging model, represents the reconstruction loss, represents the overall structural similarity loss, , , and all represent weights, represents the structural similarity loss of the reconstructed defogged image, represents the structural similarity of the reconstructed fog-free image, represents the structural similarity loss adopted for the reconstructed fog-free image, represents the structural similarity, represents the original image, represents the reconstructed image, represents the reconstructed defogging loss, represents the reconstructed fog-free loss, N represents the total number of samples, represents the i-th true value, represents the i-th predicted value.
[0035] Furthermore, the way of the Taylor expansion is as follows:
[0036]
[0037] Among them, represents the generalized attention, represents the Taylor expansion formula, represents the multi-head attention operation, , and respectively represent the Query, Key, and Value of the i-th row. Query represents the query vector, Key represents the key vector, Value represents the actual feature value, j represents the relative position index related to i, N represents the sum of the elements of the input sequence, represents the value information of the j-th row, T represents the matrix transpose, represents the vector of the j-th row , represents the vector of the j-th row .
[0038] Furthermore, the S3 includes the following steps:
[0039] S301. Based on the emergency dehazing dataset, input the haze-free images and hazy images into the single-branch image dehazing model in sequence. Among them, through the training strategy, the haze-free images and hazy images input in each batch are constrained by the loss function to obtain the mapping relationship between the haze-free images and the hazy images;
[0040] S302. Use wavelet transform to obtain the low-frequency feature LL of the haze-free image, as well as the high-frequency features LH, HL, and HH of the haze-free image. Among them, the high-frequency features LH, HL, and HH of the haze-free image are saved to the feature list, and the low-frequency feature LL of the haze-free image is used to reconstruct the haze-free image; the high-frequency features of the hazy image are not saved, and the low-frequency feature of the hazy image is used to reconstruct the dehazed image;
[0041] S303. Downsample the low-frequency feature LL of the haze-free image, and at the same time implement self-attention calculation within the local window based on the multi-head self-attention mechanism of Taylor expansion;
[0042] S304. According to the downsampled haze-free image, use wavelet convolution to adjust the resolution to the lowest resolution;
[0043] S305. Upsample the haze-free image, and fuse the upsampled features with the downsampled features;
[0044] S306. Based on the haze-free image processed by the upsampling module, perform inverse fusion in the frequency domain between the high-frequency features LH, HL, and HH of the haze-free image and the low-frequency feature of the hazy image;
[0045] S307. Cut the reconstructed dehazed image processed by S306 in the channel dimension, and use the cut one-dimensional tensor to simulate the global atmospheric light Simulate the transfer function with a three-dimensional tensor , so as to calculate the reconstructed dehazed image through the physical model prior;
[0046] S308. For the image obtained through S306 and the reconstructed haze-free image obtained through S307 processing, obtain the multi-level image features of each branch through the encoder, fuse the multi-level image features of each branch with a learnable weight, and then use the decoder to restore the final dehazed image to achieve a reconstructed dehazed image with higher perceptual quality.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows: 1) In order to extract high-frequency features and reduce the constraints of the physical model on the reconstructed image, the training strategy of the image dehazing model WT-DehazeFormer is optimized to reduce the risk of model instability; 2) The Softmax calculation of the multi-head self-attention is simplified through Taylor expansion to reduce the number of parameters in the model network structure; 3) By adding a frequency-domain feature extraction module and an inverse fusion module to the image dehazing model WT-DehazeFormer, high-frequency data guidance is provided to the model to improve the quality of the reconstructed image of the model; 4) Through the image restoration module, it is prevented that the reconstructed image guided by high-frequency data overly emphasizes high-frequency features, thereby improving the visual effect of the reconstructed image of the model; 5) In order to apply the training strategy of the present invention, the loss function of the image dehazing model WT-DehazeFormer is optimized to improve the quality and stability of the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a flowchart of the method of the present invention.
[0049] Figure 2 It is a network structure diagram of the image dehazing model WT-DehazeFormer of the present invention.
[0050] Figure 3 It is a diagram showing the training strategy of the image dehazing model WT-DehazeFormer of the present invention.
[0051] Figure 4 It is a network structure diagram of the new multi-head self-attention in the DehazeBlock of the present invention.
[0052] Figure 5 It is a diagram showing the LL, LH, HL, and HH features of the original haze-free image and the original hazy image of the present invention.
[0053] Figure 6 It is a network structure diagram of the image restoration module Image Rest of the present invention.
[0054] Figure 7 It is a diagram showing the dehazed images of potholes on the emergency event dataset of the present invention.
[0055] Figure 8 It is a diagram showing the dehazed images of fires on the emergency event dataset of the present invention.
[0056] Figure 9 It is a diagram showing the dehazed images of flying birds on the emergency event dataset of the present invention.
[0057] Figure 10 It is another diagram showing the dehazed images of flying birds on the emergency event dataset of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0058] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0059] Embodiment
[0060] As Figures 1 - 2 shown, the present invention provides a method for removing haze from emergency event images based on the frequency domain features of wavelet transform, and its implementation method is as follows:
[0061] S1. Obtain virtual dehazing data and real dehazing data respectively, and construct an emergency event dehazing data set, which is specifically:
[0062] Obtain haze-free images in multiple scenarios, and generate corresponding hazy images by performing non-uniform depth fogging processing on the haze-free images, and form an image pair of virtual hazy images and haze-free images with the original haze-free images, that is, construct virtual dehazing data;
[0063] Screen out outdoor images from the real image dehazing data set, and construct an image pair of real hazy images and haze-free images, that is, construct real dehazing data;
[0064] Mix the virtual dehazing data and the real dehazing data crosswise in proportion to form an emergency event dehazing data set;
[0065] S2. Use the emergency event dehazing data set to train an image dehazing model guided by high-frequency feature data. The image dehazing model includes:
[0066] An input layer, which is used to input the haze-free image and the hazy image into a single-branch image dehazing model in sequence based on the emergency event dehazing data set. Among them, through the training strategy, the haze-free image and the hazy image input in each batch are constrained by the loss function to obtain the mapping relationship between the haze-free image and the hazy image;
[0067] A frequency domain feature extraction module, which is used to obtain the low-frequency feature LL of the haze-free image and the high-frequency features LH, HL, and HH of the haze-free image by using wavelet transform. Among them, the high-frequency features LH, HL, and HH of the haze-free image are saved to the feature list, and the low-frequency feature LL of the haze-free image is used to reconstruct the haze-free image; the high-frequency features of the hazy image are not saved, and the low-frequency feature of the hazy image is used to reconstruct the dehazed image;
[0068] A defogging module, which is used to downsample the low-frequency features LL of the fog-free image, and at the same time, based on the multi-head self-attention mechanism of Taylor expansion, realizes self-attention calculation within a local window;
[0069] A downsampling module, which is used to adjust the resolution of the fog-free image processed by downsampling by using wavelet convolution to make it reach the lowest resolution;
[0070] An upsampling module, which is used to upsample the fog-free image and fuse the upsampled features with the downsampled features;
[0071] A frequency-domain information reconstruction module, which is used to perform inverse wavelet transform on the high-frequency features LH, HL, and HH of the fog-free image as prior data and the low-frequency features of the reconstructed defogged image based on the fog-free image processed by the upsampling module to realize the restoration of texture information;
[0072] A physical information reconstruction module, which is used to cut the reconstructed defogged image processed by the frequency-domain information reconstruction module in the channel dimension and use the cut one-dimensional tensor to simulate the global atmospheric light Use a three-dimensional tensor to simulate the transfer function , so as to calculate the reconstructed defogged image through the physical model prior;
[0073] An image restoration module, which is used to obtain the multi-level image features of each branch through the encoder for the reconstructed defogged image output by the frequency-domain information reconstruction module and the reconstructed defogged image output by the physical information reconstruction module, fuse the multi-level image features of each branch with a learnable weight, and then use the decoder to restore the final defogged image to achieve a reconstructed defogged image with higher perceptual quality.
[0074] In this embodiment, as Figure 2As shown in the figure, the fusion module in the upsampling stage is the SKF module, which is an image feature fusion module. Its purpose is to obtain the feature information in the shallow layer of the U-Net structure, and it is an operation that fuses the deep features and shallow features of a traditional U-Net network structure. The purpose of the frequency domain feature fusion module WIT is to fuse the high-frequency features LH, HL, and HH of the original haze-free image with the low-frequency feature LL of the hazy image passing through the dehazing network in the frequency domain using inverse wavelet transform, prevent the loss of high-frequency information, and finely reconstruct the dehazed image. The physical information reconstruction module is the P-Piror Recon module in the above figure, which uses the atmospheric scattering model to obtain the reconstructed dehazed image, and verifies that the dehazing model of the present invention still has an effect on the physical model of real image dehazing. The final image restoration module, namely the Image Real module, aims to fuse the output of the frequency domain feature fusion module (reconstructed dehazed image) with the output of the physical information reconstruction module (reconstructed dehazed image). Through the encoder-decoder structure, the image features of the image passing through the frequency domain feature fusion module and the image passing through the physical information reconstruction module are used to obtain the multi-level image features of each branch through the encoder structure, and the multi-level image features of each branch are fused with a learnable weight, and then the decoder structure is used to restore the final dehazed image to achieve a reconstructed dehazed image with higher perceptual quality.
[0075] S3. Use the trained image dehazing model to perform dehazing processing on the emergency event image, and its implementation method is as follows:
[0076] S301. Based on the emergency event dehazing data set, input the haze-free image and the hazy image into the single-branch image dehazing model in sequence. Among them, through the training strategy, the mapping relationship between the haze-free image and the hazy image is obtained by constraining the haze-free image and the hazy image input in each batch through the loss function.
[0077] S302. Use wavelet transform to obtain the low-frequency feature LL of the haze-free image, as well as the high-frequency features LH, HL, and HH of the haze-free image. Among them, the high-frequency features LH, HL, and HH of the haze-free image are saved to the feature list, and the low-frequency feature LL of the haze-free image is used to reconstruct the haze-free image; the high-frequency features of the hazy image are not saved, and the low-frequency feature of the hazy image is used to reconstruct the dehazed image.
[0078] S303. Downsample the low-frequency feature LL of the haze-free image, and at the same time, implement self-attention calculation within the local window based on the multi-head self-attention mechanism of Taylor expansion.
[0079] S304. According to the downsampled haze-free image, use wavelet convolution to adjust the resolution to the lowest resolution.
[0080] S305. Upsample the haze-free image, and fuse the upsampled features with the downsampled features.
[0081] S306. Based on the fog-free image processed by the upsampling module, perform inverse fusion in the frequency domain on the high-frequency features LH, HL, and HH of the fog-free image and the low-frequency features of the foggy image;
[0082] S307. Cut the reconstructed fog-removed image processed by S306 in the channel dimension, and use the cut one-dimensional tensor to simulate the global atmospheric light Simulate the transfer function with a three-dimensional tensor , so as to calculate the reconstructed fog-removed image through the physical model prior;
[0083] S308. For the image obtained through S306 and the reconstructed fog-free image obtained through S307 processing, obtain the multi-level image features of each branch through the encoder, fuse the multi-level image features of each branch with a learnable weight, and then use the decoder to restore the final fog-removed image, so as to achieve a reconstructed fog-removed image with higher perceptual quality.
[0084] In this embodiment, to ensure the image quality of the self-built emergency fog-removal dataset, the emergency data constructed by the present invention covers fog-free images in various scenarios such as potholes, fires, and flying birds. By performing non-uniform depth fogging on the fog-free images, corresponding foggy images are generated, and then the original fog-free images are combined to form an image pair of virtual foggy images and fog-free images (virtual fog-removal data). In addition, high-quality outdoor images with rich texture information are selected from real image fog-removal datasets such as RESIDE, NH-HAZE, and DENSE-HAZE to form an image pair of real foggy images and fog-free images (real fog-removal data). The virtual fog-removal data and the real fog-removal data are cross-mixed in a ratio of 3:1 to form an emergency fog-removal dataset. And this dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1. After that, the divided dataset is input into the Transformer fog-removal network model WT-DehazeFormer containing a high-frequency feature module through a series of technical means such as image enhancement, image augmentation, and image resolution adjustment for training, which can significantly improve the training efficiency of the model. And the final model obtained can quickly and accurately remove the fog occlusion of the real scene during the occurrence of an emergency and restore the real scene state. The optimized fog-removal result can also assist the target detection task and improve the detection accuracy.
[0085] In this embodiment, the present invention designs a Transformer neural network model based on a combined frequency domain feature module - WT-DehazeFormer. Its core is to use wavelet transform to obtain the high-frequency features of the fog-free image, perform data-guided reconstruction to obtain a dehazed image with higher quality, and perform Taylor expansion on the softmax calculation of the multi-head self-attention in the traditional Transformer structure, effectively reducing the number of model parameters. This model has a single-input and single-output structure, but by modifying the training strategy, both the foggy image and the fog-free image are input into the network to obtain the foggy-fog-free mapping relationship. In addition, a combined loss function is designed to match its training strategy. And an image restoration module is used to obtain a dehazed image that is more in line with the human visual effect. This dehazing network model improves the deficiencies such as the loss of edge and texture information in the reconstructed image, not only optimizing the image quality of the reconstructed image, preventing the loss of important semantic information, but also reducing the number of model parameters, thereby accelerating the inference speed of the model in emergencies, and at the same time improving the stability and convergence speed of the network during training.
[0086] In this embodiment, in terms of the training strategy of the network model, the present invention makes the following optimizations: The traditional single-branch dehazing network only inputs the foggy image into the network, and the fog-free image acts as a label for loss calculation. If trained in the traditional training method, the high-frequency features of the fog-free image cannot be obtained, and it is impossible to use it as a guiding data to guide the model to reconstruct a dehazed image with higher quality. The present invention inputs the foggy image and the fog-free image into the single-branch dehazing network WT-DehazeFormer in sequence, that is, two images (foggy image - fog-free image) are trained in each epoch. Through this training strategy, the high-frequency features of the fog-free image can be successfully obtained, and the model can reflect the dehazing process, that is, the mapping relationship from the foggy image to the fog-free image. Through this mapping relationship, a clear supervision signal can be given to the model, and the network can directly learn the conversion law between the foggy image and the fog-free image, so that the network can clearly learn what the target output of dehazing is, rather than just inferring the dehazing effect through the loss function. And it can improve the dehazing effect of the model, because the network not only needs to remove haze from the foggy image, but also needs to maintain a structure and details similar to the fog-free image, which has a great promoting effect on the training effect of the dehazing model. The most important thing is to avoid the dehazing model relying too much on the prior physical model, that is, the atmospheric scattering model, and the atmospheric scattering model is defined as follows:
[0087]
[0088] where x represents the pixel position, represents the foggy image, represents the restored fog-free image, Represents the transfer function, i.e., the influence of particles on the true target image. Its physical meaning is the proportion of light that can reach the detection system after particle attenuation. Represents the global atmospheric light.
[0089] The transfer function is defined as follows:
[0090]
[0091] Among them, Represents the atmospheric scattering coefficient, Represents the scene depth.
[0092] By deriving the atmospheric scattering model, we can obtain , which is defined as follows:
[0093]
[0094] Among them, Represents the atmospheric scattering coefficient, Represents the scene depth.
[0095] The network structure of the image dehazing model WT-DehazeFormer of the present invention estimates and , and derives through the atmospheric scattering model. To avoid over-reliance on the physical model and ineffective dehazing in some special cases, through the network model training strategy of the present invention, the dependence on the physical model can be effectively reduced, and at the same time, attention is paid to learning the dehazing characteristics through the data itself. In terms of the dehazing network training strategy of the image dehazing model WT-DehazeFormer, the training strategy of the present invention is as Figure 3 shown.
[0096] In this embodiment, in order to apply the new network model training strategy, in terms of the loss function, the present invention makes the following optimizations: using the combined loss function to assign different weights to the haze-free loss of reconstruction and the dehazing loss of reconstruction to constrain the reconstruction loss. This improvement not only enhances the learning ability of the network, reconstructs haze-free images with higher perceptual quality, but also can express the mapping relationship from hazy images to haze-free images, reduces the dependence on the physical model, and improves the dehazing ability of the model in special cases.
[0097] To determine the quality of the reconstructed dehazed image in the combined loss, the dehazing loss of reconstruction (hereinafter using The (hereinafter referred to as is defined as follows:
[0098]
[0099] the loss for reconstructing a haze-free image (hereinafter referred to as adopts the L1 loss. Since the difference between the reconstructed haze-free image and the original haze-free image may be large, the difference between two pixel points is large. The L1 loss can ensure that its gradient value will not be too large to cause gradient explosion, and the L1 loss is not only simple to calculate but also can preserve the sharp edges and textures in the image, avoiding the image from becoming too blurred. is defined as follows:
[0100]
[0101] In order to reconstruct a high-quality haze-free image and reflect the mapping relationship from hazy to haze-free, the present invention uses different weights to constrain the total reconstruction loss (hereinafter referred to as ), which is defined as follows:
[0102]
[0103] And in order to better reflect the human eye's perception of image quality, especially the changes in details such as textures and edges, the present invention uses the structural similarity loss (hereinafter referred to as ) for both the reconstructed haze-removed image and the reconstructed haze-free image. It is defined as follows:
[0104]
[0105] Then, the structural similarity loss of the reconstructed haze-removed image (hereinafter referred to as ) and the structural similarity of the reconstructed haze-free image (hereinafter referred to as ) use different weights to constrain the total structural similarity loss (hereinafter referred to as ), which is defined as follows:
[0106]
[0107] In summary, the loss function used in the Transformer haze removal network framework of the present invention is:
[0108]
[0109] wherein, Represents the combined loss function of the image defogging model, Represents the reconstruction loss, Represents the overall structural similarity loss, , , and All represent weights, Represents the structural similarity loss of the reconstructed defogged image, Represents the structural similarity of the reconstructed haze-free image, Represents the structural similarity loss adopted for the reconstructed haze-free image, Represents the structural similarity, Represents the original image, Represents the reconstructed image, Represents the reconstructed defogging loss, Represents the reconstructed haze-free loss, and N represents the total number of samples, Represents the i-th true value, Represents the i-th predicted value.
[0110] In this embodiment, the present invention optimizes the self-attention computational complexity caused by Softmax in the multi-head attention mechanism, and proposes to simplify the calculation of Softmax by using Taylor expansion, which can effectively reduce the number of model parameters and enable the defogging model to quickly and accurately remove the fog occlusion of the real scene during emergencies. The traditional attention calculation formula is as follows:
[0111]
[0112] Assume that the length and width of the input image are h and w, and D is a sequence of h × w feature vectors in its D dimensions, then , , , so the calculation of self-attention by Softmax will have a square term, which greatly increases the computational amount. In order to remove the square root and reduce the computational amount, the calculation formula of the generalized attention is written as follows:
[0113]
[0114] Among them, the matrix with subscript i represents the vector of its i-th row, Represents any similarity function. The calculation formula obtained by performing the first-order Taylor expansion on the generalized attention is as follows:
[0115]
[0116] Among them, Represents The maximum infinitesimal term obtained by performing Taylor expansion, that is, the Peano remainder term.
[0117] From the normalization of and vectors, we can obtain and approximate . When and have norms less than 1, the values of the attention map can be made positive. And in practice, it is found that normalizing the norm to 0.5 can obtain the best results. Also, is very similar to the x + 1 function between [-0.25, 0.25]. Then, in the domain of [-0.25, 0.25], eliminating the remainder form of the Peano remainder term in its first-order Taylor expansion, the calculation formula for the Taylor expansion of self-attention is as follows:
[0118]
[0119] Finally, applying the matrix multiplication associative law to the final formula obtained by calculating in the equation, the calculation formula is as follows:
[0120]
[0121] where represents generalized attention, represents the Taylor expansion, represents the multi-head attention operation, , and represent the Query, Key, and Value of the i-th row respectively. Query represents the query vector, Key represents the key vector, Value represents the actual eigenvalue, j represents the relative position index related to i, N represents the total number of elements in the input sequence. represents the value information of the j-th row, T represents matrix transpose, represents the vector of the j-th row , represents the vector of the j-th row .
[0122] Through this way of Taylor expansion, the quadratic term caused by the original Softmax calculation can be reduced to a linear term, greatly reducing the calculation parameters. Because the Peano remainder is ignored in the above calculation formula reasoning, some errors will inevitably occur. Due to the Peano remainder being an infinitesimal of order n, the matrix multiplication associative law cannot be used to linearize the calculation complexity of multi-head self-attention. However, the generation of the Peano remainder term is related to the Q and K matrices. Considering that the image has local correlation, by learning the local information of the Q and K matrices, the inaccurate output can be corrected. . In terms of using Taylor expansion to reduce the calculation parameters of the multi-head self-attention mechanism, the multi-head self-attention mechanism module of the present invention is as shown in Figure 4 .
[0123] In this embodiment, the present invention aims at the problem that the semantic information of the fog-shielded part in the original foggy image is not completely reconstructed for the reconstructed dehazed image, resulting in a large amount of edge and texture information loss. Especially in the dehazing task of emergencies, important semantic information in the real environment is easily lost. The present invention designs a frequency-domain feature extraction module and an inverse fusion module. The purpose is that the frequency-domain feature extraction module extracts the high-frequency features of the original haze-free image through wavelet transform, and in the inverse fusion module, the reconstructed dehazed image is fused with the high-frequency features of its original haze-free image through inverse wavelet transform, preventing the loss of the features of the reconstructed haze-free image and providing data guidance for the reconstructed image. The specific module optimization process is as follows:
[0124] In the frequency-domain feature extraction module, Symlet wavelets (specifically, the applied wavelet is Sym2) are used to extract low-frequency and high-frequency features through a low-pass filter and a group of high-pass filters. Because Symlet wavelets have a smoother shape and are more suitable for smoothing operations in images, they can avoid the over-ringing phenomenon caused by asymmetry. Symlet wavelets have good effects in retaining the low-frequency information of the image (hereinafter represented by LL) and details, and also in the high-frequency part, that is, the texture information and edge information in the horizontal, vertical, and diagonal directions of the image (hereinafter represented by LH, HL, and HH). The frequency-domain feature extraction in a 2D image can be described as follows:
[0125]
[0126] The specific process is as follows:
[0127]
[0128] Among them, represents the high-frequency features of the extracted haze-free image, represents the convolution operation, represents the low-pass filter, represents the high-pass filter, represents the input haze-free image.
[0129] In the inverse fusion module, the present invention performs inverse wavelet transform on the reconstructed dehazed image and the high-frequency features (LH, HL, and HH) of the original haze-free image extracted in the frequency-domain feature extraction module. The purpose is to perform fusion in the frequency domain, highlight the high-frequency features, make the reconstructed image more similar to the original haze-free image, and improve the image quality of the reconstructed image. The specific formula for the inverse transform is as follows:
[0130]
[0131] The LL, LH, HL, and HH features of the original haze-free image and the original hazy image are as follows Figure 5 shown
[0132] In this embodiment, aiming at the advantages of using frequency-domain information for image dehazing in the Transformer architecture, the present invention proposes the following specific dehazing network process:
[0133] Obtain the high-frequency features of the haze-free image. Input the haze-free image into the frequency-domain feature extraction module, that is, obtain the features LL, LH, HL, and HH of the haze-free image through wavelet transform. Among them, LH, HL, and HH are high-frequency features, and LL is a low-frequency feature
[0134] Input the low-frequency feature LL extracted by the frequency-domain feature extraction module into the image dehazing model WT-DehazeFormer network structure of the present invention (which can be regarded as a U-Net type network structure. The U-Net structure is that the features of the first downsampling are fused with the last but one upsampling, and the features of the second downsampling are fused with the features of the second last upsampling). Because a large number of studies have shown that in traditional CNN networks, its convolutional layers usually tend to respond to high-frequency features in the input, while in the Transformer structure, its attention heads are considered to be more suitable for low-frequency features
[0135] Next, adjust its input images (haze-free image and hazy image, each batch bath_size has two inputs, the number of channels of (haze-free - hazy)) through ordinary convolutional layers and enter the dehazing module (hereinafter represented by DehazeBlock) in the image dehazing model WT-DehazeFormer. In the dehazing module DehazeBlock, the Patch Merging image block merging technology is used for downsampling. At the same time, a multi-head self-attention mechanism based on Taylor expansion proposed by the present invention is combined. Through a series of feed-forward networks (MLP structure), residual connections, and layer normalization, self-attention calculation within a local window is realized. This design can not only effectively capture local and cross-window feature dependencies, but also process and transform feature representations, thereby enhancing the connectivity and feature transfer ability of the network. Through detailed self-attention calculation within a local window, the system can more accurately identify and utilize the detailed information in the image, while the capture of cross-window feature dependencies improves the overall feature coordination and consistency. The introduction of the feed-forward network further optimizes the non-linear transformation of features and improves the expression ability of the model. The application of residual connections effectively alleviates the gradient disappearance problem in deep networks and ensures the smooth transmission of information in the network. Layer normalization stabilizes the training process in each layer and accelerates the convergence speed of the model. The PatchMerging downsampling formula is as follows
[0136] First, the feature map is divided into Patch image patches, divided into non-overlapping small blocks of 2×2, and feature splicing is performed. The feature splicing formula is as follows:
[0137]
[0138] Among them, represents the spliced feature, represents the feature vectors of the four Patch image patches and then performs a linear transformation on them. The formula for the linear transformation is as follows:
[0139]
[0140] Among them, W represents the weight matrix and b represents the bias.
[0141] To sum up, the overall formula is as follows:
[0142]
[0143] Among them, represents the output, represents the linear transformation.
[0144] After that, the image resolution is adjusted through the downsampling module (hereinafter represented by Down-sample). Among them, WTConv wavelet convolution is used, and the selected wavelet is Biorthogonal wavelet (the specific wavelet used is bior1.3). Because the Biorthogonal wavelet has the ability of distortion-free reconstruction, the original image can be perfectly restored through decomposition and reconstruction, and the Biorthogonal wavelet has high flexibility and can adjust the characteristics of each pair of filters as needed. WTConv wavelet convolution has a higher receptive field than traditional convolution. The traditional convolutional layer obtains a larger receptive field by increasing the convolutional kernel of the convolution, but although the receptive field increases, a large amount of image information is lost. However, WTConv wavelet convolution can obtain a larger receptive field with a smaller convolutional kernel, which not only ensures the integrity of the image information but also increases the receptive field.
[0145] After reaching the lowest resolution part through the downsampling module Down-sample, the feature map is gradually upsampled and fused with the feature map in the downsampling stage, improving the feature representation effect. Among them, the upsampling module (hereinafter represented by Up-sample) uses depthwise separable convolution to further reduce the number of parameters, and a SKF feature fusion module is tightly connected after each upsampling module Up-sample to ensure the fusion of the upsampled features with the previous features, so as to restore high-resolution information and maintain the consistency of multi-scale features, and can better capture the key information in the image by fusing global context information, enhancing the discriminability of the features.
[0146] After the fog-free image is restored to the same resolution as the input through the upsampling module Up-sample, the dehazed image is reconstructed through a soft reconstruction layer (hereinafter represented by Soft Recon). After that, the reconstructed dehazed image is unfolded one-dimensionally and its features are divided into two parts to simulate the and (introduced in the foregoing) calculate the reconstructed dehazed image through the physical model prior.
[0147] In the present invention, the high-frequency features (LH, HL, HH) of the original fog-free image are inversely fused with the reconstructed dehazed image in the frequency domain to obtain a reconstructed dehazed image with higher perceptual quality.
[0148] In this embodiment, since the high-frequency part is too prominent after fusing the reconstructed image and the high-frequency features in the frequency domain, the visual effect is abrupt. The present invention inputs the image not reconstructed by the physical model and the reconstructed image fused with high-frequency features into a lightweight multi-branch integrated reconstruction network, that is, an image restoration module (hereinafter represented by the Image Rest module), to synthesize the weights of the high-frequency features in the reconstructed image, so that the final reconstructed image is closer to the real image and improves the visual consistency and naturalness. Its image restoration module is mainly an encoder-decoder structure. The encoder (a simple convolutional structure) is used to obtain the multi-level image features of each branch. After that, the multi-level image features of each branch are fused to achieve targeted fusion of different features. Finally, the decoder is used to restore a clearer, natural, and dehazed image that meets the visual effect. The network structure of the image restoration module is as Figure 6 shown, as Figures 7 - 10 shown shown, Figures 7 - 10 is a display diagram of dehazed images of potholes, fires, and birds.
[0149] In this embodiment, two evaluation metrics, peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), are used for evaluation. The PSNR formula is as follows:
[0150]
[0151] Among them, MAX represents the maximum possible value of the image pixels, and MSE represents the mean square error. The formula is as follows:
[0152] Among them, respectively represent the gray values or colors of the original image and the reconstructed image at the pixel position , and M and N respectively represent the height and width of the image.
[0153] The structural similarity SSIM formula is as follows:
[0154]
[0155] The formula for luminance similarity is as follows:
[0156]
[0157] The contrast similarity is as follows:
[0158]
[0159] The structural similarity is as follows:
[0160]
[0161] In summary, the comprehensive structural similarity is as follows:
[0162]
[0163] Among them, x and y represent two image patches to be compared, and respectively represent the means of x and y in the image patch, and respectively represent the variances of x and y in the image patch, is the covariance between image patch x and y, is a constant to avoid the denominator being zero. Usually, .
[0164] Verify the performance of the image dehazing model WT-DehazeFormer network on the validation set using the above two evaluation metrics and save the model with the best performance.
[0165] In summary, the basic idea of the present invention is to propose a Transformer image dehazing model (WT-DehazeFormer) guided by high-frequency feature data. This model inherits the model framework of the traditional Swin-Tranformer moving window hierarchical Transformer structure network:
[0166] 1) Reconstruct a fog-free clean image from a foggy image through the unique window shifting mechanism (Shifted Windowing Mechanism) in Swin-Transformer and the Encoder-Decoder architecture in the traditional Transformer structure; 2) The essence of the present invention is the supervised image-to-image reconstruction problem; 3) Paired data training (foggy image - fog-free image) is required, that is, fog-free images are needed for supervised learning to estimate high-quality fog-free images from foggy images; 4) It includes Multi-Head Self-Attention, which calculates self-attention for image patches within each window using the local window mechanism. The self-attention layers within each window are independent, which can effectively reduce the number of parameters and better extract the features of the fog; 5) It includes Multi-Scale Feature Representation, which captures information at different scales by merging image patches layer by layer and changing the window size. It can not only process high-resolution images but also retain the local and global information of the image, enabling the network model to better understand the relationship between the fog features and the semantic information of the image itself in the foggy image; 6) It includes a loss: mainly used to measure the difference between the reconstructed fog-free image and the original fog-free image. Its basic form is to calculate the absolute error of each pixel to evaluate the picture quality of the reconstructed image.
Claims
1. A method for defogging emergency images based on wavelet transform frequency domain features, characterized in that: The following steps are involved: S1, obtain virtual defogging data and real defogging data respectively, and build an emergency defogging dataset; S2, using the emergency defogging dataset to train the image defogging model guided by high-frequency feature data; S3. Use the trained image defogging model to defog the emergency image.
2. The method for defogging emergency images based on wavelet transform frequency domain features according to claim 1 is characterized in that: The S1 includes: Obtain fog-free images in multiple scenes, and generate corresponding foggy images by performing non-uniform depth fogging on the fog-free images, and form a virtual foggy image and fog-free image pair with the original fog-free images, that is, construct virtual defogging data; Filter out outdoor images from the real image defogging dataset and construct image pairs of real foggy images and fog-free images, that is, construct real defogging data; The virtual dehazed data and the real dehazed data are cross-mixed in proportion to form an emergency dehazed dataset.
3. The method for defogging emergency images based on wavelet transform frequency domain features according to claim 1, characterized in that: The image dehazing model includes: The input layer is used to sequentially input the fog-free images and foggy images into the single-branch image defogging model based on the emergency defogging dataset, wherein the mapping relationship between the fog-free images and the foggy images input in each batch is obtained through the loss function constraint through the training strategy; The frequency domain feature extraction module is used to obtain the low-frequency feature LL of the fog-free image and the high-frequency features LH, HL and HH of the fog-free image by wavelet transform, wherein the high-frequency features LH, HL and HH of the fog-free image are saved in the feature list, and the low-frequency feature LL of the fog-free image is used to reconstruct the fog-free image; the high-frequency features of the foggy image are not saved, and the low-frequency features of the foggy image are used to reconstruct the defogging image; The dehazing module is used to downsample the low-frequency features LL of the haze-free image and implement self-attention calculation within the local window based on the multi-head self-attention mechanism based on Taylor expansion; The downsampling module is used to adjust the resolution of the fog-free image after downsampling by using wavelet convolution to achieve the lowest resolution; The upsampling module is used to upsample the fog-free image and fuse the upsampled features with the downsampled features; The frequency domain information reconstruction module is used to perform inverse wavelet transform on the high-frequency features LH, HL and HH of the fog-free image as prior data and the low-frequency features of the reconstructed defogging image based on the fog-free image processed by the upsampling module to restore the texture information; The physical information reconstruction module is used to cut the reconstructed dehazed image processed by the frequency domain information reconstruction module in the channel dimension, and use the cut one-dimensional tensor to simulate the global atmospheric light Modeling transfer functions with 3D tensors , to reconstruct the dehazed image through physical model prior calculation; The image restoration module is used to obtain the multi-level image features of each branch of the reconstructed dehazed image output by the frequency domain information reconstruction module and the reconstructed dehazed image output by the physical information reconstruction module through the encoder, fuse the multi-level image features of each branch with a learnable weight, and then use the decoder to restore the final dehazed image to achieve a reconstructed dehazed image with higher perceptual quality.
4. The method for defogging emergency images based on wavelet transform frequency domain features according to claim 3, characterized in that: The expression of the high-frequency features of the fog-free image is as follows: in, represents the high-frequency features of the haze-free image, represents the convolution operation, represents a low-pass filter, represents a high-pass filter, represents the input haze-free image; The expression of the inverse wavelet transform is as follows: in, Represents a transposed convolution operation.
5. The method for defogging emergency images based on wavelet transform frequency domain features according to claim 3, characterized in that: The loss function of the image dehazing model is expressed as follows: in, represents the combined loss function of the image dehazing model, represents the reconstruction loss, represents the total structural similarity loss, , , and Both represent weights, represents the structural similarity loss of reconstructed dehazed image, represents the structural similarity of the reconstructed haze-free image, represents the structural similarity loss used to reconstruct the haze-free image, Indicates the structural similarity, represents the original image, represents the reconstructed image, represents the reconstruction dehazing loss, represents the reconstruction of haze-free loss, N represents the total number of samples, represents the i-th true value, represents the i-th predicted value.
6. The method for defogging emergency images based on wavelet transform frequency domain features according to claim 3, characterized in that: The Taylor expansion is as follows: in, represents generalized attention, represents the Taylor expansion, represents the multi-head attention operation, , and Represent the Query, Key and Value of the i-th row respectively. Query represents the query vector, Key represents the key vector, Value represents the actual feature value, j represents the relative position index related to i, and N represents the total number of elements in the input sequence. represents the value information of the jth row, T represents the matrix transformation, The vector representing the jth row , The vector representing the jth row .
7. The method for defogging emergency images based on wavelet transform frequency domain features according to claim 1, characterized in that: The S3 comprises the following steps: S301, based on the emergency defogging data set, inputting the fog-free images and the foggy images into the single-branch image defogging model in sequence, wherein the mapping relationship between the fog-free images and the foggy images input in each batch is obtained through the loss function constraint by the training strategy; S302, using wavelet transform to obtain the low-frequency feature LL of the fog-free image and the high-frequency features LH, HL and HH of the fog-free image, wherein the high-frequency features LH, HL and HH of the fog-free image are saved in a feature list, and the low-frequency feature LL of the fog-free image is used to reconstruct the fog-free image; the high-frequency features of the foggy image are not saved, and the low-frequency features of the foggy image are used to reconstruct the defogged image; S303, down-sampling the low-frequency feature LL of the haze-free image, and realizing self-attention calculation in the local window based on the multi-head self-attention mechanism of Taylor expansion; S304, adjusting the resolution of the fog-free image after downsampling by using wavelet convolution to achieve the lowest resolution; S305, up-sampling the fog-free image, and fusing the up-sampled features with the down-sampled features; S306, based on the fog-free image processed by the upsampling module, inversely fuse the high-frequency features LH, HL and HH of the fog-free image with the low-frequency features of the foggy image in the frequency domain; S307, cutting the reconstructed defogging image processed by S306 in the channel dimension, and using the cut one-dimensional tensor to simulate the global atmospheric light Modeling transfer functions with 3D tensors , to reconstruct the dehazed image through physical model prior calculation; S308. The image obtained by S306 and the reconstructed fog-free image obtained by S307 are obtained by an encoder to obtain the multi-level image features of each branch, and the multi-level image features of each branch are fused with a learnable weight, and then the decoder is used to restore the final defogged image to achieve a reconstructed defogged image with higher perceptual quality.
Citation Information
Patent Citations
Image motion blur removing method based on improved U-Net model
CN114549361A
Neural network algorithm and system for electric power security control image de-atomization
CN117911275A
Visual defogging method and system based on three-branch neural network
CN118521508A
Real scene image defogging method based on dual-branch CNN-Transformer and depth information fusion
CN119205569A
Image defogging method based on improved DEA-Net
CN119722525A
Cited By
Traffic image defogging method based on wavelet convolution and semantic-content guide fusion
CN121353135A
Traffic image defogging method based on wavelet convolution and semantic-content guidance fusion
CN121353135B