Image defogging method based on generative adversarial network
By adopting a two-stage training method based on a generative adversarial network in the image defog removal method, combining multi-layer comparison loss, style transfer loss and identity consistency loss, the differences between synthetic data and real scenes and color and texture distortion problems during the defog removal process are solved, and the high-precision and natural fog removal effect is achieved.
Patent Information
- Application Number
- CN202510028973.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-08
AI Technical Summary
When the existing methods use synthetic data to train the model, the model accuracy is low due to the differences between the synthetic data and the real scene; at the same time, color and texture distortion problems are prone to occur during the fog removal process.
Using an image defog removal method based on a generative adversarial network, two training stages are designed: the first stage is used to train the defog removal network using synthetic data, and the second stage is used to further train using real data, combining multi-layer comparison loss, style transfer loss and identity consistency loss to optimize the defog removal effect.
The problem of the difference between synthetic data and real scenes is effectively solved, the adaptability and accuracy of the model in the real environment is improved, the color and texture distortion problems are avoided, and the naturalness and accuracy of the defogging image is significantly improved.
Smart Images

Figure CN119941570A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to an image defogging method based on a generative adversarial network. Background Art
[0002] The research and development of image defogging technology aims to solve the problem of weakened visual information caused by haze. The scattering effect of tiny water droplets in the fog on light will cause image acquisition to be impaired, resulting in information loss and reduced brightness, contrast and overall visibility. These factors have seriously affected the performance of the image processing system. The core goal of image defogging technology is to restore a clear and visible image from an image affected by fog. In safety-sensitive application areas, such as in intelligent transportation systems, extreme meteorological conditions, especially haze weather, water mist particles suspended in the atmosphere significantly reduce the driver's visual judgment ability of the road conditions ahead through the car window, increasing the potential risk of driving. Image defogging technology is not only to improve image quality, but more importantly to ensure the accuracy and availability of image information, thereby supporting relevant systems to make effective decisions.
[0003] Traditional image dehazing methods mainly rely on non-physical models and use image enhancement techniques (such as histogram equalization and local contrast enhancement) to visually improve fog-affected images. However, these methods often ignore the physical formation process of fog, especially the concentration and distribution of fog and its complex interaction with ambient light. As a result, although there is some visual improvement, it is impossible to fundamentally restore the real scene of the image, and sometimes it will cause the loss of image details or color distortion.
[0004] In recent years, defogging algorithms based on physical models have begun to be studied and applied. These algorithms attempt to infer and restore the true state of fog-free images by establishing mathematical models of atmospheric scattering and ambient light. For example, the atmospheric scattering model assumes that fog is caused by tiny water droplets or ice crystals in the air that affect the propagation of light through scattering. By estimating ambient light and scattering parameters, the model can be used to restore the color and contrast of the image to its original state.
[0005] The introduction of deep learning technology has brought revolutionary progress in the field of image defogging. By constructing a deep neural network, the model can be trained to automatically learn the nonlinear mapping relationship between foggy images and clear images. This method uses the power of big data. Through the training of large-scale image data sets, the network can identify and simulate the impact of fog and achieve more accurate image restoration. In the prior art, some scholars use innovative end-to-end learning methods to directly learn and estimate the mapping relationship between blurred image blocks and their transmission maps from foggy images. This method uses a deep convolutional neural network structure to automatically extract image features and generate defogging images by learning a large number of foggy and non-fogging image pairs, which greatly simplifies the complex preprocessing and parameter adjustment process in traditional defogging technology. There is also a prior art that proposes a multi-scale convolutional neural network (MSCNN) to improve the image defogging effect by introducing a multi-scale structure. The network design includes processing units of multiple scales, each of which is responsible for capturing image details at different scales, and optimizing image quality layer by layer from coarse to fine. This hierarchical processing method enables MSCNN to more effectively handle the effects of different degrees of haze while retaining more image details.
[0006] In addition, the use of generative adversarial networks (GANs) has also shown great potential in improving the naturalness and detail of dehazed images. In this framework, the generative network is responsible for generating dehazed images, while the adversarial network attempts to distinguish the difference between dehazed images and real non-hazed images. Through this adversarial process, the dehazing model can generate more natural and delicate images. The prior art proposes a conditional generative adversarial network (cGAN), which successfully achieves the removal of haze by learning the process of mapping foggy images to their clear corresponding images. The Cycle-Dehaze network is also proposed, which improves the recovery of texture information by introducing cycle consistency and perceptual loss techniques. In addition, an innovative domain adaptive dehazing (DAD) method uses an image transfer model to effectively convert simulated foggy images into realistic images, thereby significantly improving the dehazing effect in actual scenes. Some scholars have also proposed a synthetic to real dehazing framework (PSD) based on physical priors. This improved method adapts the existing dehazing models to practical applications by using traditional priors or principles (such as dark channels) to guide the transition from synthetic to real. In addition, an unpaired dehazing framework called D4 was introduced to enhance the training effect of the dehazing network by mining the scattering coefficient and depth information in foggy and clear images.
[0007] However, in real environments, especially dynamically changing outdoor scenes, it is difficult to obtain foggy and fog-free images of the same scene at the same time, which limits the coverage and quality of training data. The scarcity and difficulty of data acquisition make the model training insufficient and difficult to adapt to the changing needs of practical applications. Therefore, many defogging models rely on synthetic data for training. These data create fog effects by applying atmospheric scattering models on clean images. However, this method often does not match the real haze conditions in the simulation of atmospheric light values and transmittance, resulting in a large deviation between the defogging effect of the model in practical applications and that in training. The difference in illumination and visual texture between synthetic fog images and real fog images makes it difficult for the model to handle real and complex haze scenes. In addition, although the generative adversarial network can generate visually convincing defogging images, in the process of directly converting foggy images to fog-free images, there are often problems with over-saturation or under-saturation of colors and unnatural texture details. These distortion problems weaken the naturalness and practical value of the defogging images, especially in application scenarios that are sensitive to color and texture details. Summary of the invention
[0008] To this end, the technical problem to be solved by the present invention is to overcome the problem that when the existing method uses synthetic data to train the model, the model accuracy is poor due to the difference between the synthetic data and the real scene; and the defects of color and texture distortion exist in the process of converting foggy images into fog-free images.
[0009] In order to solve the above technical problems, the present invention provides an image defogging method based on a generative adversarial network, comprising the following steps:
[0010] Acquire an image data set, the image data set comprising: different synthetic fog data and real fog data, the synthetic fog data comprising a fog-free image and its corresponding synthetic fog paired image, the real fog data comprising a real fog image and its corresponding real fog-free image, preprocess the image data set, and divide the preprocessed image data set into a training set and a verification set;
[0011] Input the synthetic fog paired image into the generator of the generative adversarial network to obtain the defogged image of the synthetic fog paired image;
[0012] By reconstructing the defogging image of the synthetic fog paired image, a reconstructed fog image of the synthetic fog paired image is obtained;
[0013] The sum of the smoothed L1 loss between the dehazed image and the haze-free image of the synthetic haze paired image and the smoothed L1 loss between the reconstructed haze image and the synthetic haze paired image is used as the total loss function of the first stage;
[0014] The generator is trained through the total loss function of the first stage to obtain the generator after the first stage training;
[0015] Input the real fog image into the generator to obtain the feature map of the real fog image and the defogging image of the real fog image;
[0016] Input the real fog-free image into the generator to obtain the dehazed image of the real fog-free image;
[0017] The defogging image of the real foggy image is input into the post-feature extractor to obtain the defogging image features of the real foggy image;
[0018] The feature map of the real fog image is passed through a two-layer perceptron network to obtain the multi-layer features of the real fog image;
[0019] The defogging image features of the real foggy image are passed through a two-layer perceptron network to obtain multi-layer defogging image features of the real foggy image;
[0020] By calculating the contrast loss function of each feature layer of the real fog image and each spatial position in the feature layer of the dehazed image of the same layer, a multi-layer contrast loss is constructed;
[0021] Based on the first dehazed image features of the real foggy image generated by the generator trained in the first stage and the second dehazed image features of the real foggy image generated by the generator after updating parameters in the second stage, a style transfer loss is constructed;
[0022] Based on the L1 norm loss between the real haze-free image and the dehazed image of the real haze-free image, an identity consistency loss is constructed;
[0023] The weighted sum of multi-layer contrast loss, style transfer loss, identity consistency loss, generator loss, and discriminator loss is used as the total loss function of the second stage;
[0024] The generative adversarial network is trained through the total loss function of the second stage, and the trained generative adversarial network is verified through the verification set to obtain the target generative adversarial network;
[0025] The image to be dehazed is input into the target generative adversarial network, and the dehazed image of the image to be dehazed is output.
[0026] Preferably, the step of inputting the defogging image of the real foggy image into a rear feature extractor to obtain the defogging image features of the real foggy image comprises:
[0027] After the dehazed image of the real foggy image passes through a 3×3 convolutional layer, it is downsampled twice by a factor of 2 to obtain a low-resolution feature map of the real foggy image.
[0028] After the low-resolution feature map of the real foggy image passes through three residual modules connected in sequence, the defogging image features of the real foggy image are obtained.
[0029] Preferably, the formula of the total loss function in the second stage is:
[0030] L 2 =λ 1 L C+ +λ 2 L S +λ 3 L I +λ 4 (L G +L D ),
[0031] Among them, L 2 is the total loss function of the second stage, λ 1 is the weight of multi-layer contrast loss, L C is the multi-layer contrast loss, λ 2 is the weight of style transfer loss, L S is the style transfer loss, λ 3 is the weight of the identity consistency loss, L I is the identity consistency loss, λ 4 is the weight of the generator loss and the discriminator loss, L G is the generator loss, L D is the discriminator loss;
[0032] Multi-layer contrast loss L C The formula is:
[0033]
[0034] Among them, E x~X It means that the mathematical expectation is obtained when the real fog image x follows the distribution X, X is the distribution of the real fog image in the sample space, L is the set number of layers, l is the selected layer index, s is the spatial position index, l′(.) is the contrast loss function, is the feature vector of the sub-block with spatial position s in the lth defogging image feature layer of the real fog image, is the feature vector of the sub-block with spatial position s in the lth feature layer of the real fog image, is the feature vector set of the sub-block with spatial position s in the lth feature layer of the real fog image;
[0035] Style transfer loss L S The formula is:
[0036]
[0037] in, The second dehazed image feature of the real haze image generated by the generator after updating the parameters in the second stage, The first dehazed image feature of the real haze image generated by the generator trained in the first stage, ∥.∥ 1 is the L1 norm;
[0038] Generator loss L G The formula is:
[0039]
[0040] Among them, E G(x)~r It means to find the mathematical expectation when the dehazed image G(x) of the real foggy image obeys the distribution r, where r is the distribution of the dehazed image generated by the generator in the sample space, D(.) is the discriminator of the generative adversarial network, and G(x) is the dehazed image of the real foggy image;
[0041] Discriminator loss L D The formula is:
[0042] L D =E y~I [(D(y)-1) 2 ]+E G(x)~r [.D(G(x)) / 2 ],
[0043] Among them, E y~I It means to find the mathematical expectation when the real haze-free image y follows the distribution I, where I is the distribution of the real haze-free image in the sample space;
[0044] The formula for the identity loss L I The formula is:
[0045] L I =E y~I ∥G(y)-y∥ 1 ,
[0046] Among them, E y~I It means to find the mathematical expectation when the real fog image x follows the distribution I, G(x) is the defogging image of the real fog image, x is the real fog image, ∥.∥ 1 is the L1 norm.
[0047] Preferably, the generator of the generative adversarial network includes: a front feature extractor, a defogging network; wherein the defogging network includes: a plurality of defogging modules connected in sequence, each defogging module includes: an image enhancement unit and a defogging unit connected in sequence, and the defogging feature map output by the last defogging module is used as the defogging image;
[0048] The input features of the defogging module are input into the defogging module, and the defogging feature map is output, including:
[0049] The input features of the defogging module are passed through the image enhancement unit to perform feature enhancement, and an output enhanced feature map of the image feature enhancement unit is obtained;
[0050] The output enhanced feature map of the image feature enhancement unit is input into the defogging unit, and the defogging feature map is output, including:
[0051] The output enhanced feature map of the image feature enhancement unit is input into the atmospheric scattering matrix estimation branch, and the atmospheric light matrix is output;
[0052] The output enhanced feature map of the image feature enhancement unit is input into the transmission matrix estimation branch, and the transmission map matrix is output;
[0053] The defogging feature map is calculated based on the atmospheric light matrix, the transmission map matrix, the output enhancement feature map of the image feature enhancement unit, and the input features of the defogging module.
[0054] Preferably, the step of performing feature enhancement on the input features of the defogging module through an image enhancement unit to obtain an output enhanced feature map of the image feature enhancement unit comprises:
[0055] The input features of the defogging module are reduced in dimension to obtain the input features of the defogging module after the dimension reduction;
[0056] The input features of the dehazing module after dimensionality reduction are passed through three linear layers to obtain the key vector, query vector, and value vector;
[0057] Calculate attention weights based on key vector, query vector, and value vector;
[0058] The formula for calculating the attention weight is:
[0059]
[0060] Among them, O i is the output at the i-th pixel position, Q is the query vector, K is the key vector, V is the value vector, and e is a natural constant. T For transposition, is the square root of the channel dimension of the key vector, i is the index of the query vector, j is the index of the value vector, k is the index of the key vector, N is H×W, H is the height of the input feature of the dehazing module, and W is the width of the input feature of the dehazing module;
[0061] After the attention weights are dimensionally converted, they are upsampled to obtain the attention weights after dimension conversion and upsampling;
[0062] After the dimension conversion and upsampling, the attention weight is multiplied by the learnable parameter and then added to the input feature of the dehazing module to obtain the output enhanced feature map of the image feature enhancement unit.
[0063] The formula for obtaining the output enhanced feature map of the image feature enhancement unit by multiplying the attention weight after dimension conversion and upsampling with the learnable parameter and adding it to the input feature of the defogging module is:
[0064]
[0065] in, is the output enhanced feature map of the image feature enhancement unit, γ is a learnable parameter, O is the attention weight after dimension conversion and upsampling, I 0 It is the input feature of the dehazing module.
[0066] Preferably, the output enhancement feature map of the image feature enhancement unit is input into an atmospheric scattering matrix estimation branch, and the atmospheric light matrix is output, including:
[0067] The output enhanced feature map of the image feature enhancement unit is subjected to mixed average pooling to obtain a pooled enhanced feature map;
[0068] The pooled enhanced feature map is processed by a 3×3 convolution layer and a ReLU activation function in sequence, and then concatenated with the pooled enhanced feature map to obtain the initial fusion feature.
[0069] The initial fusion features are processed by 3×3 convolutional layers and ReLU activation functions in sequence, and then concatenated with the initial fusion features to obtain secondary fusion features.
[0070] The secondary fusion features are processed by 1×1 convolution layer and Sigmoid activation function in sequence, and then added element by element with the initial fusion features to obtain the atmospheric light pre-generation features;
[0071] After upsampling the pre-generated features of atmospheric light, an atmospheric light matrix is obtained.
[0072] Preferably, the step of inputting the output enhanced feature map of the image feature enhancement unit into a transmission matrix estimation branch and outputting a transmission map matrix comprises:
[0073] The output enhanced feature map of the image feature enhancement unit passes through a 3×3 convolution layer to obtain a feature map processed by a 3×3 convolution layer;
[0074] The feature map processed by the 3×3 convolution layer is processed by the 3×3 convolution layer and the ReLU activation function in sequence, and then concatenated with the feature map processed by the 3×3 convolution layer to obtain the first spliced feature;
[0075] The first stitched feature is processed by a 3×3 convolution layer and a ReLU activation function in sequence, and then concatenated with the first stitched feature to obtain a second stitched feature;
[0076] The second stitched feature is processed by a 3×3 convolution layer and a ReLU activation function in sequence, and then concatenated with the second stitched feature to obtain a third stitched feature;
[0077] The third stitching feature is processed by a 1×1 convolution layer and a Sigmoid activation function in sequence, and then added element by element to the feature map processed by a 3×3 convolution layer to obtain a transmission map matrix.
[0078] Preferably, the defogging feature map is calculated based on the atmospheric light matrix, the transmission map matrix, the output enhancement feature map of the image feature enhancement unit, and the input features of the defogging module, and the calculation formula is:
[0079]
[0080] in, is the dehazing feature map, is the output enhanced feature map of the image feature enhancement unit, ⊙ is the element-by-element multiplication, is the transmission map matrix, is the atmospheric light matrix, It is the input feature of the dehazing module.
[0081] Preferably, the defogging image of the synthetic fog paired image, the atmospheric light matrix output by the atmospheric scattering matrix estimation branch in the defogging unit of the last defogging module, and the transmission map matrix output by the transmission matrix estimation branch in the defogging unit of the last defogging module are passed through an image reconstruction module to obtain a reconstructed fog image. The calculation formula for the reconstructed fog image is:
[0082]
[0083] Among them, I rec To reconstruct the fog image, is the defogging image of the synthetic fog paired image, ⊙ is the element-by-element multiplication, is the transmission map matrix output by the transmission matrix estimation branch in the defogging unit of the last defogging module, It is the atmospheric light matrix output by the atmospheric scattering matrix estimation branch in the defogging unit of the last defogging module. for The inverse matrix of It is the dehazing feature map output by the last dehazing module.
[0084] Preferably, the image to be defogged is input into a target generative adversarial network, and a defogged image of the image to be defogged is output, including:
[0085] The image to be dehazed is input into the front feature extractor to extract the feature map of the image to be dehazed, including:
[0086] After the image to be dehazed passes through a 3×3 convolutional layer, it is downsampled twice by a factor of 2 to obtain a low-resolution feature map of the image to be dehazed;
[0087] After the low-resolution feature map of the image to be dehazed passes through three residual modules connected in sequence, it is upsampled twice by a factor of 2 to obtain the feature map of the image to be dehazed;
[0088] The feature map of the image to be dehazed is input into the dehazing network, and the dehazed image of the image to be dehazed is output.
[0089] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0090] The present invention discloses an image defogging method based on a generative adversarial network. The present invention effectively solves the problem of the difference between synthetic data and real scenes by designing two training stages. In the first stage, the defogging network is trained by synthetic data. The generation process of synthetic data is relatively simple. Compared with real foggy images, the complexity of the data is lower. In the first stage, the defogging network is trained with synthetic data, which not only solves the problem of data scarcity, but also enhances the target effectiveness, interpretability and preliminary generalization ability of the model. In the second stage, the generative adversarial network is trained with real data. This method allows the model to deeply learn the characteristics of fog in real scenes, effectively bridging the gap between synthetic data and real scenes, and improving the adaptability and accuracy of the model in real environments. It can ensure that effective defogging can be achieved even when the real fog-free image data is limited.
[0091] The present invention constructs a total loss function of the second stage including multi-layer contrast loss, style transfer loss, generator loss, discriminator loss and identity consistency loss in the second stage loss function of the present invention, thereby avoiding the problem of color and texture distortion. Multi-layer contrast loss is introduced, and the multi-layer features of the real fog image and the defogged image features of the real fog image are obtained by passing the feature map of the real fog image and the defogged image features of the real fog image through two layers of perceptron networks respectively. The contrast loss function of each feature layer of the real fog image and each spatial position in the defogged image feature layer of the real fog image is calculated to construct a multi-layer contrast loss. This enables the model to learn the subtle differences between the images before and after defogging in different feature layers and spatial positions, thereby more accurately adjusting the image during the defogging process, avoiding color and texture distortion, and improving the naturalness and accuracy of the defogged image; the style transfer loss is introduced, and the defogging loss function of the real fog image is used to obtain the multi-layer contrast loss. By comparing the dehazed image features of the real foggy images obtained by the dehazing network after training in the first and second stages, the model can better utilize the knowledge learned in the first stage when processing real foggy images, further optimize the dehazing effect, and reduce the color and texture distortion problems that may be caused by changes in the training stage; the identity consistency loss is introduced. The identity consistency loss can improve the generalization ability of the dehazing generator by calculating the L1 norm between the real fog-free image and the dehazed image of the real fog-free image, thereby further making the generated dehazed image maintain the color, structure and texture characteristics of the real foggy image, further improving the accuracy of the dehazed image.
[0092] The present invention also designs a defogging network, which includes: a plurality of defogging modules connected in sequence, each of which includes: an image enhancement unit and a defogging unit connected in sequence, and through a self-attention mechanism, the image enhancement unit can capture the relationship between features at different positions in the image, thereby improving the image defogging capability, and can better retain the original details and structure of the image during the defogging process, reduce image distortion caused by defogging, and improve the quality of the defogging image. The defogging unit outputs an atmospheric light matrix and a transmission map matrix respectively through an atmospheric scattering matrix estimation branch and a transmission matrix estimation branch, and combines them with the enhanced feature map to calculate the defogging feature map, thereby realizing a deep fusion of the physical model and deep learning. This combination makes the defogging process more consistent with the real physical principles, thereby improving the accuracy and reliability of defogging, making the generated defogging image closer to the real fog-free state, and improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:
[0094] Figure 1 It is a flowchart of the steps of an image defogging method based on a generative adversarial network of the present invention.
[0095] Figure 2 It is a schematic diagram of multi-layer contrast loss.
[0096] Figure 3 It is a structural diagram of the generative adversarial network.
[0097] Figure 4 It is a structural diagram of the defogging unit.
[0098] Figure 5 This is a comparison of the effect of defogging by generating an adversarial network. Figure 5 (a) is a foggy picture. Figure 5 (b) in the figure is the dehazed image. Figure 5 (c) in the figure is the label map. DETAILED DESCRIPTION
[0099] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.
[0100] like Figure 1 As shown, Figure 1 This is a flowchart of the steps of an image dehazing method based on a generative adversarial network of the present invention.
[0101] Embodiment 1 of the present invention provides an image defogging method based on a generative adversarial network, comprising the following steps:
[0102] Acquire an image data set, the image data set comprising: different synthetic fog data and real fog data, the synthetic fog data comprising a fog-free image and its corresponding synthetic fog paired image, the real fog data comprising a real fog image and its corresponding real fog-free image, preprocess the image data set, and divide the preprocessed image data set into a training set and a verification set;
[0103] Input the synthetic fog paired image into the generator of the generative adversarial network to obtain the defogged image of the synthetic fog paired image;
[0104] By reconstructing the defogging image of the synthetic fog paired image, a reconstructed fog image of the synthetic fog paired image is obtained;
[0105] The sum of the smoothed L1 loss between the dehazed image and the haze-free image of the synthetic haze paired image and the smoothed L1 loss between the reconstructed haze image and the synthetic haze paired image is used as the total loss function of the first stage;
[0106] In this embodiment, specifically, the formula of the total loss function in the first stage is:
[0107]
[0108] Among them, L 1 is the total loss function of the first stage, is the smooth L1 loss between the dehazed image and the haze-free image of the synthetic haze paired image, The smooth L1 loss between the reconstructed fog image and the synthetic fog paired image;
[0109] Taking the smooth L1 loss between the dehazed image and the haze-free image of the synthetic haze paired image as an example, the calculation formula of the smooth L1 loss is:
[0110]
[0111] Among them, n f is the total number of pixels of the defogging image or the fog-free image of the synthetic fog paired image, σ is the pixel index, is the haze-free image y f The value of the σ-th pixel in , is the value of the σth pixel in the defogging image of the synthetic fog paired image, Pair images for synthetic fog.
[0112] Compared with the L2 loss, the smooth L1 loss adopted in this embodiment can prevent potential gradient explosion.
[0113] The generator is trained through the total loss function of the first stage to obtain the generator after the first stage training;
[0114] Input the real fog image into the generator to obtain the feature map of the real fog image and the defogging image of the real fog image;
[0115] Input the real fog-free image into the generator to obtain the dehazed image of the real fog-free image;
[0116] The defogging image of the real foggy image is input into the post-feature extractor to obtain the defogging image features of the real foggy image;
[0117] The feature map of the real fog image is passed through a two-layer perceptron network M to obtain the L-layer feature map of the real fog image {f l} L ;
[0118] The defogging image features of the real foggy image are passed through a two-layer perceptron network M to obtain the L-layer defogging image features of the real foggy image.
[0119] By calculating the contrast loss function of each feature layer of the real fog image and each spatial position in the feature layer of the defogging image of the same layer, a multi-layer contrast loss L is constructed. C ;
[0120] like Figure 2 As shown, Figure 2It is a schematic diagram of multi-layer contrast loss.
[0121] The spatial position set of each feature layer of the real fog image and the feature layer of the defogging image at the same layer is N l represents the total number of spatial positions in each layer, is the last spatial position of each layer;
[0122] In this embodiment, specifically, the multi-layer contrast loss L C The formula is:
[0123]
[0124] Among them, E x~X It means that the mathematical expectation is obtained when the real fog image x follows the distribution X, X is the distribution of the real fog image in the sample space, L is the set number of layers, l is the selected layer index, s is the spatial position index, l′(.) is the contrast loss function, is the feature vector of the sub-block with spatial position s in the lth defogging image feature layer of the real fog image, is the feature vector of the sub-block with spatial position s in the lth feature layer of the real fog image, is the feature vector set of the sub-block with spatial position s in the lth feature layer of the real fog image;
[0125] In this embodiment, specifically, the process of obtaining a single contrast loss function is:
[0126] A sub-block is randomly selected from the dehazed image y of the real foggy image as an anchor point, the sub-block corresponding to the anchor point in the real foggy image x is represented as a positive sample, and the other sub-blocks in the real foggy image x are represented as negative samples. The correlation between the anchor point and the positive sample is maximized through the noise contrast estimation module.
[0127] The anchor point, positive sample and V negative samples are converted into feature vectors and represented as c, c respectively. + 、c - , then the single contrast loss function can be expressed as the cross entropy loss, the formula is:
[0128]
[0129] Among them, l′(c,c + ,c - ) is a single contrast loss function, exp(.) represents an exponential function with e as the base, and d(c,c + ) represents the feature space distance between the anchor point and the positive sample, that is, the cosine similarity between the anchor point and the positive sample, τ is the adjustment factor, is the feature space distance between the anchor point and the vth negative sample.
[0130] The introduction of multi-layer contrast loss enables the model to learn the subtle differences between the images before and after dehazing at different feature layers and spatial positions, thereby adjusting the image more accurately during the dehazing process, avoiding color and texture distortion, and improving the naturalness and accuracy of the dehazed image.
[0131] Based on the first dehazed image features of the real foggy image generated by the generator trained in the first stage and the second dehazed image features of the real foggy image generated by the generator after updating the parameters in the second stage, the style transfer loss L is constructed. S ;
[0132] In this embodiment, specifically, the style transfer loss L S The formula is:
[0133]
[0134] in, The second dehazed image feature of the real haze image generated by the generator after updating the parameters in the second stage, The first dehazed image feature of the real haze image generated by the generator trained in the first stage, ∥.∥ 1 is the L1 norm;
[0135] The introduction of style transfer loss enables the model to better utilize the knowledge learned in the first stage when processing real foggy images, further optimize the dehazing effect, and reduce color and texture distortion problems that may be caused by changes in the training stage.
[0136] Based on the L1 norm loss between the dehazed image and the real haze-free image, we construct the identity consistency loss L I ;
[0137] In this embodiment, specifically, the identity consistency loss formula L I The formula is:
[0138] L I =E y~I ∥G(y)-y∥ 1 ,
[0139] Among them, E y~I It means to find the mathematical expectation when the real haze-free image y follows the distribution I, G(y) is the dehazed image of the real haze-free image, y is the real haze-free image, ∥.∥ 1 is the L1 norm.
[0140] The weighted sum of multi-layer contrast loss, style transfer loss, identity consistency loss, generator loss, and discriminator loss is used as the total loss function of the second stage;
[0141] In this embodiment, preferably, the formula of the total loss function in the second stage is:
[0142] L 2 =λ 1 L C+ +λ 2 L S +λ 3 L I +λ 4 (L G +L D ),
[0143] Among them, L 2 is the total loss function of the second stage, λ 1 is the weight of multi-layer contrast loss, L C is the multi-layer contrast loss, λ 2 is the weight of style transfer loss, L S is the style transfer loss, λ 3 is the weight of the identity consistency loss, L I is the identity consistency loss, λ 4 is the weight of the generator loss and the discriminator loss, L G is the generator loss, L D is the discriminator loss;
[0144] In this embodiment, specifically, the generator loss L G The formula is:
[0145] L G =E G(x)~r 0(D(G(x))-1) 2 1,
[0146] Among them, E G(x)~r It means to find the mathematical expectation when the dehazed image G(x) of the real foggy image obeys the distribution r, where r is the distribution of the dehazed image generated by the generator in the sample space, D(.) is the discriminator of the generative adversarial network, and G(x) is the dehazed image of the real foggy image;
[0147] Discriminator loss L D The formula is:
[0148] L D =E y~I [(D(y)-1) 2 ]+E G(x)~r [.D(G(x)) / 2 ],
[0149] Among them, E y~I It means to find the mathematical expectation when the real haze-free image y follows the distribution I, where I is the distribution of the real haze-free image in the sample space;
[0150] The generative adversarial network is trained through the total loss function of the second stage, and the trained generative adversarial network is verified through the verification set to obtain the target generative adversarial network;
[0151] The image to be dehazed is input into the target generative adversarial network, and the dehazed image of the image to be dehazed is output.
[0152] In this embodiment, preferably, Figure 3 As shown, Figure 3 This is the structural diagram of the generated adversarial network.
[0153] Constructing a generative adversarial network, wherein the generator of the generative adversarial network includes: a front feature extractor, a defogging network; wherein the defogging network includes: a plurality of defogging modules connected in sequence, each defogging module includes: an image enhancement unit and a defogging unit connected in sequence, and the defogging feature map output by the last defogging module is used as the defogging image;
[0154] In this embodiment, preferably, the input features of the defogging module are input into the defogging module, and the defogging feature map is output, including:
[0155] The input features of the defogging module are enhanced through the image enhancement unit to obtain the output enhanced feature map of the image feature enhancement unit, including:
[0156] The input features of the dehazing module Perform dimensionality reduction to obtain the input features of the defogging module after dimensionality reduction Among them, B is the batch size, C is the number of channels of the input feature of the defogging module, H is the height of the input feature of the defogging module, and W is the width of the input feature of the defogging module;
[0157] The input features of the defogging module after dimensionality reduction are transformed into The key linear layer is obtained to obtain the key vector K;
[0158] The input features of the defogging module after dimensionality reduction are transformed into The query linear layer is used to obtain the query vector Q;
[0159] The input features of the defogging module after dimensionality reduction are transformed into The value linear layer obtains the value vector V;
[0160] Calculate attention weights based on key vector, query vector, and value vector;
[0161] The formula for calculating the attention weight is:
[0162]
[0163] Among them, Oi is the output at the i-th pixel position, Q is the query vector, K is the key vector, V is the value vector, and e is a natural constant. T For transposition, is the square root of the channel dimension of the key vector, i is the index of the query vector, j is the index of the value vector, k is the index of the key vector, N is H×W, H is the height of the input feature of the dehazing module, and W is the width of the input feature of the dehazing module;
[0164] After the attention weights are dimensionally converted, they are upsampled to obtain the attention weights after dimension conversion and upsampling;
[0165] After the dimension conversion and upsampling, the attention weight is multiplied by the learnable parameter and then added to the input feature of the dehazing module to obtain the output enhanced feature map of the image feature enhancement unit.
[0166] The formula for obtaining the output enhanced feature map of the image feature enhancement unit by multiplying the attention weight after dimension conversion and upsampling with the learnable parameter and adding it to the input feature of the defogging module is:
[0167]
[0168] in, is the output enhanced feature map of the image feature enhancement unit, γ is a learnable parameter, O is the attention weight after dimension conversion and upsampling, I 0 It is the input feature of the dehazing module.
[0169] The feature map extracted by the front feature extractor is sent to the dehazing network based on the atmospheric scattering model. The atmospheric scattering model is usually simplified into the following formula:
[0170] I(χ)=J(χ)T(χ)+A(1-T(χ)),
[0171] Where χ is the position of the pixel, T(χ) is the medium transmission map, A is the global atmospheric light, I(χ) is the pixel value of the foggy image at pixel χ, and J(χ) is the pixel value of the fog-free image at pixel χ.
[0172] For each pixel, the above formula is transformed to obtain the reconstruction model J(x) of the clean image:
[0173]
[0174] Let κ be the convolution kernel used for feature extraction, then the above formula is:
[0175]
[0176] Among them, Θ represents the convolution operation, ⊙ represents the Hadamard multiplication;
[0177] Then it can be expressed as a matrix-vector:
[0178]
[0179] in, is the matrix representation of κ, is the matrix representation of J, is the matrix representation of I, for The matrix representation of is the matrix representation of A;
[0180] Let F 1 As the former feature extractor is equivalent to The same assumption F 2 As The characteristic representation of , the formula is further rewritten as:
[0181]
[0182] set up for The estimate, F 2 The estimate, for The characteristic representation of for The characteristic representation of , then the formula is rewritten as:
[0183]
[0184] Since the main features before and after defogging remain unchanged, the output defogging feature map of the defogging unit is It is expressed as:
[0185]
[0186] According to the above reasoning process, a defogging unit based on the physical model can be constructed. The key to the defogging effect is to accurately estimate the global atmospheric light and medium transmission map, which in turn come from the output enhanced feature map of the image feature enhancement unit. Therefore, one branch of the defogging unit is to estimate the atmospheric light matrix from the output enhanced feature map of the image feature enhancement unit. The other branch is to estimate the transmission map matrix from the output enhanced feature map of the image feature enhancement unit
[0187] like Figure 4 As shown, Figure 4 This is the structural diagram of the defogging unit.
[0188] The output enhanced feature map of the image feature enhancement unit is input into the defogging unit, and the defogging feature map is output, including:
[0189] The output enhanced feature map of the image feature enhancement unit is input into the atmospheric scattering matrix estimation branch, and the atmospheric light matrix is output;
[0190] In this embodiment, preferably, the output enhancement feature map of the image feature enhancement unit is input into the atmospheric scattering matrix estimation branch, and the atmospheric light matrix is output, including:
[0191] The output enhanced feature map of the image feature enhancement unit is subjected to mixed average pooling to obtain a pooled enhanced feature map;
[0192] The pooled enhanced feature map is processed by a 3×3 convolution layer and a ReLU activation function in sequence, and then concatenated with the pooled enhanced feature map to obtain the initial fusion feature.
[0193] The initial fusion features are processed by 3×3 convolutional layers and ReLU activation functions in sequence, and then concatenated with the initial fusion features to obtain secondary fusion features.
[0194] The secondary fusion features are processed by 1×1 convolution layer and Sigmoid activation function in sequence, and then added element by element with the initial fusion features to obtain the atmospheric light pre-generation features;
[0195] After upsampling the pre-generated features of atmospheric light, an atmospheric light matrix is obtained.
[0196] Normally, the distribution of atmospheric light values in the image feature space is relatively uniform, so some unimportant information can be ignored through pooling. However, in actual practice, due to the density and distribution of fog, the direction and intensity of light sources, the atmospheric light values are not completely uniform in some areas. Therefore, the present invention adopts hybrid pooling, which captures some average information while retaining significant features to improve the comprehensiveness of the features. Then, through a dense residual connection, the atmospheric light matrix is finally output. The above series of processes can be expressed by the following formula:
[0197]
[0198] in, is the enhanced feature map after pooling, MP(.) is mixed average pooling, concat(.) is the concatenation operation on the feature dimension, ReLU(.) is the ReLU activation function, is the atmospheric light matrix, up(.) is the upsampling operation, Sigmoid(.) is the Sigmoid activation function, conv 1×1 (.) is a 1×1 convolutional layer, conv 3×3(.) is a 3×3 convolutional layer, is the initial fusion feature, It is a secondary fusion feature.
[0199] The output enhanced feature map of the image feature enhancement unit is input into the transmission matrix estimation branch, and the transmission map matrix is output;
[0200] In this embodiment, preferably, the step of inputting the output enhanced feature map of the image feature enhancement unit into a transmission matrix estimation branch and outputting a transmission map matrix comprises:
[0201] The output enhanced feature map of the image feature enhancement unit passes through a 3×3 convolution layer to obtain a feature map processed by a 3×3 convolution layer;
[0202] The feature map processed by the 3×3 convolution layer is processed by the 3×3 convolution layer and the ReLU activation function in sequence, and then concatenated with the feature map processed by the 3×3 convolution layer to obtain the first spliced feature;
[0203] The first stitched feature is processed by a 3×3 convolution layer and a ReLU activation function in sequence, and then concatenated with the first stitched feature to obtain a second stitched feature;
[0204] The second stitched feature is processed by a 3×3 convolution layer and a ReLU activation function in sequence, and then concatenated with the second stitched feature to obtain a third stitched feature;
[0205] The third stitching feature is processed by a 1×1 convolution layer and a Sigmoid activation function in sequence, and then added element by element to the feature map processed by a 3×3 convolution layer to obtain a transmission map matrix.
[0206] Since the transmission map matrix is closely related to the transmittance, and the transmittance is mainly affected by the depth of field of the entire scene, we need to make full use of the overall and local information of the image and use a 3×3 convolution layer to replace the mixed pooling of the atmospheric scattering matrix estimation branch.
[0207] The acquisition process of the transmission map matrix is expressed by the following formula:
[0208]
[0209] in, is the feature map after processing by 3×3 convolution layer, is the first stitching feature, is the second stitching feature, is the third stitching feature, is the transmission map matrix.
[0210] The defogging feature map is calculated based on the atmospheric light matrix, the transmission map matrix, the output enhancement feature map of the image feature enhancement unit, and the input features of the defogging module.
[0211] In this embodiment, preferably, the defogging feature map is calculated based on the atmospheric light matrix, the transmission map matrix, the output enhancement feature map of the image feature enhancement unit, and the input features of the defogging module, and the calculation formula is:
[0212]
[0213] in, is the dehazing feature map, is the output enhanced feature map of the image feature enhancement unit, ⊙ is the element-by-element multiplication, is the transmission map matrix, is the atmospheric light matrix, It is the input feature of the dehazing module.
[0214] Since the atmospheric scattering model is equivalent to A, and It is essentially obtained by inverse transformation of T, thus obtaining the calculation formula for reconstructing the fog image.
[0215] In this embodiment, preferably, the defogging image of the synthetic fog paired image, the atmospheric light matrix output by the atmospheric scattering matrix estimation branch in the defogging unit of the last defogging module, and the transmission map matrix output by the transmission matrix estimation branch in the defogging unit of the last defogging module are passed through an image reconstruction module to obtain a reconstructed fog image. The calculation formula for the reconstructed fog image is:
[0216]
[0217] Among them, I rec To reconstruct the fog image, is the defogging image of the synthetic fog paired image, ⊙ is the element-by-element multiplication, is the transmission map matrix output by the transmission matrix estimation branch in the defogging unit of the last defogging module, It is the atmospheric light matrix output by the atmospheric scattering matrix estimation branch in the defogging unit of the last defogging module. for The inverse matrix of It is the dehazing feature map output by the last dehazing module.
[0218] In this embodiment, preferably, the image to be defogged is input into a target generative adversarial network, and a defogged image of the image to be defogged is output, including:
[0219] The image to be dehazed is input into the front feature extractor to extract the feature map of the image to be dehazed, including:
[0220] After the image to be dehazed passes through a 3×3 convolutional layer, it is downsampled twice by a factor of 2 to obtain a low-resolution feature map of the image to be dehazed;
[0221] After the low-resolution feature map of the image to be dehazed passes through three residual modules connected in sequence, it is upsampled twice by a factor of 2 to obtain the feature map of the image to be dehazed;
[0222] The feature map of the image to be dehazed is input into the dehazing network, and the dehazed image of the image to be dehazed is output.
[0223] The front feature extractor captures the local features of the image to be defogged, such as edges and textures, through a 3×3 convolutional layer, and then downsamples twice by a factor of 2, which not only reduces the data dimension and the amount of subsequent calculations, but also highlights the main features and avoids overfitting. The low-resolution feature map of the image to be defogged passes through three residual modules connected in sequence, and uses jump connections to solve the gradient problem, achieve deep feature learning, and improve the defogging network's ability to handle different fog interferences. Two times of upsampling by a factor of 2 restore the low-resolution feature map of the image to be defogged to the original resolution, while retaining key features and restoring the image structure, providing high-quality input for the defogging network. These operations enhance the adaptability and generalization ability of the defogging network to different scenes and fog levels, and also simplify the learning task of the defogging network, making it converge faster, optimizing the training process, improving training efficiency, and shortening training time.
[0224] The present invention effectively solves the problem of the difference between synthetic data and real scenes by designing two training stages. In the first stage, the defogging network is trained by synthetic data. The generation process of synthetic data is relatively simple. Compared with real foggy images, the complexity of the data is lower. Using synthetic data to train the defogging network in the first stage not only solves the problem of data scarcity, but also enhances the target effectiveness, interpretability and preliminary generalization ability of the model. In the second stage, the generative adversarial network is trained by real data. This method allows the model to deeply learn the characteristics of fog in real scenes, effectively bridging the gap between synthetic data and real scenes, and improving the adaptability and accuracy of the model in real environments. It can ensure that effective defogging can be achieved even when the real fog-free image data is limited.
[0225] Specifically, in the first stage of pre-training, the Adam optimizer is empirically used and its momentum parameter is set to β 1 =0.9,β 2 = 0.999 and train for 100 cycles. We set the initial learning rate to 1×10 -4, and the decay rate is 0.8 every 20 cycles. In the fine-tuning stage, due to the possible instability of training, this embodiment sets the same 100 cycles as the first stage. The initial learning rate of the generator and discriminator is set to 2×10 -4 , and the linear decay strategy is used every 30 epochs. The trade-off weight in the generator loss function is set to λ according to the trial and error method. 1 =1,λ 2 =0.5,λ 3 =0.02,λ 4 =1.
[0226] As shown in Table 1, Table 1 is a schematic diagram of various evaluation indicators of the present invention on the SOTS dataset, URHI dataset and RTTS dataset.
[0227] Table 1
[0228]
[0229] In the experiment on the synthetic dataset SOTS, the performance of the image dehazing method based on generative adversarial network proposed in this paper is comprehensively evaluated from both qualitative and quantitative dimensions.
[0230] In the qualitative evaluation stage, the present invention can effectively eliminate the residual haze in the image, no artifacts are generated in the sky area, the overall brightness of the image is balanced, and the colors appear real and natural. At the same time, the present invention can significantly restore the texture details obscured by haze, greatly improving the visual clarity and realism of the image.
[0231] In terms of quantitative evaluation, the signal-to-noise ratio (PSNR) and structural similarity (SSIM), two standard indicators widely used in the field of synthetic data dehazing research, are introduced. Experimental results show that on the SOTS dataset, the present invention achieves an average peak signal-to-noise ratio (PSNR) of 34.13 dB and an average structural similarity (SSIM) of 0.9863 on the SOTS dataset, both of which reach the current optimal level. In addition, the standard deviation of the model on these indicators is low, which indicates that the performance of the model on different samples is highly stable, which fully proves the robustness of the model.
[0232] Experiments on real haze datasets URHI and RTTS also verified the effectiveness of the present invention in actual scenarios. During the experiment, 300 images were randomly selected from URHI and RTTS for testing. The results show that the present invention can remove most of the haze residues in real haze images and significantly improve the texture detail performance of the image. Especially in the sky area, it is almost impossible to detect the presence of artifacts. The overall color of the image is natural, the brightness distribution is balanced, and the visual effect is significantly better than other methods. In the RTTS dataset, the dehazing effect of the image is also clear and excellent. Even if there is high-density haze in some areas, the present invention can still be effectively processed while retaining the original detail information of the image to the greatest extent.
[0233] In the quantitative analysis stage, the frequency domain-based analysis method is used as the main evaluation index to measure the dehazing quality of the image. The experimental results show that on the URHI dataset, the score of the present invention reaches 27.8128; on the RTTS dataset, the score reaches 27.7937, both at a high level. These results fully demonstrate that the present invention can effectively remove haze in actual complex environments while taking into account the details and realism of the image.
[0234] In addition, two key indicators, edge strength and frequency domain clarity, are introduced to further verify the dehazing effect. The edge strength is calculated by the edge detection algorithm of the Canny operator, and the frequency domain clarity is obtained based on the clarity evaluation function of the frequency domain analysis. These two indicators can effectively reflect the degree of edge detail retention and overall clarity of the image after dehazing.
[0235] In order to fully demonstrate the effect of the image dehazing method based on generative adversarial network proposed in the present invention in practical applications, this embodiment 2 verifies its great potential in improving the performance of downstream tasks by taking dehazing as a key preprocessing step of the target detection task.
[0236] Figure 5 This is a comparison of the effect of defogging through the generation of adversarial networks. Figure 5 (a) is a foggy picture. Figure 5 (b) in the figure is the dehazed image. Figure 5 (c) in the figure is the label map.
[0237] In the second embodiment, 300 challenging haze images are randomly selected from the RTTS dataset, and the target detection model YOLOv8 is used to evaluate the detection performance of the original haze images and the images processed by different dehazing algorithms.
[0238] By calculating the average precision (AP) of each processing method, this Example 2 reveals the actual impact of different defogging technologies on target detection tasks, and the experimental data fully demonstrates the excellent performance of this method. Compared with unprocessed haze images, the defogging method based on the generative adversarial network proposed by the present invention improves the average accuracy of target detection by 4.6%, which is significantly ahead of other defogging algorithms. In addition, the present invention can effectively solve the problem of detection blind spots caused by haze occlusion, so that many targets that were originally undetectable can be successfully identified after defogging. At the same time, it can also correct the misclassification and positioning caused by haze interference, providing higher accuracy and reliability for target detection tasks.
[0239] What is more noteworthy is that the image dehazing method based on generative adversarial network proposed in the present invention not only improves the detection accuracy, but also greatly optimizes the recall rate and overall performance of the model.
[0240] Experiments show that the detection accuracy and recall rate are improved by 4.5% and 3.3% respectively, indicating that the present invention can significantly enhance the visibility and information integrity of the image and provide clearer and more detailed input data for the target detection task.
[0241] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0242] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0243] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0244] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0245] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.
Claims
1. An image dehazing method based on a generative adversarial network, characterized in that: The following steps are involved: Acquire an image data set, the image data set comprising: different synthetic fog data and real fog data, the synthetic fog data comprising a fog-free image and its corresponding synthetic fog paired image, the real fog data comprising a real fog image and its corresponding real fog-free image, preprocess the image data set, and divide the preprocessed image data set into a training set and a verification set; Input the synthetic fog paired image into the generator of the generative adversarial network to obtain the defogged image of the synthetic fog paired image; By reconstructing the defogging image of the synthetic fog paired image, a reconstructed fog image of the synthetic fog paired image is obtained; The sum of the smoothed L1 loss between the dehazed image and the haze-free image of the synthetic haze paired image and the smoothed L1 loss between the reconstructed haze image and the synthetic haze paired image is used as the total loss function of the first stage; The generator is trained through the total loss function of the first stage to obtain the generator after the first stage training; Input the real fog image into the generator to obtain the feature map of the real fog image and the defogging image of the real fog image; Input the real fog-free image into the generator to obtain the dehazed image of the real fog-free image; The defogging image of the real foggy image is input into the post-feature extractor to obtain the defogging image features of the real foggy image; The feature map of the real fog image is passed through a two-layer perceptron network to obtain the multi-layer features of the real fog image; The defogging image features of the real foggy image are passed through a two-layer perceptron network to obtain the multi-layer defogging image features of the real foggy image; By calculating the contrast loss function of each feature layer of the real fog image and each spatial position in the feature layer of the dehazed image of the same layer, a multi-layer contrast loss is constructed; Based on the first dehazed image features of the real foggy image generated by the generator trained in the first stage and the second dehazed image features of the real foggy image generated by the generator after updating parameters in the second stage, a style transfer loss is constructed; Based on the L1 norm loss between the real haze-free image and the dehazed image of the real haze-free image, an identity consistency loss is constructed; The weighted sum of multi-layer contrast loss, style transfer loss, identity consistency loss, generator loss, and discriminator loss is used as the total loss function of the second stage; The generative adversarial network is trained through the total loss function of the second stage, and the trained generative adversarial network is verified through the verification set to obtain the target generative adversarial network; The image to be dehazed is input into the target generative adversarial network, and the dehazed image of the image to be dehazed is output.
2. The image dehazing method based on a generative adversarial network according to claim 1, characterized in that: The defogging image of the real foggy image is input into the rear feature extractor to obtain the defogging image features of the real foggy image, including: After the dehazed image of the real foggy image passes through a 3×3 convolutional layer, it is downsampled twice by a factor of 2 to obtain a low-resolution feature map of the real foggy image. After the low-resolution feature map of the real foggy image passes through three residual modules connected in sequence, the defogging image features of the real foggy image are obtained.
3. The image dehazing method based on generative adversarial network according to claim 1, characterized in that: The formula for the total loss function in the second stage is: L2=λ1L C+ +λ2L S +λ3L I +λ4(L G +L D ), Among them, L2 is the total loss function of the second stage, λ1 is the weight of multi-layer contrast loss, and L C is the multi-layer contrast loss, λ2 is the weight of the style transfer loss, L S is the style transfer loss, λ3 is the weight of the identity consistency loss, L I is the identity consistency loss, λ4 is the weight of the generator loss and the discriminator loss, L G is the generator loss, L D is the discriminator loss; Multi-layer contrast loss L C The formula is: Among them, E x~X It means that the mathematical expectation is obtained when the real fog image x follows the distribution X, where X is the distribution of the real fog image in the sample space, L is the set number of layers, l is the selected layer index, s is the spatial position index, and l ′ (.) is the contrast loss function, is the feature vector of the sub-block with spatial position s in the lth dehazed image feature layer of the real fog image, f l s is the feature vector of the sub-block with spatial position s in the lth feature layer of the real fog image, {f l o } o≠s is the feature vector set of the sub-block with spatial position s in the lth feature layer of the real fog image; Style transfer loss L S The formula is: in, The second dehazed image feature of the real haze image generated by the generator after updating the parameters in the second stage, It is the first dehazed image feature of the real foggy image generated by the generator trained in the first stage, and ∥.∥1 is the L1 norm; Generator loss L G The formula is: L G =E G(x)~r 0(D(G(x))-1) 2 1, Among them, E G(x)~r It means to find the mathematical expectation when the dehazed image G(x) of the real foggy image obeys the distribution r, where r is the distribution of the dehazed image generated by the generator in the sample space, D(.) is the discriminator of the generative adversarial network, and G(x) is the dehazed image of the real foggy image; Discriminator loss L D The formula is: L D =E y~I [(D(y)-1) 2 ]+E G(x)~r [.D(G(x)) / 2 ], Among them, E y~I It means to find the mathematical expectation when the real haze-free image y follows the distribution I, where I is the distribution of the real haze-free image in the sample space; The formula for the identity loss L I The formula is: L I =E y~I ∥G(y)-y∥1, Among them, E y~I It means finding the mathematical expectation when the real haze-free image y follows the distribution I, G(y) is the dehazed image of the real haze-free image, y is the real haze-free image, and ∥.∥1 is the L1 norm.
4. The image dehazing method based on a generative adversarial network according to claim 1, characterized in that: The generator of the generative adversarial network includes: a front feature extractor and a defogging network; wherein the defogging network includes: a plurality of defogging modules connected in sequence, each of which includes: an image enhancement unit and a defogging unit connected in sequence, and the defogging feature map output by the last defogging module is used as the defogging image; The input features of the defogging module are input into the defogging module, and the defogging feature map is output, including: The input features of the defogging module are passed through the image enhancement unit to perform feature enhancement, and an output enhanced feature map of the image feature enhancement unit is obtained; The output enhanced feature map of the image feature enhancement unit is input into the defogging unit, and the defogging feature map is output, including: The output enhanced feature map of the image feature enhancement unit is input into the atmospheric scattering matrix estimation branch, and the atmospheric light matrix is output; The output enhanced feature map of the image feature enhancement unit is input into the transmission matrix estimation branch, and the transmission map matrix is output; The defogging feature map is calculated based on the atmospheric light matrix, the transmission map matrix, the output enhancement feature map of the image feature enhancement unit, and the input features of the defogging module.
5. The image defogging method based on generative adversarial network according to claim 4, characterized in that: The step of performing feature enhancement on the input features of the defogging module through an image enhancement unit to obtain an output enhanced feature map of the image feature enhancement unit includes: The input features of the defogging module are reduced in dimension to obtain the input features of the defogging module after the dimension reduction; The input features of the dehazing module after dimensionality reduction are passed through three linear layers to obtain the key vector, query vector, and value vector; Calculate attention weights based on key vector, query vector, and value vector; The formula for calculating attention weight is: Among them, O i is the output at the i-th pixel position, Q is the query vector, K is the key vector, V is the value vector, and e is a natural constant. T For transposition, is the square root of the channel dimension of the key vector, i is the index of the query vector, j is the index of the value vector, k is the index of the key vector, N is H×W, H is the height of the input feature of the dehazing module, and W is the width of the input feature of the dehazing module; After the attention weights are dimensionally converted, they are upsampled to obtain the attention weights after dimension conversion and upsampling; After the dimension conversion and upsampling, the attention weight is multiplied by the learnable parameter and then added to the input feature of the dehazing module to obtain the output enhanced feature map of the image feature enhancement unit. The formula for obtaining the output enhanced feature map of the image feature enhancement unit by multiplying the attention weight after dimension conversion and upsampling with the learnable parameter and adding it to the input feature of the defogging module is: in, is the output enhanced feature map of the image feature enhancement unit, γ is a learnable parameter, O is the attention weight after dimensional conversion and upsampling, and I0 is the input feature of the dehazing module.
6. The image defogging method based on generative adversarial network according to claim 4, characterized in that: The output enhancement feature map of the image feature enhancement unit is input into the atmospheric scattering matrix estimation branch, and the atmospheric light matrix is output, including: The output enhanced feature map of the image feature enhancement unit is subjected to mixed average pooling to obtain a pooled enhanced feature map; The pooled enhanced feature map is processed by a 3×3 convolution layer and a ReLU activation function in sequence, and then concatenated with the pooled enhanced feature map to obtain the initial fusion feature. The initial fusion features are processed by 3×3 convolutional layers and ReLU activation functions in sequence, and then concatenated with the initial fusion features to obtain secondary fusion features. The secondary fusion features are processed by 1×1 convolution layer and Sigmoid activation function in sequence, and then added element by element with the initial fusion features to obtain the atmospheric light pre-generation features; After upsampling the pre-generated features of atmospheric light, an atmospheric light matrix is obtained.
7. The image defogging method based on generative adversarial network according to claim 4, characterized in that: The step of inputting the output enhanced feature map of the image feature enhancement unit into the transmission matrix estimation branch and outputting the transmission map matrix comprises: The output enhanced feature map of the image feature enhancement unit passes through a 3×3 convolution layer to obtain a feature map processed by a 3×3 convolution layer; The feature map processed by the 3×3 convolution layer is processed by the 3×3 convolution layer and the ReLU activation function in sequence, and then concatenated with the feature map processed by the 3×3 convolution layer to obtain the first spliced feature; The first stitched feature is processed by a 3×3 convolution layer and a ReLU activation function in sequence, and then concatenated with the first stitched feature to obtain a second stitched feature; The second stitched feature is processed by a 3×3 convolution layer and a ReLU activation function in sequence, and then concatenated with the second stitched feature to obtain a third stitched feature; The third stitching feature is processed by a 1×1 convolution layer and a Sigmoid activation function in sequence, and then added element by element to the feature map processed by a 3×3 convolution layer to obtain a transmission map matrix.
8. The image defogging method based on generative adversarial network according to claim 4, characterized in that: The defogging feature map is calculated based on the atmospheric light matrix, the transmission map matrix, the output enhancement feature map of the image feature enhancement unit, and the input features of the defogging module. The calculation formula is: in, is the dehazing feature map, is the output enhanced feature map of the image feature enhancement unit, ⊙ is the element-by-element multiplication, is the transmission map matrix, is the atmospheric light matrix, It is the input feature of the dehazing module.
9. The image defogging method based on generative adversarial network according to claim 4, characterized in that: The defogging image of the synthetic fog paired image, the atmospheric light matrix output by the atmospheric scattering matrix estimation branch in the defogging unit of the last defogging module, and the transmission map matrix output by the transmission matrix estimation branch in the defogging unit of the last defogging module are passed through the image reconstruction module to obtain the reconstructed fog image. The calculation formula for the reconstructed fog image is: Among them, I rec To reconstruct the fog image, is the defogging image of the synthetic fog paired image, ⊙ is the element-by-element multiplication, is the transmission map matrix output by the transmission matrix estimation branch in the defogging unit of the last defogging module, It is the atmospheric light matrix output by the atmospheric scattering matrix estimation branch in the defogging unit of the last defogging module. for The inverse matrix of It is the dehazing feature map output by the last dehazing module.
10. The image defogging method based on generative adversarial network according to claim 4, characterized in that: The image to be defogged is input into the target generative adversarial network, and the defogged image of the image to be defogged is output, including: The image to be dehazed is input into the front feature extractor to extract the feature map of the image to be dehazed, including: After the image to be dehazed passes through a 3×3 convolutional layer, it is downsampled twice by a factor of 2 to obtain a low-resolution feature map of the image to be dehazed; After the low-resolution feature map of the image to be dehazed passes through three residual modules connected in sequence, it is upsampled twice by a factor of 2 to obtain the feature map of the image to be dehazed; The feature map of the image to be dehazed is input into the dehazing network, and the dehazed image of the image to be dehazed is output.
Citation Information
Patent Citations
End-to-end image defogging method based on multi-feature fusion
CN114742719A
Image defogging method and system based on generative adversarial network and multi-scale fusion
CN115457265A