Image Dehazing Method and System Based on Improved Generative Adversarial Network
Through the improved generative adversarial network combined with dark channel prior algorithm and generative adversarial learning method, the problems of insufficient feature extraction capabilities and details loss during image defog removal are solved, and efficient image defog removal effect is achieved, including real color recovery and detailed texture retention.
Patent Information
- Application Number
- CN202410336917.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-03-22
AI Technical Summary
The prior art has problems such as insufficient feature extraction capability, difficulty in restoring the real color, loss of detailed texture information, and noise artifacts in the process of image defogging.
The image defog method based on an improved generative adversarial network is adopted, and a two-stage defog strategy is adopted: first, the dark channel prior algorithm is used for preliminary defog removal, and then the generative adversarial learning method is used to refine the processing, combining the dense connection module, shallow feature fusion module, spatial attention module and mixed loss function, the model parameters are optimized to improve the defog removal effect.
It effectively improves the color authenticity, retention of detail texture information and halo artifact removal of the defog result map, improves the image defog effect, and is suitable for subsequent advanced visual tasks.
Smart Images

Figure CN118333898B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image defogging, and specifically to an image defogging method and system based on an improved generative adversarial network. Background Art
[0002] Haze is a common atmospheric phenomenon, mainly generated when light scatters on the surface of aerosol particles in the atmosphere. Images taken in haze weather will have problems such as blurring, information loss, and contrast decline, which will pose potential hazards to subsequent advanced vision tasks. Therefore, image defogging technology has important research value and is of great significance in fields such as video surveillance, drone photography, and autonomous driving assistance technology. In recent years, relevant scholars have made many efforts to better restore clear images from foggy images.
[0003] For example, the existing invention patent application document "A Two-Step Hyperspectral Image Defogging Method Based on Physics and Frequency Domain Guidance" with the publication number CN117522739A. The existing method includes: constructing a training data set; constructing a defogging model for the first-step training and an improved RGB defogging model for the second-step training; designing loss functions for the models of the first step and the second step and training the two-step models; performing image defogging through the trained two-step models. And the existing invention patent application document "Foggy Weather Traffic Sign Recognition Method, Device, Equipment and Storage Medium Based on Deep Learning" with the publication number CN117593724A. The existing method includes: dividing a foggy picture into a sky area and a non-sky area, and calculating the transmittance of the two areas respectively; then fusing the transmittance of the above two areas, and performing defogging processing on the picture according to the atmospheric scattering model; finally, inputting the picture after defogging processing into the EfficientDet target detection network to identify the traffic sign. However, the foregoing existing image defogging methods mainly focus on research based on deep learning, lacking research on the physical essence of foggy images, and there are still some problems: there are still deficiencies in color authenticity and image feature extraction ability, and the restored clean images are prone to problems such as loss of detailed texture information and halo artifacts. Therefore, how to construct an efficient image defogging method that restores true colors, retains detailed textures, and removes noise artifacts is a key problem that needs to be solved urgently by researchers in this field.
[0004] The existing public literature "Low-light Dehazing Algorithm Based on Attention Mechanism, Dense Residual Fusion and Spatial Local Filtering". In this existing method, the dense residual block is first used to increase the depth of the neural network, enabling the network to learn more advanced feature information. Then, the spatial and channel attention mechanisms are introduced to filter and screen the features, enabling the network to distinguish uneven lighting areas and solve problems such as color distortion. The method of enhancing spatial local filtering is adopted to improve the contrast, clarity and visibility of the dehazing result. Finally, a joint loss function is designed to constrain the learning of the network. Among them, this existing technology uses DAF-DehazeNet to process low-light haze images, solving problems such as color distortion, patches and artifacts. Although this method can make the dehazing result close to a clear image, there is a problem that the target is not clear due to poor contrast.
[0005] The existing public literature "Vehicle Detection in Haze Environment Based on Multi-scale Feature Fusion". In this existing method, the conditional generative adversarial network is first used to preprocess the haze image. Then, aiming at the characteristics of unclear target features in the haze environment, a multi-scale feature fusion module is proposed. Based on YOLOv3, when extracting features from the backbone network, a shallow branch and deep features are added for upsampling and splicing fusion to obtain a feature map with a scale of 104×104, which is used to enhance the semantic information of the shallow layer. And a feature enhancement strategy guided by the CBAM attention mechanism is adopted to ensure the integrity of the context information to improve the detection accuracy. Finally, the dehazed image is sent into the improved YOLOv3 network for detection. In order to reduce the missed detection rate of target detection in a foggy environment, the conditional generative adversarial network is used for image dehazing. However, due to the relatively simple network structure, its learning ability is limited, and there are problems of incomplete dehazing effect and very serious edge artifacts and checkerboard artifacts.
[0006] The existing public literature "Image Dehazing Based on Channel Attention and Conditional Generative Adversarial Network". The generator of this existing method designs a multi-scale dense residual network to extract image features with convolutional kernels of different scales, reducing information loss. The dense residual network deepens the network depth, fully excavates the depth features of the foggy image, adds skip connections to improve the feature utilization rate, and avoids gradient disappearance. The attention mechanism is introduced to dynamically adjust the weights of different-scale information during feature fusion, strengthening the propagation of effective features and reducing the model complexity. The discriminator uses a fully convolutional network to block and discriminate images, paying attention to the detailed information of each region of the image to improve the discrimination ability. However, there are still obvious edge artifact problems in the dehazing results of the aforementioned existing solutions. In addition, 5×5 large-scale convolutional kernels are used in the multi-scale residual module, resulting in too many network training parameters and high training costs.
[0007] In summary, the existing technologies have technical problems such as insufficient feature extraction ability, difficulty in restoring true colors, loss of detailed texture information, and the presence of noise artifacts. Summary of the Invention
[0008] The technical problem to be solved by the present invention is: how to solve the technical problems of insufficient feature extraction ability, difficulty in restoring true colors, loss of detailed texture information, and the existence of noise artifacts in the prior art.
[0009] The present invention solves the above technical problems by adopting the following technical solutions: An image dehazing method based on an improved generative adversarial network includes:
[0010] S1. Collect and construct a hazy image dataset, preprocess the hazy images in the hazy image dataset to a preset size, and thereby establish a training set. The training set includes: hazy images and corresponding clear images;
[0011] S2. Input the hazy images in the training set into the dark channel prior algorithm for processing to obtain a first dehazed image;
[0012] S3. Construct a generative adversarial network model based on the U-Net network, retain 4 sets of upsampling and downsampling to retain more key feature information in the image while learning feature maps of different resolutions. The generative adversarial network model includes: a generator and a discriminator, wherein the generator includes: an encoder with an added dense connection module DCB, a shallow feature fusion module SFFM, a spatial attention module SAM, and a decoder;
[0013] S4. Input the first dehazed image and the corresponding clear image into the generative adversarial network model for refinement processing, and perform a training operation on the model parameters of the generative adversarial network model to obtain a dehazed result image;
[0014] S5. Construct an image dehazing hybrid loss function to supervise the training operation and optimize the model parameters to update and obtain an applicable image dehazing model, wherein the hybrid loss function includes: a generative adversarial loss, a reconstruction loss, and a consistency loss;
[0015] S6. Collect and input a real-time hazy image into the applicable image dehazing model to obtain and output a real-time dehazed image.
[0016] The present invention adopts a two-stage dehazing strategy. First, the dark channel prior algorithm is used to improve the visibility of the image, and then the generative adversarial learning method is used to improve the authenticity of the image. It combines the advantages of the prior-based and generative adversarial learning methods and can effectively improve the image dehazing effect. The present invention improves the color authenticity, the retention degree of detailed texture information, and the removal degree of halo artifacts of the dehazed result image, which is helpful for the realization of subsequent high-level vision tasks.
[0017] In a more specific technical solution, step S3 includes:
[0018] S31. Input the foggy image into the generator. In the encoding stage, use the dense connection module DCB to perform no less than 2 downsampling operations to adjust the channel dimension of the foggy image, and obtain no less than 2 feature maps with different downsampling sizes respectively;
[0019] S32. Pass the foggy image with the original size through a convolutional channel transformation to obtain and input the shallow feature information into the shallow feature fusion module SFFM for learning operations, so as to obtain and fuse the shallow output features and the upsampled output of the model, and use the different receptive fields to process and obtain the first rich semantic information;
[0020] S33. Send the feature maps with different downsampling sizes into the spatial attention module SAM respectively for processing the high-frequency regions of the image, and fuse them with the corresponding scale feature maps of each layer of the decoder, so as to obtain the second rich semantic information and the global context information according to the first rich semantic information, and obtain the dense output features;
[0021] S34. In the decoder, perform a recovery operation on the dense output features through no less than 2 upsampling modules to obtain the de-checkerboard artifact image with the original size. Among them, use the combination of bilinear interpolation and 3×3 convolution to replace the conventional transposed convolution operation in the decoder for upsampling operations to avoid the checkerboard artifact phenomenon.
[0022] The SFFM module adopted by the present invention adaptively generates different weights by inputting multi-scale information to guide the Softmax activation function. The weights generated by the two groups of branches obtain different receptive fields in the feature fusion layer, thus making up for the lack of semantic information of the shallow features; the SAM module adopted by the present invention uses a parallel strategy, convolutional layers and operations of multiplying with the input features, making the network pay more attention to the high-frequency regions of the image.
[0023] The fog model provided by the present invention uses the spatial attention module SAM (Spatial Attention Module), making the model pay more attention to the high-frequency regions of the image during the defogging process, effectively reducing the loss of image texture and details.
[0024] The present invention adopts the combination of bilinear interpolation and 3×3 convolution to replace the conventional transposed convolution operation in the decoder, thereby reducing the checkerboard artifact phenomenon and obtaining better visual effects.
[0025] In a more specific technical solution, the dense connection components in the dense connection module DCB include: a convolutional layer and a multi-scale convolutional unit MSCU, which are respectively used for channel transformation operations and multi-scale feature learning operations.
[0026] The dehazing model proposed by the present invention uses the method of stacking multiple small-sized convolutional kernels instead of large-sized convolutional kernels in the densely connected block (DCB). On the premise of the same receptive field, this method effectively reduces the number of parameters of the model, thus saving the training cost. Through this improvement, the method provided by the present invention effectively reduces the training burden of the network while improving the network performance.
[0027] In a more specific technical solution, the shallow feature fusion module (SFFM) guides the Softmax activation function to adaptively generate the differential branch weights according to the multi-scale information, so as to obtain different receptive fields in the feature fusion layer.
[0028] The fog model of the present invention proposes a shallow feature fusion module (SFFM). This module selects and learns the features of the original-size image before compression, and fuses the shallow output features with the upsampled output of the model to obtain richer global structure information. This module effectively improves the contrast of the image, makes the overall color more real and natural, and further improves the dehazing effect.
[0029] In a more specific technical solution, step S5 includes:
[0030] S51. Define and update the operations by supervising the generator and discriminator according to the generative adversarial loss;
[0031] S52. Define the reconstruction loss according to the reconstructed hazy image to update the generator;
[0032] S53. Define the consistency loss according to the artifact and noise parameters to keep the dehazed image consistent with the corresponding clear image;
[0033] S54. Determine the hybrid loss function according to the generative adversarial loss, reconstruction loss and consistency loss.
[0034] The present invention first preliminarily dehazes the input hazy image through the dark channel prior algorithm to obtain the atmospheric light value A and the transmittance t. Then, the preliminary dehazing result map is obtained by using the atmospheric scattering model, and a generative adversarial network model is constructed. By constructing a new hybrid loss function, the model is supervised for training. Finally, a stable optimal image dehazing model is obtained. Specifically, the present invention adds a consistency loss function to the hybrid loss function. This loss function can reduce the visual difference between the dehazing result and the clear image, help to alleviate the edge artifact problem, and thus obtain a dehazing result closer to the clear image.
[0035] In a more specific technical solution, in step S51, the generative adversarial loss L is defined by using the following logic G :
[0036]
[0037] Wherein, J real represents the clear image corresponding to the input foggy image, and J DCP represents the initially de-fogged image obtained by dark channel prior. E represents the mathematical expectation, G represents the generator, and D represents the discriminator.
[0038] In a more specific technical solution, in step S52, the distance L real between the input foggy image I rec and the reconstructed foggy image I 1 is minimized to define the reconstruction loss:
[0039] L rec = ||I real - I rec ||.
[0040] In a more specific technical solution, in step S53, the consistency loss L idt is defined using the following logic:
[0041] L idt = ||J real - G(J real )|| + ||J real - J ref ||;
[0042] Wherein, J ref represents the de-fogged result image.
[0043] In view of the problem that the generative adversarial network introduced artifacts and noise, which visually affected the image de-fogging, the present invention proposes a consistency loss function to make the de-fogged image consistent with the reference clear image.
[0044] In a more specific technical solution, in step S54, the mixed loss function is determined using the following logic:
[0045] L all = λ 1 L G + λ 2 L rec + λ 3 L idt ;
[0046] Wherein, λ i (i = 1, 2, 3) are hyperparameters representing weights, and λ 1 , λ 2 , λ 3 are empirical values.
[0047] In a more specific technical solution, an image defogging system based on an improved generative adversarial network includes:
[0048] A training set construction module for collecting and constructing a foggy image data set, preprocessing the foggy images in the foggy image data set to a preset size, and thereby establishing a training set, where the training set includes: foggy images and corresponding clear images;
[0049] A preliminary defogging module for inputting the foggy images in the training set into the dark channel prior algorithm for processing to obtain a first defogged image. The preliminary defogging module is connected to the training set construction module;
[0050] A generative adversarial network model construction module for constructing a generative adversarial network model using a U-Net network, retaining 4 sets of upsampling and downsampling to retain more key feature information in the image while learning feature maps of different resolutions. The generative adversarial network model includes: a generator and a discriminator, where the generator includes: an encoder with an added dense connection module DCB, a shallow feature fusion module SFFM, a spatial attention module SAM, and a decoder. The generative adversarial network model construction module is connected to the training set construction module;
[0051] A model training module for inputting the first defogged image and the corresponding clear image into the generative adversarial network model for refinement processing, and performing a training operation on the model parameters of the generative adversarial network model to obtain a defogging result image. The model training module is connected to the generative adversarial network model construction module and the preliminary defogging module;
[0052] A model training supervision and optimization module for constructing an image defogging hybrid loss function, supervising the training operation, and optimizing the model parameters to update and obtain a suitable image defogging model. Among them, the hybrid loss function includes: a generative adversarial loss, a reconstruction loss, and a consistency loss. The model training supervision and optimization module is connected to the model training module;
[0053] A real-time defogging module for collecting and inputting real-time foggy images into the suitable image defogging model to obtain and output real-time defogged images. The real-time defogging module is connected to the model training supervision and optimization module.
[0054] The present invention has the following advantages compared with the prior art:
[0055] The present invention adopts a two-stage defogging strategy. First, it uses the dark channel prior algorithm to improve the image visibility, and then uses the generative adversarial learning method to improve the image authenticity. It combines the advantages of the prior-based and generative adversarial learning methods and can effectively improve the image defogging effect. The present invention improves the color authenticity, detail texture information retention degree, and halo artifact removal degree of the defogging result image, which is helpful for the realization of subsequent high-level vision tasks.
[0056] The upsampling module of the present invention uses a combination of bilinear interpolation and 3×3 convolution to replace the conventional transposed convolution operation in the decoder, avoiding the checkerboard artifact phenomenon.
[0057] The SFFM module adopted by the present invention adaptively generates different weights by guiding the Softmax activation function with input multi-scale information. The weights generated by the two groups of branches obtain different receptive fields in the feature fusion layer, thus making up for the lack of semantic information in the shallow features. The SAM module adopted by the present invention uses a parallel strategy, convolutional layers, and operations of multiplying with the input features, making the network pay more attention to the high-frequency regions of the image.
[0058] The present invention first preliminarily removes haze from the input hazy image through the dark channel prior algorithm to obtain the atmospheric light value A and the transmittance t. Then, the preliminary haze-removal result image is obtained using the atmospheric scattering model, and a generative adversarial network model is constructed. By constructing a new hybrid loss function, the model is supervised for training. Finally, a stable and optimal image haze-removal model is obtained.
[0059] Aiming at the problem that the generative adversarial network introduces artifacts and noise, which visually affects image haze removal, the present invention proposes a consistency loss function to make the haze-removal image consistent with the reference clear image.
[0060] The present invention solves the technical problems existing in the prior art, such as insufficient feature extraction ability, difficulty in restoring true colors, loss of detailed texture information, and the existence of noise artifacts. Brief Description of the Drawings
[0061] Figure 1 It is a schematic diagram of the basic steps of the image haze-removal method based on an improved generative adversarial network according to Embodiment 1 of the present invention;
[0062] Figure 2 It is a schematic diagram of the overall model framework of the image haze-removal method based on an improved generative adversarial network according to Embodiment 1 of the present invention;
[0063] Figure 3 It is a schematic diagram of the model structure of the generator according to Embodiment 1 of the present invention;
[0064] Figure 4 It is a schematic diagram of the structure of the dense connection module DCB in the generative adversarial network model according to Embodiment 1 of the present invention;
[0065] Figure 5 It is a schematic diagram of the structure of the multi-scale convolutional unit MSCU in the dense connection module of the generative adversarial network model according to Embodiment 1 of the present invention;
[0066] Figure 6 It is a schematic diagram of the structure of the shallow feature fusion module SFFM in the generative adversarial network model according to Embodiment 1 of the present invention;
[0067] Figure 7 Schematic diagram of the structure of the Spatial Attention Fusion Module SAM in the generative adversarial network model of Embodiment 1 of the present invention;
[0068] Figure 8 Schematic diagram of the specific implementation steps of the defogging process in Embodiment 1 of the present invention;
[0069] Figure 9 Schematic diagram of an example of an outdoor foggy image and its defogged result in Embodiment 2 of the present invention. Specific implementation manners
[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0071] Embodiment 1
[0072] As Figure 1 shown, the image defogging method based on an improved generative adversarial network provided by the present invention includes the following basic steps:
[0073] Step S1: To reduce the calculation time and memory consumption, the obtained public indoor and outdoor foggy image datasets are preprocessed to a size of 3×256×256, and a training set is established. The training set includes paired foggy and corresponding clear images;
[0074] Step S2: The foggy images in the training set are input into the dark channel prior algorithm to obtain a preliminary defogged image;
[0075] In this embodiment, first, the atmospheric light value A and the transmittance t are obtained using the dark channel prior algorithm, and then the preliminary defogged image is obtained using the atmospheric scattering model;
[0076] Step S3: An improved generative adversarial network model including a generator and a discriminator is constructed;
[0077] As Figure 2 shown, in this embodiment, the generator is composed of an encoder, a Shallow Feature Fusion Module (SFFM), a Spatial Attention Module (SAM), and a decoder. In this embodiment, the U-Net network model is used as the basic model framework of the generator, and 4 times of upsampling and downsampling are retained to retain more key feature information in the image while learning feature maps of different resolutions. On this basis, the DCB, SFFM, and SAM modules are added;
[0078] In this embodiment, after the input image enters the generator, it first undergoes four downsamplings with dense connection blocks (DCBs) during the encoding stage, obtaining feature maps with sizes of 128×128×128, 256×64×64, 512×32×32, and 1024×16×16 respectively.
[0079] As Figure 3 shown, in this embodiment, in order to better achieve multi-level information interaction between the encoder and decoder and compensate for the lost detailed information during the encoding process, a shallow feature fusion module (SFFM) and a spatial attention module (SAM) are designed; in this embodiment, after the original-size image features undergo channel transformation through a single convolutional layer, they are input into the SFFM module for learning, and the shallow output features are fused with the upsampled output of the model to obtain richer semantic information; in this embodiment, the image features with sizes of 128×128×128, 256×64×64, and 512×32×32 are respectively fed into the SAM module and fused with the corresponding-scale feature maps of each layer of the decoder, thereby obtaining richer semantic information and global context information;
[0080] In this embodiment, the dense output features are restored to an RGB image with the same size as the input image through four upsamplings; in this embodiment, a combination of bilinear interpolation and 3×3 convolution is used to replace the conventional transposed convolution operation in the decoder for upsampling, avoiding the checkerboard artifact phenomenon;
[0081] As Figure 4 and Figure 5 shown, in this embodiment, the DCB module includes four identical parts, each of which includes but is not limited to: a convolutional layer and a multi-scale convolutional unit (MSCU), which are respectively used for channel transformation and learning of multi-scale features. For the DCB module structure, refer to Figure 4 and for the MSCU module structure, refer to Figure 5 ;
[0082] In this embodiment, when the entire module processes the feature map, it does not change its size but only changes the channel dimension; in this embodiment, if the number of input channels is N, the output of each part is N / 4, and the output of the final dense connection module is 2N;
[0083] As Figure 6 shown, in this embodiment, the SFFM module adaptively generates different weights by guiding the Softmax activation function with input multi-scale information, and the weights generated by the two groups of branches obtain different receptive fields in the feature fusion layer, thereby compensating for the deficiency of semantic information in the shallow features;
[0084] As Figure 7As shown, in this embodiment, the SAM module makes the network pay more attention to the high-frequency regions of the image, thus compensating for the detailed information lost during the encoding process of the generator.
[0085] Step S4: Input the preliminary haze-removed image and the corresponding clear image in the training set into the generative adversarial network for refinement, train the model parameters, and obtain the haze-removed result image.
[0086] In this embodiment, the clear image corresponding to the hazy image and the haze-removed result image are simultaneously input into the discriminator, enabling adversarial learning with the generator to iteratively update the model parameters to obtain a haze-removed result image containing more detailed information.
[0087] Step S5: Construct an image haze-removal hybrid loss function to supervise the model training, optimize the model parameters, and thereby update and obtain an optimal image haze-removal model that tends to be stable.
[0088] In this embodiment, the hybrid loss function includes: generative adversarial loss, reconstruction loss, and consistency loss. The sub-loss functions are as follows:
[0089] In this embodiment, the generative adversarial loss is used to supervise the generator G and the discriminator D to update in an adversarial manner, expressed as:
[0090]
[0091] where J real represents the clear image corresponding to the input hazy image, and J DCP represents the preliminary haze-removed image obtained through dark channel prior.
[0092] In this embodiment, the reconstruction loss is defined as the L real distance between the input hazy image I rec and the reconstructed hazy image I 1 . The generator G is updated by minimizing the distance between I real and I rec , expressed as:
[0093] L rec = ||I real - I rec ||;
[0094] In this embodiment, aiming at the problem that introducing artifacts and noise in the generative adversarial network affects image haze removal visually, a consistency loss function is proposed to make the haze-removed image consistent with the reference clear image. The consistency loss is expressed as:
[0095] L idt = ||J real - G(J real )|| + ||J real-J ref ||;
[0096] Among them, J ref represents the defogged result image;
[0097] In this embodiment, the formula of the total hybrid loss function is as follows:
[0098] L all = λ 1 L G + λ 2 L rec + λ 3 L idt ;
[0099] Among them, λ i (i = 1, 2, 3) is a hyperparameter representing the weight, λ 1 and λ 2 take empirical values of 0.02, and λ 3 takes an empirical value of 1.
[0100] Step S6: Input the new single foggy image into the optimal image defogging network model and directly output the defogged image;
[0101] As Figure 8 shown, in this embodiment, by inputting the foggy image, image defogging can be achieved without other information, and it can adapt to input images of different sizes, and the size of the output image will be consistent with it. In this embodiment, in the operation of using the optimal image defogging network model to process and output the defogged image;
[0102] Step S61: Input the new single foggy image into the dark channel prior algorithm for preliminary defogging;
[0103] Step S62: Input the preliminarily defogged image into the encoder network of the generator. After four downsamplings including the DCB module, the obtained multi-scale features are respectively passed through the SFFM and SAM modules, and feature fusion is performed with the upsampling outputs of each layer of the decoder;
[0104] Step S63: Perform channel transformation through the convolutional layer to obtain a defogged result map with a size of 3×256×256.
[0105] Embodiment 2
[0106] As Figure 9As shown, in this embodiment, to further illustrate the effectiveness of the present invention, the Indoor Training Set (ITS) and Outdoor Training Set (OTS) of RESIDE are preprocessed respectively in this embodiment as the training set to train the network model in the present invention, and the trained network is used to perform image defogging test operations on two groups of indoor test sets (SOTS-indoor, D-HAZY) and one group of outdoor test set (SOTS-outdoor).
[0107] In this embodiment, 90 rounds of training are preset in the experiment, the Adam optimizer is used, and the initial learning rate of the neural network is 2×10 -4 , which decays exponentially and gradually decreases to 2×10 -6 during the iteration process. When the input image is randomly cropped to 3×256×256 pixels and the batch size is 1, the indexes obtained by the method of the present invention (denoted by DCFFA-Net, Densely Connected Feature Fusion Attention Network) and seven methods including FVR, BCCR, DCP, CAP, AOD-Net, MSCNN, and EPDN are compared on the indoor and outdoor test sets of SOTS. The PSNR and SSIM indexes obtained by using the method of the present invention on the indoor test set reach 25.95dB and 0.9225 respectively, which are 4.4dB and 0.0154 higher than the best performer among the comparison methods respectively; the PSNR and SSIM indexes obtained on the outdoor test set reach 24.21dB and 0.9260 respectively, which are 1.89dB and 0.0577 higher than the best performer among the comparison methods respectively. It shows that the image defogging performance of the method of the present invention is the best, the defogged image has the lowest distortion degree compared with the original clear image, has the best performance in terms of color, brightness and contrast, and has the strongest robustness and generalization ability.
[0108] In summary, the present invention adopts a two-stage defogging strategy. First, the dark channel prior algorithm is used to improve the image visibility, and then the generative adversarial learning method is used to improve the image authenticity. It combines the advantages of the prior-based and generative adversarial learning methods and can effectively improve the image defogging effect. The present invention improves the color authenticity, the retention degree of detail texture information and the removal degree of halo artifacts of the defogging result map, which is helpful for the realization of subsequent high-level vision tasks.
[0109] The upsampling module of the present invention uses a combination of bilinear interpolation and 3×3 convolution to replace the conventional deconvolution operation in the decoder, avoiding the checkerboard artifact phenomenon.
[0110] The SFFM module adopted by the present invention guides the Softmax activation function to adaptively generate different weights by inputting multi-scale information. The weights generated by the two groups of branches obtain different receptive fields in the feature fusion layer, thereby making up for the lack of semantic information in the shallow features; the SAM module adopted by the present invention uses a parallel strategy, convolutional layers and operations of multiplying with the input features, making the network pay more attention to the high-frequency regions of the image.
[0111] The present invention first preliminarily removes the haze from the input hazy image through the dark channel prior algorithm to obtain the atmospheric light value A and the transmittance t. Then, the atmospheric scattering model is used to obtain the preliminary haze-removal result image, and a generative adversarial network model is constructed. By constructing a new hybrid loss function, the model is supervised for training. Finally, a stable optimal image haze-removal model is obtained.
[0112] Aiming at the problem that the generative adversarial network introduces artifacts and noises, which affect the haze removal of the image visually, the present invention proposes a consistency loss function to make the haze-removed image consistent with the reference clear image.
[0113] The present invention solves the technical problems existing in the prior art, such as insufficient feature extraction ability, difficulty in restoring true colors, loss of detailed texture information, and the existence of noise artifacts.
[0114] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. The image dehazing method based on improved generative adversarial network is characterized by: include: S1. Collect and construct a foggy image dataset, preprocess the foggy images in the foggy image dataset to a preset size, and establish a training set, where the training set includes: foggy images and corresponding clear images; S2, passing the foggy image in the training set into the dark channel prior algorithm for processing to obtain a first defogging image; S3. Construct a generative adversarial network model based on the U-Net network, retaining no less than 2 groups of up and down sampling, so as to retain the key feature information in the image while learning feature maps of different resolutions. The generative adversarial network model includes: a generator and a discriminator, wherein the generator includes: an encoder with a dense connection module DCB, a shallow feature fusion module SFFM, a spatial attention module SAM and a decoder; Step S3 includes: S31, inputting the foggy image into the generator, and in the encoding stage, performing downsampling operations at least twice using the dense connection module DCB, adjusting the channel dimension of the foggy image, and obtaining at least two difference downsampling size feature maps respectively; S32, transforming the foggy image of the original size through a layer of convolution channel to obtain and input shallow feature information into the shallow feature fusion module SFFM for learning operation, so as to obtain and fuse the shallow output features and the model upsampled output, so as to obtain the first rich semantic information by using the difference receptive field processing; S33, sending the difference downsampled size feature maps to the spatial attention module SAM for image high-frequency area processing, so as to be fused with the corresponding scale feature maps of each layer of the decoder, thereby obtaining the second rich semantic information and the global context information according to the first rich semantic information, so as to obtain dense output features; S34. In the decoder, performing a restoration operation on the dense output features through at least two upsampling modules to obtain a checkerboard artifact-free image of the original size, wherein the upsampling operation is performed using a combination of bilinear interpolation and 3×3 convolution; The densely connected components in the densely connected module DCB include: convolutional layers and multi-scale convolutional units MSCU, which are used to perform channel transformation operations and multi-scale feature learning operations respectively; S4, passing the first defogging image and the corresponding clear image into the generative adversarial network model for refinement processing, performing a training operation on the model parameters of the generative adversarial network model, and obtaining a defogging result image; S5. Construct an image defogging hybrid loss function, supervise the training operation, and optimize the model parameters to update and obtain a suitable image defogging model. The hybrid loss function includes: generating adversarial loss L G , reconstruction loss L rec And the consistency loss L idt ; Wherein, step S5 comprises: S51. Define and update operations based on the generated adversarial loss to supervise the generator and the discriminator; S52, defining a reconstruction loss according to the reconstructed foggy image to update the generator; S53, defining a consistency loss according to artifact and noise parameters to keep the defogging image consistent with the corresponding clear image; wherein the consistency loss is defined using the following logic: L idt =||J real -G(J real )||+||J real -J ref ||; In the formula, J ref represents the dehazed image, J real Indicates the clear image corresponding to the input foggy image; S54, determining a hybrid loss function according to the generated adversarial loss, the reconstruction loss, and the consistency loss; S6. Collect and input a real-time foggy image into an applicable image defogging model to obtain and output a real-time defogging image.
2. The image dehazing method based on the improved generative adversarial network according to claim 1 is characterized in that: The shallow feature fusion module SFFM guides the Softmax activation function to adaptively generate difference branch weights according to multi-scale information, so as to obtain the difference receptive field in the feature fusion layer.
3. The image dehazing method based on the improved generative adversarial network according to claim 1 is characterized in that: In step S51, the generative adversarial loss is defined using the following logic: In the formula, J real represents the clear image corresponding to the input foggy image, J DCP Represents the preliminary dehazed image after dark channel prior.
4. The image dehazing method based on the improved generative adversarial network according to claim 1, characterized in that: In step S52, the input foggy image I is minimized. real and reconstruct the foggy image I rec The L1 distance between them is used to define the reconstruction loss: L rec =||I real -I rec ||。 5. The image dehazing method based on the improved generative adversarial network according to claim 1, characterized in that: In step S54, the hybrid loss function is determined using the following logic: L all =λ1L G +λ2L rec +λ3L idt ; In the formula, λ i (i=1,2,3) is a hyperparameter representing the weight, and λ1, λ2, and λ3 are empirical values.
6. An image defogging system based on an improved generative adversarial network, used to execute the image defogging method based on an improved generative adversarial network according to any one of claims 1 to 5, characterized in that: The system comprises: A training set construction module is used to collect and construct a foggy image data set, preprocess the foggy images in the foggy image data set to a preset size, and establish a training set, wherein the training set includes: the foggy images and corresponding clear images; A preliminary defogging module, used for inputting the foggy image in the training set into a dark channel priori algorithm for processing to obtain a first defogging image, wherein the preliminary defogging module is connected to the training set construction module; A generative adversarial network model construction module, which uses a U-Net network to construct a generative adversarial network model, retaining no less than 2 groups of up and down sampling, so as to retain more key feature information in the image while learning feature maps of different resolutions. The generative adversarial network model includes: a generator and a discriminator, wherein the generator includes: an encoder with a dense connection module DCB, a shallow feature fusion module SFFM, a spatial attention module SAM and a decoder, and the generative adversarial network model construction module is connected to the training set construction module; A model training module, used for transferring the first defogging image and the corresponding clear image into the generative adversarial network model for refinement, performing training operations on the model parameters of the generative adversarial network model, and obtaining a defogging result image, wherein the model training module is connected to the generative adversarial network model construction module and the preliminary defogging module; A model training supervision optimization module is used to construct an image defogging mixed loss function, supervise the training operation, and optimize the model parameters to update and obtain a suitable image defogging model, wherein the mixed loss function includes: generating adversarial loss, reconstruction loss, and consistency loss, and the model training supervision optimization module is connected to the model training module; The real-time defogging module is used to collect and input the real-time foggy image into the applicable image defogging model to obtain and output the real-time defogging image. The real-time defogging module is connected to the model training supervision optimization module.
Citation Information
Patent Citations
Two-step hyperspectral image defogging method based on physical and frequency domain guidance
CN117522739A
Foggy day traffic sign identification method and device based on deep learning, equipment and storage medium
CN117593724A
Image defogging method and system based on generative adversarial network and multi-scale fusion
CN115457265A