Two-Stage Unsupervised Dehazing Method Based on Pseudo-Haze Images

By constructing a pseudo-haze image generation framework and a three-branch network, combining atmospheric scattering model and dark channel prior, the problem of poor fog removal effect of haze image is solved, and more realistic fog removal effect and network generalization ability are improved.

CN120047360BActive Publication Date: 2025-07-18HUNAN UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510526550.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-18
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing defog removal method has poor effect on haze images, especially due to the domain differences between the synthetic domain and the real domain, and the existing unsupervised methods have failed to effectively consider the change in haze concentration with depth, resulting in the lack of authenticity of the generated pseudo-haze images.

Method used

A two-stage unsupervised fog removal method based on pseudo-haze images is constructed, including a pseudo-haze image generation framework and a three-branch network. Through steps such as depth estimation, haze transfer and transmittance estimation, combined with atmospheric scattering model and dark channel priors, the unpaired real haze-clear image pairs are used to learn and train, reduce domain differences and improve network generalization capabilities.

Benefits of technology

The defog removal effect of haze images and the generalization ability of the network are improved, the generated pseudo-haze images are more realistic, and the network performance and interpretability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047360B_ABST
    Figure CN120047360B_ABST
Patent Text Reader

Abstract

The present invention discloses a two-stage unsupervised dehazing method based on pseudo-haze images, belonging to the technical field of computer image processing. The method includes the following steps: establishing a training data set containing haze images and unpaired clear images; constructing a two-stage unsupervised dehazing network based on pseudo-haze images; using the training data set to train the dehazing network until a pre-set loss function converges; inputting the image to be dehazed into the trained dehazing network to obtain a dehazed image. By learning and training with unpaired real haze-clear image pairs, the present invention not only improves the dehazing effect of the network on haze images, but also improves the generalization ability of the network; the present invention introduces the decomposition and reconstruction of the traditional dark channel prior and the atmospheric scattering model into the deep learning framework, which not only improves the dehazing performance of the network, but also improves the interpretability of the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer image processing, and particularly relates to a two-stage unsupervised dehazing method based on pseudo-haze images. Background Art

[0002] Existing image dehazing methods can be roughly divided into two categories, namely, dehazing methods based on image priors and dehazing methods based on deep learning. Traditional dehazing methods based on image priors propose various priors as additional constraints to recover haze-free images based on image statistical analysis and empirical observations. However, this method is not applicable to all dehazing scenarios, and their performance is limited by the accuracy of the handcrafted priors adopted in various real-world scenarios.

[0003] The development of deep learning and large-scale synthetic datasets has enabled fully supervised dehazing methods to get rid of the limitations of traditional priors and gain a dominant position. Although network architectures including convolutional neural networks (CNNs), encoder-decoder networks, and transformers have achieved considerable success, the fact that haze images lack paired clear images has led to the existing dehazing methods usually using paired synthetic datasets for learning and training. Due to the significant domain differences between the synthetic domain and the real domain, these methods have poor dehazing effects on haze images.

[0004] In view of the problems existing in fully supervised dehazing methods, some semi-supervised or unsupervised methods have been proposed one after another. Semi-supervised methods are first pre-trained on paired synthetic datasets, and then the pre-trained weights are shared and re-trained on haze images. Unsupervised methods usually establish a cyclic dehazing and fogging process between clear images and haze images based on the principle of cycle consistency. However, these methods do not take into account that the haze concentration changes with depth and do not interact with haze images, resulting in the generated pseudo-haze images lacking authenticity. Summary of the Invention

[0005] In order to solve the above technical problems existing in the prior art, the present invention provides a two-stage unsupervised dehazing method based on pseudo-haze images.

[0006] The technical solution of the present invention to solve the above technical problems is: a two-stage unsupervised dehazing method based on pseudo-haze images, comprising the following steps:

[0007] Step S1, establishing a training dataset including haze images and unpaired clear images.

[0008] Step S2: Construct a two-stage unsupervised dehazing network based on pseudo-hazy images. Specifically, it includes: constructing a pseudo-hazy image generation framework and a three-branch network. The pseudo-hazy image generation framework consists of a depth estimation network and a haze transfer network. The three-branch network consists of a transmission rate estimation network, an atmospheric light estimation network, and a clear image estimation network.

[0009] Step S3: Use the training dataset to train the dehazing network until the pre-set loss function converges.

[0010] Step S4: Take the real hazy image from the real scene as the image to be dehazed, input it into the trained dehazing network, and obtain a clear dehazed image after processing.

[0011] Furthermore, the depth estimation network adopts an hourglass model, which consists of a symmetric downsampling module and an upsampling module. The downsampling module gradually reduces the spatial resolution of the feature map through convolutional layers and pooling operations, while increasing the number of feature channels to capture the high-level semantic information of the image. The upsampling module gradually restores the spatial resolution of the feature map through transposed convolution or upsampling operations to reconstruct the details of the image and predict the depth information map.

[0012] Furthermore, the haze transfer network consists of a pseudo-haze image generator and a discriminator. The generator uses four consecutive convolutional kernels of different sizes, 11×11, 9×9, 7×7, and 1×1, to extract features from the input haze image, capture the fog density information in different regions of the image, and uses the Sigmoid activation function to map the output value between 0 and 1 to obtain the image feature X. A parallel 7×7 convolution is used to extract more local detail information from the input haze image, and the extracted feature map is concatenated with feature X in the channel dimension. Three consecutive 5×5, 3×3, and 1×1 convolution operations are used to fuse the features after concatenation, and the Sigmoid activation function is used to map the fused features to the range of 0 to 1 to obtain the transmittance estimate F. Two cascaded 3×3 convolutions and the RELU activation function are used to extract features from the depth information map obtained by the depth estimation network to obtain the feature map C1, and then the same convolution operations and activation functions are used to obtain the feature map C2. The feature map C1 is multiplied by the transmittance estimate F, and the result is added to the feature image C2 to obtain the transmittance estimate that fuses the depth information of the clear image and the haze features. The 0.1% brightest pixels in the dark channel of the input clear image are selected, and the average value of the RGB channels at the corresponding positions of these pixels is calculated and used as the atmospheric light. Combining the transmittance map and the clear image, the final pseudo-haze image is obtained using the atmospheric scattering model. The discriminator consists of multiple convolutional layers, the LeakyReLU activation function, and a Sigmoid output layer. The discriminator divides the input image into small blocks of 30×30 pixels and makes independent true / false judgments on each block.

[0013] Furthermore, construct the adversarial loss function:

[0014]

[0015]

[0016] where represents the haze image, represents the pseudo-haze image from the generator, represents the discriminator, represents the generator.

[0017] Furthermore, the transmittance estimation network adopts a densely connected encoder-decoder structure, obtains the transmittance map of the haze image through the dark channel prior, and inputs the obtained transmittance map into the transmittance estimation network for correction and optimization to obtain the optimized transmittance map.

[0018] Furthermore, for any input clear image , its dark channel is expressed as:

[0019]

[0020] Among them, represents the intensity observed at the pixel in color channel c, represents a window of size w centered on pixel x; the atmospheric scattering model is obtained The dark channels on both sides:

[0021]

[0022] Where and are the dark channels of images and at pixel x respectively; according to the dark channel prior, 0, then the transmittance is calculated as:

[0023]

[0024] Among them, the brightest 0.1% pixels in are selected, and the average values of the RGB channels at the corresponding positions of these pixels are calculated, and used as the atmospheric light A, then is obtained.

[0025] Furthermore, the atmospheric light estimation network is composed of four consecutive convolutional combinations and global max pooling; the convolutional combination is composed of "7×7 max pooling + 3×3 convolution + group normalization + RELU activation + 3×3 convolution + group normalization + RELU activation"; the global max pooling performs a pooling operation on each feature channel in the spatial dimension to extract the maximum value within the channel, and the sigmoid activation function maps the output value to a three-element atmospheric light vector between 0 and 1, and each element corresponds to a color channel.

[0026] The beneficial effects of the present invention are:

[0027] The present invention performs learning and training through unpaired real haze - clear image pairs, which not only improves the haze removal effect of the network on haze images, but also improves the generalization ability of the network.

[0028] In the first stage of the present invention, a haze transfer network is used to realize the conversion of clear images into pseudo - haze images to reduce the domain difference between the synthetic domain and the real domain. This network combines the physical characteristics of haze images in the real world, that is, haze changes with depth, and introduces depth information into the haze transfer network, and finally obtains a pseudo - haze image highly similar to the haze image.

[0029] In the second stage of the present invention, a three-branch training network consisting of transmittance map estimation, atmospheric light estimation, and clear image estimation is adopted. By introducing the decomposition and reconstruction of traditional dark channel prior and atmospheric scattering model into the deep learning framework, not only the haze removal performance of the network is improved, but also the interpretability of the network is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic diagram of the network structure of a two-stage unsupervised haze removal method based on pseudo-hazy images in the present invention.

[0031] Figure 2 It is a schematic diagram of the structure of the haze transfer network in the present invention.

[0032] Figure 3 It is an example of the haze removal result of an embodiment of the present invention. (a)-(e) are five randomly selected hazy images from real hazy images collected from real scenes, and (f)-(j) are the results of testing (a)-(e) using the present invention.

[0033] Figure 4 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] As Figures 1-4 shown, a two-stage unsupervised haze removal method based on pseudo-hazy images includes the following steps:

[0036] Step S1, establish a training data set including hazy images and unpaired clear images; the training data set includes two types of images: one is real hazy images collected from real scenes, and the other is unpaired clear images.

[0037] Step S2, construct a two-stage unsupervised haze removal network based on pseudo-hazy images; specifically including:

[0038] Construct a haze removal network composed of a pseudo-hazy image generation framework in the first stage and a three-branch network in the second stage, called TUF_Net, and the structure is as Figure 1 shown.

[0039] The pseudo-hazy image generation framework is composed of a depth estimation network (abbreviated as D_Net) and a haze transfer network (abbreviated as Haze-Transfer Net), and the structure is as Figure 2 shown.

[0040] Construct the pseudo-hazy image generation network in the first stage, specifically including:

[0041] Step S21: Pre-train the hourglass model using a large-scale depth map dataset to obtain a trained depth estimation network. Input a clear image into D_Net to obtain the corresponding depth information map.

[0042] Step S22: Construct a haze transfer network, the structure of which is as Figure 2 shown, specifically including:

[0043] Step S221: Construct a pseudo-haze image generator based on the atmospheric scattering model, called Gh. In the generator, first use four consecutive convolutional kernels of different sizes (11×11, 9×9, 7×7, 1×1) to extract features from the input haze image to capture the fog density information in different regions of the image, and use the Sigmoid activation function to map the output value to between 0 and 1 to obtain the image feature X. At the same time, use a parallel 7×7 convolution to extract more local detail information from the input haze image, and splice the extracted feature map with feature X in the channel dimension to enable the network to fully consider the context information of the image. Then use three consecutive 5×5, 3×3, 1×1 convolution operations to perform feature fusion on the spliced features, and use the Sigmoid activation function to map the fused features to the range of 0 to 1 to obtain a rough transmission rate estimate F. Secondly, use 2 cascaded 3×3 convolutions and the RELU activation function to extract features from the obtained depth information map to obtain the feature map C1, and then use the same convolution operation and activation function to obtain the feature map C2. Then, multiply the feature map C1 by the rough transmission rate estimate F, and add the obtained result to the feature image C2 to obtain a transmission rate estimate that fuses the clear image depth information and the real haze features. Finally, according to the dark channel prior knowledge, select the brightest 0.1% pixels in the dark channel of the input clear image, and calculate the average value of the RGB channels at the corresponding positions of these pixels, and use it as the atmospheric light. Combine the transmission rate map and the clear image, and use the atmospheric scattering model to obtain the final pseudo-haze image.

[0044] Step S222: Construct a discriminator composed of multiple convolutional layers, LeakyReLU activation functions, and a Sigmoid output layer, called Dh. The discriminator divides the input image into small blocks of 30×30 pixels and makes independent true / false judgments on each block. The role of the discriminator is to determine whether the obtained pseudo-haze image is a real haze image or a fake image generated by the generator, so as to guide the generator to generate more realistic images. Use the LSGAN loss as the adversarial loss function:

[0045]

[0046]

[0047] Among them,​ represents the haze image, represents the pseudo-haze image obtained by the generator Gh.

[0048] Construct a three-branch network in the second stage, specifically including:

[0049] Step S23, use the pyramid dense connection network in DCPDN to construct a transmission rate estimation network, called T_Net. Specifically including:

[0050] Step S231, obtain the transmission rate map of the input haze image through the dark channel prior. The dark channel prior means that in the non-sky area of most clear images, there is at least one color channel called the dark channel, where the pixel value is very low and may even be close to 0. For any input clear image J, its dark channel is expressed as:

[0051] ;

[0052] where, represents the intensity observed at pixel y in color channel c, represents a window of size w centered on pixel x. In the present invention, w = 15 is preferably selected; then the dark channels on both sides of the atmospheric scattering model described in step (2) are obtained:

[0053] ;

[0054] where and are respectively the dark channels of the images and at pixel . According to the dark channel prior, , then the transmission rate is calculated as:

[0055]

[0056] where, the brightest 0.1% pixels in are selected, and the average value of the RGB channels at the corresponding positions of these pixels is calculated and used as the atmospheric light , then can be obtained.

[0057] Step S232, input the transmission rate obtained in step S231 into T_Net for correction and optimization.

[0058] Step S24: Construct an atmospheric light estimation network, denoted as A_Net. A_Net is specifically composed of four consecutive convolutional combinations and global max pooling; the convolutional combination is composed of "7×7 max pooling + 3×3 convolution + group normalization + RELU activation + 3×3 convolution + group normalization + RELU activation"; global max pooling performs pooling operations on each feature channel in the spatial dimension to extract the maximum value within the channel, and the sigmoid activation function maps the output value to a three-element atmospheric light vector between 0 and 1, with each element corresponding to a color channel.

[0059] Step S25: Construct a clear image estimation network, simply denoted as J_Net, using a multi-scale enhanced haze removal network based on dense feature fusion, namely MSBDN. Input the pseudo-hazy image generated in the first stage and the real hazy images in the training dataset into J_Net to obtain the corresponding dehazed images.

[0060] Step S251: Construct a consistency loss function: ;

[0061] where, represents the pseudo-hazy image generated in the first stage, represents the clear image corresponding to , and represents the L1 norm.

[0062] Step S252: Construct a discriminator composed of multiple convolutional layers, LeakyReLU activation functions, and a Sigmoid output layer, denoted as Dc. The discriminator divides the input image into small blocks of 30×30 pixels and independently determines the authenticity of each block. Using the judgment results of the discriminator, the dehazed image of the hazy image and the clear image have the same distribution, and the LSGAN loss is used as the adversarial loss function:

[0063] ;

[0064] ;

[0065] where, represents the clear image, represents the hazy image.

[0066] Step S26: Input the hazy image into a three-branch network to respectively output the transmission rate map, atmospheric light, and dehazed image.

[0067] Reconstruct the hazy image using the atmospheric scattering model; the atmospheric scattering model is:

[0068]

[0069] Among them, x represents the position of the pixel. I(x) represents the collected hazy image, J(x) represents the clear image, A represents the global atmospheric light, and T(x) is the transmittance, indicating the attenuation degree of light when propagating in the atmosphere.

[0070] Step S27, construct a reconstruction loss function: ;

[0071] Where represents the input hazy image, represents the reconstructed hazy image.

[0072] Step S3, use the training dataset to train the dehazing network until the pre-set loss function converges, and obtain the trained dehazing network.

[0073] Use the training dataset to train TUF_Net, and the specific training method is as follows:

[0074] Step S31, first use the depth information map obtained in step S21 and the training dataset to train the hazy transfer network in the first stage until the adversarial loss function set in step S222 converges.

[0075] Step S32, fix the pre-trained hazy transfer network, and use the trained hazy transfer network to obtain the pseudo-hazy image matching the clear image in the training dataset. At the same time, train the three-branch network in the second stage with the pseudo-hazy image and the hazy image in the training dataset until the consistency loss set in step S251, the adversarial loss function set in step S252, and the reconstruction loss function set in step S27 converge.

[0076] Step S4, use the real hazy image from the real scene as the image to be dehazed, input it into the trained dehazing network, and obtain the clear dehazed image after processing.

[0077] The experiments were conducted on the RESIDE public dataset using the present invention. RESIDE contains subsets of ITS / OTS, SOTS - indoor / outdoor, HSTS, RTTS, and URHI. ITS / OTS contains 13990 / 313950 synthetic indoor / outdoor hazy images with ground - truth clear images for training. SOTS - indoor / outdoor includes 500 synthetic indoor / outdoor hazy images and ground - truth clear images for testing. HSTS consists of 10 synthetic hazy images and 10 hazy images obtained from the real world. The training samples were randomly cropped to 256 × 256 and data augmentation was performed by random rotation and horizontal flipping. The random rotation angles were set to 90 degrees, 180 degrees, and 270 degrees. We obtained 3577 real - world outdoor clear images and 2903 higher - quality hazy images from RESIDE to form an unpaired training dataset for training, and used the hazy images as the test set.

[0078] Figure 3 The defogging results of the hazy images are shown. Among them, (a) - (e) are five randomly selected hazy images collected from real - world scenes, and (f) - (j) are the results of testing (a) - (e) using the present invention. It can be seen that clear defogged images are obtained.

Claims

1. A two-stage unsupervised dehazing method based on pseudo-haze images, characterized in that, The steps include the following: Step S1: Establish a training dataset including haze images and unpaired clear images; Step S2: Construct a two-stage unsupervised dehazing network based on pseudo-haze images; specifically including: constructing a pseudo-haze image generation framework and a three-branch network; the pseudo-haze image generation framework consists of a depth estimation network and a haze transfer network; the three-branch network consists of a transmission rate estimation network, an atmospheric light estimation network, and a clear image estimation network; the haze transfer network consists of a pseudo-haze image generator and a discriminator; the pseudo-haze image generator uses four consecutive convolutional kernels of different sizes, 11×11, 9×9, 7×7, and 1×1, to extract features from the input haze image to capture the fog density information in different regions of the image, and uses the Sigmoid activation function to map the output value to between 0 and 1 to obtain the image feature X; a parallel 7×7 convolution is used to extract local detail information from the input haze image, and the extracted feature map is concatenated with the feature X in the channel dimension; three consecutive 5×5, 3×3, and 1×1 convolution operations are used to perform feature fusion on the concatenated features, and the fused features are mapped to the range of 0 to 1 using the Sigmoid activation function to obtain the transmission rate estimation F; two cascaded 3×3 convolutions and the RELU activation function are used to extract features from the depth information map obtained by the depth estimation network to obtain the feature map C1, and then the same convolution operations and activation functions are used to obtain the feature map C2; the feature map C1 is multiplied by the transmission rate estimation F, and the result obtained is added to the feature image C2 to obtain the transmission rate estimation that fuses the clear image depth information and the haze features; the brightest 0.1% pixels in the dark channel of the input clear image are selected, and the average value of the RGB channels at the corresponding positions of these pixels is calculated and used as the atmospheric light; combining the transmission rate map and the clear image, the final pseudo-haze image is obtained using the atmospheric scattering model; the discriminator consists of multiple convolutional layers, the LeakyReLU activation function, and a Sigmoid output layer; the discriminator divides the input image into small blocks of 30×30 pixels and makes independent true / false judgments on each small block; Step S3: Train the dehazing network using the training dataset until the pre-set loss function converges to obtain a trained dehazing network; Step S4: Input the image to be dehazed into the trained dehazing network to obtain the dehazed image.

2. The two-stage unsupervised dehazing method based on pseudo-haze images according to claim 1, wherein, The depth estimation network consists of a symmetric downsampling module and an upsampling module; the downsampling module gradually reduces the spatial resolution of the feature map through convolutional layers and pooling operations while increasing the number of feature channels to capture the high-level semantic information of the image; the upsampling module gradually restores the spatial resolution of the feature map through deconvolution or upsampling operations to reconstruct the details of the image and predict the depth information map.

3. The two-stage unsupervised dehazing method based on pseudo-haze images according to claim 1, characterized in that, The adversarial loss function used by the discriminator is as follows: ]; ; Among them, represents the haze image, represents the pseudo-haze image from the generator, represents the discriminator, represents the generator.

4. The two-stage unsupervised dehazing method based on pseudo-haze images according to claim 1, wherein, The described transmittance estimation network adopts a densely connected encoder-decoder structure, obtains the transmittance map of the hazy image through the dark channel prior, inputs the obtained transmittance map into the transmittance estimation network for correction and optimization, and obtains the optimized transmittance map.

5. The two-stage unsupervised dehazing method based on pseudo-haze images according to claim 1, characterized in that For any input clear image J, its dark channel is expressed as: ; Among them, represents the intensity observed at the pixel in the color channel, represents a window of size w centered on pixel x; the atmospheric scattering model is: ; The dark channels on both sides: ; where and are the dark channels of images I and J at pixel x, respectively; according to the dark channel prior, 0, then the transmittance is calculated as: ; Among them, select the brightest 0.1% of the pixels, and calculate the average value of the RGB channels at the corresponding positions of these pixels, which is used as the atmospheric light A, then obtain .

6. The two-stage unsupervised dehazing method based on pseudo-haze images according to claim 1, wherein, The described atmospheric light estimation network consists of four consecutive convolutional combinations and global max pooling; the convolutional combination consists of "7×7 max pooling + 3×3 convolution + group normalization + RELU activation + 3×3 convolution + group normalization + RELU activation"; the global max pooling performs pooling operations on each feature channel in the spatial dimension to extract the maximum value within the channel, and the sigmoid activation function maps the output value to a three-element atmospheric light vector between 0 and 1, and each element corresponds to a color channel.

Citation Information

Patent Citations

  • Rapid image defogging method based on depth estimation prior and electronic equipment

    CN114240766A