Two-stage unsupervised defogging method based on pseudo haze image

By constructing a pseudo-haze image generation framework that includes depth estimation network and haze transfer network, as well as a three-branch network composed of transmittance estimation, atmospheric light estimation and clear image estimation, the problem of the poor effect of existing fog removal methods on haze images is solved, and more realistic pseudo-haze image generation and more efficient fog removal effects are achieved.

CN120047360AActive Publication Date: 2025-05-27HUNAN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510526550.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing defog removal methods are not effective on haze images, especially due to the domain differences between the synthetic domain and the real domain, resulting in the lack of authenticity of the generated pseudo-haze images.

Method used

A two-stage unsupervised fog removal method based on pseudo-haze images is adopted to reduce the domain difference between the synthetic domain and the real domain and introduce depth information and physical characteristics by constructing a pseudo-haze image generation framework that includes depth estimation network and haze transfer network, as well as a three-branch network composed of transmittance estimation, atmospheric light estimation and clear image estimation.

Benefits of technology

The network's defog removal effect and generalization ability of haze images is improved, and the generated pseudo-haze images are more realistic, and the defog removal performance and interpretability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047360A_ABST
    Figure CN120047360A_ABST
Patent Text Reader

Abstract

The invention discloses a two-stage unsupervised defogging method based on a pseudo haze image, and belongs to the technical field of computer image processing, and the method comprises the following steps: building a training data set containing a haze image and an unpaired clear image; constructing a two-stage unsupervised defogging network based on the pseudo haze image; training the defogging network by adopting the training data set until a preset loss function is converged; and inputting a to-be-defogged image into the trained defogging network to obtain a defogged image. According to the invention, learning training is carried out on the unpaired real haze-clear image pair, so that the defogging effect of the network on the haze image is improved, and the generalization ability of the network is improved; according to the method, traditional dark channel prior and decomposition and reconstruction of an atmospheric scattering model are introduced into a deep learning framework, so that the network defogging performance is improved, and the network interpretability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of computer image processing, and in particular relates to a two-stage unsupervised defogging method based on pseudo haze images. Background Art

[0002] Existing image dehazing methods can be roughly divided into two categories, namely, image prior-based dehazing methods and deep learning-based dehazing methods. Traditional image prior-based dehazing methods propose a variety of priors as additional constraints to restore haze-free images based on image statistical analysis and empirical observations, but this method is not suitable for all dehazing scenarios, and their performance is limited by the accuracy of the hand-crafted priors adopted in various real-world scenarios.

[0003] The development of deep learning and large-scale synthetic datasets has enabled fully supervised dehazing methods to break free from the limitations of traditional priors and gain a dominant position. Although network structures including convolutional neural networks (CNNs), encoder-decoder networks (Encoder-Decoder), and Transformer have achieved considerable success, the fact that haze images lack paired clear images has led to existing dehazing methods usually using paired synthetic datasets for learning and training. Due to the significant domain difference between the synthetic domain and the real domain, these methods have poor dehazing effects on haze images.

[0004] In response to the problems of fully supervised dehazing methods, some semi-supervised or unsupervised methods have been proposed one after another. Semi-supervised methods are first pre-trained on paired synthetic datasets, and then the pre-trained weights are shared and re-trained on haze images. Unsupervised methods are usually based on the principle of cycle consistency, establishing a cyclic dehazing and hazing process between clear images and haze images. However, these methods do not take into account that the haze concentration will change with depth, nor do they interact with haze images, resulting in the lack of authenticity of the generated pseudo haze images. Summary of the invention

[0005] In order to solve the above technical problems existing in the prior art, the present invention provides a two-stage unsupervised dehazing method based on pseudo haze images.

[0006] The technical solution of the present invention to solve the above technical problem is: a two-stage unsupervised defogging method based on pseudo haze images, comprising the following steps:

[0007] Step S1, establishing a training data set including haze images and unpaired clear images.

[0008] Step S2, constructing a two-stage unsupervised dehazing network based on pseudo haze images; specifically, constructing a pseudo haze image generation framework and a three-branch network; the pseudo haze image generation framework is composed of a depth estimation network and a haze transfer network; the three-branch network is composed of a transmittance estimation network, an atmospheric light estimation network and a clear image estimation network;

[0009] Step S3, training the dehazing network using the training data set until the preset loss function converges;

[0010] Step S4, taking the real haze image from the real scene as the image to be dehazed, inputting it into the trained dehazing network, and obtaining a clear dehazed image after processing.

[0011] Furthermore, the depth estimation network adopts an hourglass model, which is composed of a symmetrical downsampling module and an upsampling module; the downsampling module gradually reduces the spatial resolution of the feature map through convolution layers and pooling operations, while increasing the number of feature channels to capture high-level semantic information of the image; the upsampling module gradually restores the spatial resolution of the feature map through deconvolution or upsampling operations to reconstruct the details of the image and predict the depth information map;

[0012] Furthermore, the haze transfer network is composed of a pseudo haze image generator and a discriminator; the generator uses four consecutive convolution kernels of different sizes, 11×11, 9×9, 7×7, and 1×1, to extract features from the input haze image to capture the fog density information in different areas of the image, and uses the Sigmoid activation function to map the output value to between 0 and 1 to obtain the image feature X; a parallel 7×7 convolution is used to extract more local detail information from the input haze image, and the extracted feature map is spliced ​​with the feature X in the channel dimension; three consecutive 5×5, 3×3, and 1×1 convolution operations are used to fuse the spliced ​​features, and the Sigmoid activation function is used to map the fused features to the range of 0 to 1 to obtain Transmittance estimation F; use two cascaded 3×3 convolutions and RELU activation functions to extract features from the depth information map obtained by the depth estimation network to obtain feature map C1, and then use the same convolution operation and activation function to obtain feature map C2; multiply feature map C1 with the transmittance estimate F, and add the result to feature image C2 to obtain a transmittance estimate that integrates the depth information of the clear image and the haze features; select the brightest 0.1% pixels in the dark channel of the input clear image, and calculate the average value of the RGB channels at the corresponding positions of these pixels, which is used as atmospheric light; combine the transmittance map and the clear image, and use the atmospheric scattering model to obtain the final pseudo haze image; the discriminator consists of multiple convolutional layers, LeakyReLU activation function and Sigmoid output layer; the discriminator divides the input image into multiple 30×30 pixel blocks, and makes independent true or false judgments on each block.

[0013] Furthermore, we construct an adversarial loss function:

[0014]

[0015]

[0016] in, represents a haze image, represents the pseudo haze image from the generator, represents the discriminator, Represents a generator.

[0017] Furthermore, the transmittance estimation network adopts a densely connected compiler-decoder structure, obtains the transmittance map of the haze image through a dark channel prior, and inputs the obtained transmittance map into the transmittance estimation network for correction and optimization to obtain an optimized transmittance map.

[0018] Furthermore, for any input clear image , and its dark channel is expressed as:

[0019]

[0020] in, represents the intensity observed at the pixel in color channel c, Represents a window of size w centered on pixel x; get the atmospheric scattering model Dark channels on both sides:

[0021]

[0022] in and The images are and The dark channel at pixel x; according to the dark channel prior, 0, then the transmittance Calculated as:

[0023]

[0024] Among them, select The brightest 0.1% pixels in the image are selected, and the average values ​​of the RGB channels at the corresponding positions of these pixels are calculated, which is regarded as the atmospheric light A. Then, .

[0025] Furthermore, the atmospheric light estimation network is composed of four consecutive convolution combinations and global maximum pooling; the convolution combination is composed of "7×7 maximum pooling + 3×3 convolution + group normalization + RELU activation + 3×3 convolution + group normalization + RELU activation"; the global maximum pooling performs a pooling operation on each feature channel in the spatial dimension, extracts the maximum value in the channel, and the sigmoid activation function maps the output value to a three-element atmospheric light vector between 0 and 1, where each element corresponds to a color channel.

[0026] The beneficial effects of the present invention are:

[0027] The present invention uses unpaired real haze-clear image pairs for learning and training, which not only improves the defogging effect of the network on haze images, but also improves the generalization ability of the network.

[0028] In the first stage of the present invention, a haze transfer network is used to realize the conversion of a clear image into a pseudo haze image to reduce the domain difference between the synthetic domain and the real domain. The network combines the physical properties of haze images in the real world, that is, haze changes with depth, introduces depth information into the haze transfer network, and finally obtains a pseudo haze image that is highly realistic with the haze image.

[0029] The second stage of the present invention adopts a three-branch training network composed of transmittance map estimation, atmospheric light estimation and clear image estimation, and introduces the decomposition and reconstruction of the traditional dark channel prior and atmospheric scattering model into the deep learning framework, which not only improves the network dehazing performance, but also improves the network interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Schematic diagram of the network structure of a two-stage unsupervised dehazing method based on pseudo haze images in the present invention.

[0031] Figure 2 It is a schematic diagram of the structure of the haze transfer network in the present invention.

[0032] Figure 3 2 is an example of the defogging result of an embodiment of the present invention, (a) to (e) are five haze images randomly selected from real haze images collected from real scenes, and (f) to (j) are the results of testing (a) to (e) using the present invention.

[0033] Figure 4 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0034] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0035] like Figure 1-4 As shown, a two-stage unsupervised dehazing method based on pseudo haze images includes the following steps:

[0036] Step S1, establishing a training data set including haze images and unpaired clear images; the training data set includes two types of images: one is real haze images collected from real scenes, and the other is unpaired clear images.

[0037] Step S2, constructing a two-stage unsupervised dehazing network based on the pseudo haze image; specifically comprising:

[0038] Construct a dehazing network consisting of the pseudo haze image generation framework in the first stage and the three-branch network in the second stage, called TUF_Net, with the following structure: Figure 1 shown.

[0039] The pseudo haze image generation framework consists of a depth estimation network (abbreviated as D_Net) and a haze transfer network (abbreviated as Haze-Transfer Net). The structure is as follows: Figure 2 shown.

[0040] Construct the first stage of pseudo haze image generation network, including:

[0041] Step S21, use a large-scale depth map dataset to pre-train the hourglass model to obtain a trained depth estimation network, and input the clear image into D_Net to obtain the corresponding depth information map.

[0042] Step S22, constructing a haze transfer network, the structure of which is as follows Figure 2 As shown, specifically including:

[0043] Step S221, construct a pseudo haze image generator based on the atmospheric scattering model, called Gh. In the generator, firstly, four consecutive convolution kernels of different sizes (11×11, 9×9, 7×7, 1×1) are used to extract features from the input haze image to capture the fog density information of different areas of the image, and the Sigmoid activation function is used to map the output value to between 0 and 1 to obtain the image feature X. At the same time, a parallel 7×7 convolution is used to extract more local detail information from the input haze image, and the extracted feature map is spliced ​​with the feature X in the channel dimension, so that the network fully considers the context information of the image. Then, three consecutive 5×5, 3×3, and 1×1 convolution operations are used to fuse the spliced ​​features, and the Sigmoid activation function is used to map the fused features to the range of 0 to 1 to obtain a rough transmittance estimate F. Secondly, two cascaded 3×3 convolutions and RELU activation functions are used to extract features from the depth information map to obtain feature map C1, and the same convolution operation and activation function are used to obtain feature map C2. Next, feature map C1 is multiplied by a rough transmittance estimate F, and the result is added to feature image C2 to obtain a transmittance estimate that combines the depth information of the clear image and the real haze features. Finally, based on the dark channel prior knowledge, the brightest 0.1% pixels in the dark channel of the input clear image are selected, and the average value of the RGB channels at the corresponding positions of these pixels is calculated as the atmospheric light. Combining the transmittance map and the clear image, the atmospheric scattering model is used to obtain the final pseudo haze image.

[0044] Step S222, construct a discriminator consisting of multiple convolutional layers, LeakyReLU activation function and Sigmoid output layer, called Dh. The discriminator divides the input image into multiple small blocks of 30×30 pixels and makes independent true and false judgments for each small block. The role of the discriminator is to determine whether the obtained pseudo haze image is a real haze image or a fake image generated by the generator, so as to guide the generator to produce more realistic images. LSGAN loss is used as the adversarial loss function:

[0045] ]

[0046]

[0047] in, represents a haze image, Represents the pseudo haze image obtained by the generator Gh.

[0048] Construct the three-branch network of the second phase, including:

[0049] Step S23, using the pyramid densely connected network in DCPDN to build a transmittance estimation network, called T_Net. Specifically including:

[0050] Step S231, obtain the transmittance map of the input haze image through the dark channel prior. The dark channel prior means that in the non-sky area of ​​most clear images, there is at least one color channel called the dark channel, in which the pixel value is very low, and may even be close to 0. For any input clear image J, its dark channel is expressed as:

[0051] ;

[0052] in, represents the intensity observed at pixel y in color channel c, represents a window of size w centered on pixel x, and w=15 is preferred in the present invention; then the dark channels on both sides of the atmospheric scattering model in step (2) are obtained:

[0053] ;

[0054] in and The images are and In pixels According to the dark channel prior, , then the transmittance Calculated as:

[0055]

[0056] Among them, select The brightest 0.1% pixels are selected and the average value of the RGB channels at the corresponding positions of these pixels is calculated as the atmospheric light. , we can get .

[0057] Step S232: The transmittance obtained in step S231 is Input to T_Net for correction and optimization.

[0058] Step S24, construct an atmospheric light estimation network, called A_Net. A_Net is specifically composed of four consecutive convolution combinations and global maximum pooling; the convolution combination is composed of "7×7 maximum pooling + 3×3 convolution + group normalization + RELU activation + 3×3 convolution + group normalization + RELU activation"; the global maximum pooling performs a pooling operation on each feature channel in the spatial dimension, extracts the maximum value in the channel, and the sigmoid activation function maps the output value to a three-element atmospheric light vector between 0 and 1, and each element corresponds to a color channel.

[0059] Step S25, a multi-scale enhanced dehazing network based on dense feature fusion, MSBDN, is used to construct a clear image estimation network, referred to as J_Net. The pseudo haze image generated in the first stage and the real haze image in the training data set are input into J_Net to obtain the corresponding dehazed image.

[0060] Step S251, construct a consistency loss function: ;

[0061] in, represents the pseudo haze image generated in the first stage, Representation and The corresponding clear image, represents the L1 norm.

[0062] Step S252, construct a discriminator consisting of multiple convolutional layers, LeakyReLU activation function and Sigmoid output layer, called Dc. The discriminator divides the input image into multiple small blocks of 30×30 pixels and makes independent true and false judgments for each small block. Using the judgment results of the discriminator, the dehazed image of the haze image has the same distribution as the clear image, and LSGAN loss is used as the adversarial loss function:

[0063] ;

[0064] ;

[0065] in, Indicates a clear image. Represents a haze image.

[0066] Step S26, input the haze image into a three-branch network, and output a transmittance map, an atmospheric light map, and a dehazed image respectively.

[0067] The haze image is reconstructed using the atmospheric scattering model; the atmospheric scattering model is:

[0068]

[0069] Among them, x represents the position of the pixel. I(x) represents the collected haze image, J(x) represents the clear image, A represents the global atmospheric light, and T(x) is the transmittance, which represents the attenuation of light when it propagates in the atmosphere.

[0070] Step S27, construct the reconstruction loss function: ;

[0071] in represents the input haze image, Represents the reconstructed haze image.

[0072] Step S3, using the training data set to train the defogging network until the preset loss function converges to obtain a trained defogging network.

[0073] The training data set is used to train TUF_Net. The specific training method is as follows:

[0074] In step S31, the first stage haze transfer network is trained using the depth information map and training data set obtained in step S21 until the adversarial loss function set in step S222 converges.

[0075] Step S32, fix the pre-trained haze transfer network, and use the trained haze transfer network to obtain a pseudo haze image that matches the clear image in the training data set. At the same time, the pseudo haze image and the haze image in the training data set are used to train the three-branch network of the second stage until the consistency loss set in step S251, the adversarial loss function set in step S252, and the reconstruction loss function set in step S27 converge.

[0076] Step S4, taking the real haze image from the real scene as the image to be dehazed, inputting it into the trained dehazing network, and obtaining a clear dehazed image after processing.

[0077] The present invention is used to conduct experiments on the RESIDE public dataset. RESIDE includes ITS / OTS, SOTS-indoor / outdoor, HSTS, RTTS and URHI subsets. ITS / OTS contains 13990 / 313950 synthetic indoor / outdoor haze images with real clear images (ground-truth) for training. SOTS-indoor / outdoor includes 500 synthetic indoor / outdoor haze images and real clear images (ground-truth) for testing. HSTS consists of 10 synthetic haze images and 10 haze images obtained from the real world. The training samples are randomly cropped to 256 × 256, and data enhancement is performed by random rotation and horizontal flipping, and the random rotation angles are set to 90 degrees, 180 degrees and 270 degrees. We obtain 3577 real outdoor clear images and 2903 higher quality haze images from RESIDE to form an unpaired training dataset for training, and use haze images as the test set.

[0078] Figure 3 The defogging results of haze images are shown. (a) to (e) are five haze images randomly selected from haze images collected from real scenes, and (f) to (j) are the results of testing (a) to (e) using the present invention. It can be seen that clear defogging images are obtained.

Claims

1. A two-stage unsupervised dehazing method based on pseudo haze images, characterized in that: The following steps are involved: Step S1, establishing a training data set including haze images and unpaired clear images; Step S2, constructing a two-stage unsupervised dehazing network based on pseudo haze images; specifically, constructing a pseudo haze image generation framework and a three-branch network; the pseudo haze image generation framework is composed of a depth estimation network and a haze transfer network; the three-branch network is composed of a transmittance estimation network, an atmospheric light estimation network and a clear image estimation network; Step S3, training the defogging network using the training data set until the preset loss function converges to obtain a trained defogging network; Step S4: input the image to be defogged into the trained defogging network to obtain a defogged image.

2. The two-stage unsupervised dehazing method based on pseudo haze images according to claim 1 is characterized in that: The depth estimation network consists of a symmetrical downsampling module and an upsampling module; the downsampling module gradually reduces the spatial resolution of the feature map through convolution layers and pooling operations, while increasing the number of feature channels to capture high-level semantic information of the image; the upsampling module gradually restores the spatial resolution of the feature map through deconvolution or upsampling operations to reconstruct the details of the image and predict the depth information map.

3. The two-stage unsupervised dehazing method based on pseudo haze images according to claim 1 is characterized in that: The haze transfer network is composed of a pseudo haze image generator and a discriminator; the pseudo haze image generator uses four consecutive convolution kernels of different sizes, 11×11, 9×9, 7×7, and 1×1, to extract features from the input haze image to capture the fog density information of different areas of the image, and uses the Sigmoid activation function to map the output value to between 0 and 1 to obtain the image feature X; a parallel 7×7 convolution is used to extract local detail information from the input haze image, and the extracted feature map is spliced ​​with the feature X in the channel dimension; three consecutive 5×5, 3×3, and 1×1 convolution operations are used to fuse the spliced ​​features, and the Sigmoid activation function is used to map the fused features to the range of 0 to 1 to obtain the transmittance estimation F; two cascaded 3×3 convolutions and RELU activation functions are used to extract features from the depth information map obtained by the depth estimation network to obtain the feature map C1, and then the same convolution operation and activation function are used to obtain the feature map C2; The feature map C1 is multiplied by the transmittance estimate F, and the result is added to the feature image C2 to obtain the transmittance estimate that integrates the depth information of the clear image and the haze feature; the brightest 0.1% pixels in the dark channel of the input clear image are selected, and the average value of the RGB channels at the corresponding positions of these pixels is calculated and used as the atmospheric light; the transmittance map and the clear image are combined, and the atmospheric scattering model is used to obtain the final pseudo haze image; the discriminator consists of multiple convolutional layers, LeakyReLU activation function and Sigmoid output layer; the discriminator divides the input image into multiple small blocks of 30×30 pixels, and makes independent true and false judgments on each small block.

4. The two-stage unsupervised dehazing method based on pseudo haze images according to claim 3 is characterized in that: The adversarial loss function used by the discriminator is as follows: ]; ; in, represents a haze image, represents the pseudo haze image from the generator, represents the discriminator, Represents a generator.

5. The two-stage unsupervised dehazing method based on pseudo haze images according to claim 1, characterized in that: The transmittance estimation network adopts a densely connected compiler-decoder structure, obtains the transmittance map of the haze image through a dark channel prior, and inputs the obtained transmittance map into the transmittance estimation network for correction and optimization to obtain an optimized transmittance map.

6. The two-stage unsupervised dehazing method based on pseudo haze images according to claim 5, characterized in that: For any input clear image J, its dark channel is expressed as: ; in, represents the intensity observed at the pixel in the color channel, represents a window of size w centered on pixel x; the atmospheric scattering model is: ; Dark channels on both sides: ; in and are the dark channels of images I and J at pixel x respectively; according to the dark channel prior, 0, then the transmittance Calculated as: ; Among them, select The brightest 0.1% pixels in the image are selected, and the average values ​​of the RGB channels at the corresponding positions of these pixels are calculated, which is regarded as the atmospheric light A. Then, .

7. The two-stage unsupervised dehazing method based on pseudo haze images according to claim 1, characterized in that: The atmospheric light estimation network is composed of four consecutive convolution combinations and global maximum pooling; the convolution combination is composed of "7×7 maximum pooling + 3×3 convolution + group normalization + RELU activation + 3×3 convolution + group normalization + RELU activation"; the global maximum pooling performs a pooling operation on each feature channel in the spatial dimension to extract the maximum value in the channel, and the sigmoid activation function maps the output value to a three-element atmospheric light vector between 0 and 1, where each element corresponds to a color channel.

Citation Information

Patent Citations

  • Image defogging method, electronic device, storage medium and computer program product

    CN114004760A

  • Rapid image defogging method based on depth estimation prior and electronic equipment

    CN114240766A

  • Non-supervision image defogging method and system based on symbiotic double models

    CN114359107A

  • Image defogging method based on deep learning and traditional priori knowledge fusion

    CN115660998A

  • Image defogging method and device based on deep learning staged training

    CN116452470A

Cited By

  • Bidirectional decoupling translation unsupervised defogging network with feature and pixel comparative representation

    CN120725920A

  • Remote sensing image haze removal detection method and device and electronic equipment

    CN120931629A

  • Real haze image defogging method based on haze degradation model

    CN121392148A

  • A real fog and haze image defogging method based on a haze degradation model

    CN121392148B

  • Defogging adaptive method for synthesizing remote sensing image based on pseudo fog

    CN121937323A