Zero-shot learning dehazing image enhancement method and device based on image decomposition

Through the zero-shot learning method of image decomposition, the initial network reconstruction and optimization layers are used to generate dehazed images, which solves the problems of long training time and difficulty in sample acquisition in the existing technology and achieves efficient image dehazing effect.

CN120298269BActive Publication Date: 2025-09-23QUANZHOU INST OF EQUIP MFG +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510793853.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-23
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Existing image dehazing methods have the problems of high training time cost and difficulty in obtaining training samples.

Method used

A zero-shot learning method based on image decomposition is adopted to reconstruct the haze image through the joint training of the initial clean layer generation network, the initial atmospheric light layer generation network and the initial transmittance layer generation network. The atmospheric scattering model is used to calculate the loss and the network is optimized to generate the target layer, finally solving the dehazed image.

Benefits of technology

Without the need for a large number of training samples, the training time is significantly shortened, the amount of calculation is reduced, and the efficiency and effect of image dehazing are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298269B_ABST
    Figure CN120298269B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of computer vision image enhancement, and provides a zero-shot learning defogging image enhancement method and device based on image decomposition. Three initial layer generation networks, namely an initial clean layer generation network, an initial atmospheric light layer generation network, and an initial transmittance layer generation network, are introduced. A target atmospheric light layer generation network and a target transmittance layer generation network are obtained through joint training. These networks are used to determine the target atmospheric light layer and target transmittance layer of the image to be defogged, respectively. Furthermore, the defogged image corresponding to the image to be defogged is obtained by combining the network with an atmospheric scattering model. During the joint training of the three initial layer generation networks, only the image to be defogged is required, without the need for a large number of training samples. This not only avoids the difficulty in obtaining training samples for image defogging models in the prior art, but also reduces the amount of computation required during training, shortens the network training time, and reduces the network training time cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision image enhancement, and in particular to a zero-sample learning defogging image enhancement method and device based on image decomposition. Background Art

[0002] In recent years, declining air quality and reduced visibility have significantly impacted people's lives and industrial development. Image quality of outdoor scenes often deteriorates under adverse weather conditions, such as haze, fog, and smoke. This light mixes with ambient light reflected from other directions and enters the field of view through atmospheric particles. As a result, objects captured in adverse weather conditions exhibit low contrast, dull colors, and shifted brightness. Image dehazing can significantly increase object contrast.

[0003] However, existing image dehazing methods all have problems such as high training time cost of image dehazing models and difficulty in obtaining training samples. Summary of the Invention

[0004] The present invention provides a zero-sample learning defogging image enhancement method and device based on image decomposition, which are used to solve the defects existing in the prior art.

[0005] The present invention provides a zero-shot learning defogging image enhancement method based on image decomposition, comprising:

[0006] Acquire an image to be defogged, and input the image to be defogged into an initial clean layer generation network, an initial atmospheric light layer generation network, and an initial transmittance layer generation network, respectively, to obtain an initial clean layer output by the initial clean layer generation network, an initial atmospheric light layer output by the initial atmospheric light layer generation network, and an initial transmittance layer output by the initial transmittance layer generation network;

[0007] Based on the initial clean layer, the initial atmospheric light layer, and the initial transmittance layer, an atmospheric scattering model is applied to reconstruct a haze image, and a reconstruction loss is calculated based on the image to be dehazed and the reconstructed haze image;

[0008] Determining a dark channel value of each pixel in the image to be defogged, determining a real atmospheric light layer corresponding to the image to be defogged based on the dark channel value, and calculating atmospheric light loss based on the real atmospheric light layer and the initial atmospheric light layer;

[0009] Based on the reconstruction loss and the atmospheric light loss, a total loss is calculated, and based on the total loss, the initial clean layer generation network, the initial atmospheric light layer generation network, and the initial transmittance layer generation network are jointly optimized to obtain a target clean layer generation network, a target atmospheric light layer generation network, and a target transmittance layer generation network;

[0010] The image to be defogged is input into the target atmospheric light layer generation network and the target transmittance layer generation network to obtain a target atmospheric light layer and a target transmittance layer respectively. Based on the target atmospheric light layer and the target transmittance layer, an atmospheric scattering model is applied to solve the defogged image corresponding to the image to be defogged.

[0011] According to a zero-shot learning defogging image enhancement method based on image decomposition provided by the present invention, the initial atmospheric light layer generation network is used to encode the image to be defogged to obtain latent variables, and determine the normal distribution satisfied by the latent variables, reconstruct the latent variables based on the normal distribution to obtain reconstructed variables, and decode the reconstructed variables to obtain the initial atmospheric light layer;

[0012] The method further comprises:

[0013] Calculate the KL divergence loss based on the normal distribution and the standard normal distribution;

[0014] The calculating of the total loss based on the reconstruction loss and the atmospheric light loss specifically includes:

[0015] The total loss is calculated based on the reconstruction loss, the atmospheric light loss, and the KL divergence loss.

[0016] According to a zero-shot learning defogging image enhancement method based on image decomposition provided by the present invention, the initial atmospheric light layer generation network includes an encoder, an intermediate block and a decoder, and the encoder and the decoder both include the same number of convolution blocks;

[0017] Each convolution block in the encoder includes a convolution layer, an activation layer, and a maximum pooling layer connected in sequence, and the encoder is used to encode the image to be dehazed to obtain a latent variable;

[0018] The intermediate block is used to determine the normal distribution satisfied by the latent variable, and reconstruct the latent variable based on the normal distribution to obtain a reconstructed variable;

[0019] Each convolution block in the decoder includes an upsampling layer, a convolution layer, a batch normalization layer and an activation layer connected in sequence, and the decoder is used to decode the reconstructed variables to obtain the initial atmospheric light layer.

[0020] According to a zero-shot learning defogging image enhancement method based on image decomposition provided by the present invention, the total loss is calculated based on the reconstruction loss and the atmospheric light loss, further comprising:

[0021] Calculating a dark channel prior loss based on the L1 norm of the dark channel value;

[0022] The total loss is calculated based on the reconstruction loss, the atmospheric light loss and the dark channel prior loss.

[0023] According to a zero-shot learning defogging image enhancement method based on image decomposition provided by the present invention, the total loss is calculated based on the reconstruction loss and the atmospheric light loss, further comprising:

[0024] Regularizing the initial atmospheric light layer and the initial transmittance layer respectively to obtain a regularized result of the atmospheric light layer and a regularized result of the transmittance layer;

[0025] Calculating a regularization loss based on the regularization result of the atmospheric light layer and the regularization result of the transmittance layer;

[0026] A total loss is calculated based on the reconstruction loss, the atmospheric light loss, and the regularization loss.

[0027] According to a zero-shot learning dehazing image enhancement method based on image decomposition provided by the present invention, regularizing the initial atmospheric light layer and the initial transmittance layer respectively to obtain the atmospheric light layer regularization result and the transmittance layer regularization result includes:

[0028] Based on the Laplace norm, the initial atmospheric light layer and the initial transmittance layer are regularized respectively to obtain a regularized result of the atmospheric light layer and a regularized result of the transmittance layer.

[0029] According to a zero-shot learning dehazing image enhancement method based on image decomposition provided by the present invention, the initial clean layer generation network is a fully convolutional neural network, and the initial transmittance layer generation network is a U-net architecture.

[0030] The present invention also provides a zero-sample learning defogging image enhancement device based on image decomposition, comprising:

[0031] A layer generation module is used to obtain an image to be defogged, and input the image to be defogged into an initial clean layer generation network, an initial atmospheric light layer generation network, and an initial transmittance layer generation network, respectively, to obtain an initial clean layer output by the initial clean layer generation network, an initial atmospheric light layer output by the initial atmospheric light layer generation network, and an initial transmittance layer output by the initial transmittance layer generation network;

[0032] a loss calculation module, configured to apply an atmospheric scattering model to reconstruct a haze image based on the initial clean layer, the initial atmospheric light layer, and the initial transmittance layer, and calculate a reconstruction loss based on the image to be dehazed and the reconstructed haze image;

[0033] The loss calculation module is further configured to determine a dark channel value of each pixel in the image to be defogged, determine a real atmospheric light layer corresponding to the image to be defogged based on the dark channel value, and calculate the atmospheric light loss based on the real atmospheric light layer and the initial atmospheric light layer;

[0034] a network optimization module, configured to calculate a total loss based on the reconstruction loss and the atmospheric light loss, and jointly optimize the initial clean layer generation network, the initial atmospheric light layer generation network, and the initial transmittance layer generation network based on the total loss to obtain a target clean layer generation network, a target atmospheric light layer generation network, and a target transmittance layer generation network;

[0035] The image defogging module is used to input the image to be defogged into the target atmospheric light layer generation network and the target transmittance layer generation network to obtain the target atmospheric light layer and the target transmittance layer respectively, and based on the target atmospheric light layer and the target transmittance layer, apply the atmospheric scattering model to solve the defogged image corresponding to the image to be defogged.

[0036] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the zero-sample learning dehazing image enhancement method based on image decomposition as described above is implemented.

[0037] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described zero-sample learning defogging image enhancement methods based on image decomposition.

[0038] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described zero-shot learning dehazing image enhancement methods based on image decomposition.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] The present invention provides a zero-shot learning defogging image enhancement method and device based on image decomposition. Three initial layer generation networks are introduced: an initial clean layer generation network, an initial atmospheric light layer generation network, and an initial transmittance layer generation network. Joint training yields a target atmospheric light layer generation network and a target transmittance layer generation network, which are used to determine the target atmospheric light layer and target transmittance layer, respectively, for the image to be defogged. Furthermore, the defogged image corresponding to the image to be defogged is obtained by combining the network with an atmospheric scattering model. Joint training of the three initial layer generation networks only requires the image to be defogged, eliminating the need for a large number of training samples. This not only avoids the difficulty in obtaining training samples for image defogging models in the prior art, but also reduces the computational effort during training, significantly shortening the network training time and cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on the drawings in the following description without any creative work.

[0042] Figure 1 This is one of the flow charts of the zero-shot learning dehazing image enhancement method based on image decomposition provided by the present invention;

[0043] Figure 2 It is a schematic diagram of the process of generating a target clean layer, a target atmospheric light layer, and a target transmittance layer in the zero-shot learning defogging image enhancement method based on image decomposition provided by the present invention;

[0044] Figure 3 Schematic diagram of the structure of the initial A-Net in the zero-shot learning dehazing image enhancement method based on image decomposition provided by the present invention;

[0045] Figure 4 Schematic diagram of the structure of the initial Clear-Net in the zero-shot learning defogging image enhancement method based on image decomposition provided by the present invention;

[0046] Figure 5 Schematic diagram of the structure of the initial T-Net in the zero-shot learning dehazing image enhancement method based on image decomposition provided by the present invention;

[0047] Figure 6 Schematic diagram of the structure of the zero-shot learning defogging image enhancement device based on image decomposition provided by the present invention;

[0048] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0050] Figure 1 FIG. 1 is a flow chart of a zero-shot learning defogging image enhancement method based on image decomposition provided in an embodiment of the present invention. Figure 1 As shown, the method includes:

[0051] S1, obtaining an image to be dehazed, and inputting the image to be dehazed into an initial clean layer generation network, an initial atmospheric light layer generation network, and an initial transmittance layer generation network, respectively, to obtain an initial clean layer output by the initial clean layer generation network, an initial atmospheric light layer output by the initial atmospheric light layer generation network, and an initial transmittance layer output by the initial transmittance layer generation network;

[0052] S2, applying an atmospheric scattering model to reconstruct a haze image based on the initial clean layer, the initial atmospheric light layer, and the initial transmittance layer, and calculating a reconstruction loss based on the image to be dehazed and the reconstructed haze image;

[0053] S3, determining a dark channel value of each pixel in the image to be defogged, determining a real atmospheric light layer corresponding to the image to be defogged based on the dark channel value, and calculating atmospheric light loss based on the real atmospheric light layer and the initial atmospheric light layer;

[0054] S4, calculating a total loss based on the reconstruction loss and the atmospheric light loss, and jointly optimizing the initial clean layer generation network, the initial atmospheric light layer generation network, and the initial transmittance layer generation network based on the total loss to obtain a target clean layer generation network, a target atmospheric light layer generation network, and a target transmittance layer generation network;

[0055] S5, inputting the image to be defogged into the target atmospheric light layer generation network and the target transmittance layer generation network, obtaining the target atmospheric light layer and the target transmittance layer respectively, and applying the atmospheric scattering model based on the target atmospheric light layer and the target transmittance layer to solve the defogged image corresponding to the image to be defogged.

[0056] Specifically, the zero-shot learning dehazing and image enhancement method based on image decomposition provided in the embodiments of the present invention is implemented by a zero-shot learning dehazing and image enhancement device based on image decomposition. The device can be configured in an electronic device, which can be a computer or an image acquisition device. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., and is not specifically limited here. The image acquisition device can be a camera, and the camera can have a zero-shot learning dehazing and image enhancement function based on image decomposition.

[0057] It can be understood that the zero-shot learning dehazing image enhancement function refers to achieving image enhancement through the zero-shot learning image dehazing process to obtain a dehazed image corresponding to the image to be dehazed. This dehazed image is the result of image enhancement.

[0058] First, step S1 is performed to obtain an image to be defogged. The image to be defogged refers to a foggy image that needs to be defogged, and the image to be defogged is a color image.

[0059] The image to be dehazed can be input into the initial clean layer generation network (Clear-Net), the initial atmospheric light layer generation network (A-Net), and the initial transmittance layer generation network (T-Net), respectively. The initial clean layer is output by the initial clean layer generation network, the initial atmospheric light layer is output by the initial atmospheric light layer generation network, and the initial transmittance layer is output by the initial transmittance layer generation network. Thus, the image to be dehazed can be disentangled into the initial clean layer, the initial atmospheric light layer, and the initial transmittance layer through the initial Clear-Net, the initial A-Net, and the initial T-Net.

[0060] Thereafter, step S2 is executed to reconstruct a haze image using the initial clean layer, the initial atmospheric light layer, and the initial transmittance layer and applying the atmospheric scattering model.

[0061] The atmospheric scattering model can be expressed by the following formula:

[0062] ;

[0063] in, For foggy images, For a clean layer of a foggy image, is the transmittance layer of the foggy image, is the atmospheric light layer of the foggy image, and a is a pixel.

[0064] Using the image to be dehazed x and the haze image f, calculate the reconstruction loss. That is:

[0065]

[0066] Among them, n is the number of pixels in the image to be dehazed x and the haze image f, is the pixel value of pixel i in the haze image f, is the pixel value of pixel i in the image x to be dehazed, is the reconstruction loss.

[0067] The purpose of this reconstruction loss is to decompose the image to be dehazed into three different image layers: clean layer, atmospheric light layer, and transmittance layer by minimizing the difference between the haze image obtained by constrained reconstruction and the image to be dehazed.

[0068] Then, step S3 is executed to determine the dark channel value of each pixel in the image to be defogged. That is, ;

[0069] in, is the dark channel value of pixel i in the image to be dehazed, is a local image block centered on pixel i in the image to be dehazed, that is, a small patch around pixel i, c is the color channel, r is the red channel, g is the green channel, b is the blue channel, is the pixel value of pixel y in the color channel c of the image to be dehazed, y is Pixels in .

[0070] Using the dark channel value, we can determine the real atmospheric light layer corresponding to the image to be defogged, which is:

[0071] ;

[0072] in, is the real atmospheric light layer corresponding to the image to be dehazed, is a set of preset number of pixels with high dark channel values, is the pixel value of pixel j in the image to be dehazed.

[0073] Using Real Atmospheric Light Layers And the initial atmospheric light layer , calculate the atmospheric light loss, we have:

[0074] ;

[0075] in, is the atmospheric light loss, which belongs to the mean square error loss.

[0076] Finally, step S4 is executed to calculate the total loss using the reconstruction loss and the atmospheric light loss. Here, the total loss can be obtained by weighted summing the reconstruction loss and the atmospheric light loss. Alternatively, other losses can be introduced as needed based on the reconstruction loss and the atmospheric light loss, and the weighted sum of all losses can be calculated to obtain the total loss.

[0077] Using the total loss, the initial clean layer generation network, the initial atmospheric light layer generation network, and the initial transmittance layer generation network can be jointly optimized to obtain the target clean layer generation network, the target atmospheric light layer generation network, and the target transmittance layer generation network.

[0078] The target clean layer generation network has the performance of generating the target clean layer of the image to be dehazed, the target atmospheric light layer generation network has the performance of generating the target atmospheric light layer of the image to be dehazed, and the target transmittance layer generation network has the performance of generating the target transmittance layer of the image to be dehazed.

[0079] Finally, step S5 is executed to input the image to be defogged into the target atmospheric light layer generation network and the target transmittance layer generation network to obtain the target atmospheric light layer and the target transmittance layer respectively. The target atmospheric light layer and the target transmittance layer are used to apply the atmospheric scattering model to solve the defogged image corresponding to the image to be defogged. Figure 2 shown.

[0080] It can be understood that although the target clean layer generation network has the performance of generating the target clean layer of the image to be dehazed, in the embodiment of the present invention, the target clean layer generation network is not directly used to determine the dehazed image corresponding to the image to be dehazed. The target clean layer generation network is only an intermediate product obtained by jointly optimizing the initial clean layer generation network, the initial atmospheric light layer generation network and the initial transmittance layer generation network.

[0081] The zero-shot learning dehazing image enhancement method based on image decomposition provided in an embodiment of the present invention introduces three initial layer generation networks: an initial clean layer generation network, an initial atmospheric light layer generation network, and an initial transmittance layer generation network. Jointly training these networks yields a target atmospheric light layer generation network and a target transmittance layer generation network, which are used to determine the target atmospheric light layer and target transmittance layer, respectively, for the image to be dehazed. Furthermore, combined with an atmospheric scattering model, the dehazed image corresponding to the image to be dehazed is obtained. Joint training of the three initial layer generation networks only requires the image to be dehazed, eliminating the need for a large number of training samples. This not only avoids the difficulty in obtaining training samples for image dehazing models in the prior art, but also reduces the computational effort during training, significantly shortening the network training time and cost.

[0082] Based on the above embodiment, the initial atmospheric light layer generation network is used to encode the image to be dehazed to obtain latent variables, determine the normal distribution satisfied by the latent variables, reconstruct the latent variables based on the normal distribution to obtain reconstructed variables, and decode the reconstructed variables to obtain the initial atmospheric light layer;

[0083] The method further comprises:

[0084] Calculate the KL divergence loss based on the normal distribution and the standard normal distribution;

[0085] The calculating of the total loss based on the reconstruction loss and the atmospheric light loss specifically includes:

[0086] The total loss is calculated based on the reconstruction loss, the atmospheric light loss, and the KL divergence loss.

[0087] Specifically, after the image to be dehazed is input into the initial atmospheric light layer generation network, the initial atmospheric light layer generation network can encode the image to be dehazed to obtain latent variables, and can convert the latent variables into the mean and variance of the normal distribution by minimizing the equation, determine the normal distribution satisfied by the latent variables, and then reparameterize the normal distribution, reconstruct the latent variables by resampling in the normal distribution to obtain reconstructed variables, and decode the reconstructed variables to obtain the initial atmospheric light layer.

[0088] Here, the variational lower bound is reparameterized in order to be optimized end-to-end using standard stochastic gradient methods, resulting in a lower bound estimator that can be directly optimized using standard stochastic gradient methods.

[0089] Furthermore, this method can also introduce KL divergence loss, which can be calculated by normal distribution and standard normal distribution, that is:

[0090] ;

[0091] in, is the KL (Kullback-Leibler) divergence loss, used to enforce latent variables Conforms to the standard normal distribution , Normal distribution The mean of Normal distribution KL stands for KL divergence, which is used to measure the difference between two probability distributions, and || represents the KL divergence calculation between two probability distributions. represents the i-th dimension of the latent variable z.

[0092] Afterwards, the total loss can be calculated by weighted summing the reconstruction loss, atmospheric light loss, and KL divergence loss.

[0093] In the embodiment of the present invention, the KL divergence loss is included in the total loss, which can improve the performance of the target atmospheric light layer generation network in generating the target atmospheric light layer.

[0094] Based on the above embodiment, the initial atmospheric light layer generation network includes an encoder, an intermediate block and a decoder, and the encoder and the decoder both include the same number of blocks;

[0095] Each block in the encoder includes a convolution layer, an activation layer, and a maximum pooling layer connected in sequence, and the encoder is used to encode the image to be dehazed to obtain a latent variable;

[0096] The intermediate block is used to determine the normal distribution satisfied by the latent variable, and reconstruct the latent variable based on the normal distribution to obtain a reconstructed variable;

[0097] Each block in the decoder includes an upsampling layer, a convolution layer, a batch normalization layer and an activation layer connected in sequence, and the decoder is used to decode the reconstructed variables to obtain the initial atmospheric light layer.

[0098] Specifically, the initial A-Net can be a variational autoencoder structure, including an encoder, an intermediate block, and a decoder. The encoder can be built based on a convolutional neural network (CNN), and the decoder can be symmetrical with the encoder. The encoder and decoder both include the same number of convolutional blocks, for example Figure 3 As shown, both the encoder and decoder include 4 convolution blocks.

[0099] Each convolution block in the encoder consists of a 3×3 convolutional layer (conv), an activation layer (ReLU), and a maximum pooling layer (MaxPooling) connected in sequence. The encoder is used to encode the dehazed image to obtain a latent variable.

[0100] The middle block is used to determine the normal distribution satisfied by the latent variable, and reconstruct the latent variable based on the normal distribution to obtain the reconstructed variable.

[0101] Each convolution block in the decoder includes an upsampling layer (UpSample), a 3×3 convolution layer, a batch normalization layer (BatchNorm), and an activation layer (ReLU) connected in sequence. The decoder is used to decode the reconstructed variables to obtain the initial atmospheric light layer.

[0102] In addition, the initial A-Net may further include an output layer, which may include an upsampling layer (UpSample), a 3×3 convolution layer, and an activation layer (Sigmoid) connected in sequence to output an initial atmospheric light layer.

[0103] Based on the above embodiment, the calculating of the total loss based on the reconstruction loss and the atmospheric light loss further includes:

[0104] Calculating a dark channel prior loss based on the L1 norm of the dark channel value;

[0105] The total loss is calculated based on the reconstruction loss, the atmospheric light loss and the dark channel prior loss.

[0106] Specifically, this method can also introduce dark channel prior loss to constrain the Clear-Net network to form clean layers with dark channel statistical characteristics.

[0107] The dark channel prior loss can be calculated by the dark channel value, that is:

[0108] ;

[0109] in, is the dark channel prior loss, is the L1 norm.

[0110] After that, the total loss can be calculated by weighted summing the reconstruction loss, atmospheric light loss, and dark channel prior loss. It can be understood that in the presence of KL divergence loss, the reconstruction loss, atmospheric light loss, KL divergence loss, and dark channel prior loss can be weighted summed to calculate the total loss, thereby improving the overall performance of the three target layer generation networks.

[0111] Based on the above embodiment, the calculating of the total loss based on the reconstruction loss and the atmospheric light loss further includes:

[0112] Regularizing the initial atmospheric light layer and the initial transmittance layer respectively to obtain a regularized result of the atmospheric light layer and a regularized result of the transmittance layer;

[0113] Calculating a regularization loss based on the regularization result of the atmospheric light layer and the regularization result of the transmittance layer;

[0114] A total loss is calculated based on the reconstruction loss, the atmospheric light loss, and the regularization loss.

[0115] Specifically, this method can also introduce regularization loss to improve the stability of the three target layer generation networks.

[0116] To calculate the regularization loss, we can first regularize the initial atmospheric light layer and the initial transmittance layer respectively to obtain the regularized results of the atmospheric light layer and the transmittance layer. Then, we can perform a weighted sum of the regularized results of the atmospheric light layer and the transmittance layer to obtain the regularized loss.

[0117] ;

[0118] in, is the regularization loss, is the initial atmospheric light layer, is the initial transmittance layer, is the regularization result of the atmospheric light layer, is the regularization result of the transmittance layer, 、 is the balance factor, as the weighted weight, and there is , .

[0119] After that, the reconstruction loss, atmospheric light loss, and regularization loss are weighted and summed to calculate the total loss. It can be understood that in the presence of KL divergence loss, the reconstruction loss, atmospheric light loss, regularization loss, and KL divergence loss can be weighted and summed to obtain the total loss. In the presence of KL divergence loss and dark channel prior loss, the reconstruction loss, atmospheric light loss, regularization loss, KL divergence loss, and dark channel prior loss can be weighted and summed to calculate the total loss, thereby further improving the comprehensive performance of the three target layer generation networks. At this point, the total loss can be expressed as:

[0120] .

[0121] in, is the total loss, is atmospheric light loss.

[0122] On the basis of the above embodiment, regularizing the initial atmospheric light layer and the initial transmittance layer respectively to obtain the atmospheric light layer regularization result and the transmittance layer regularization result includes:

[0123] Based on the Laplace norm, the initial atmospheric light layer and the initial transmittance layer are regularized respectively to obtain a regularized result of the atmospheric light layer and a regularized result of the transmittance layer.

[0124] Specifically, the regularization result of the atmospheric light layer can be expressed as:

[0125]

[0126] The regularization result of the transmittance layer can be expressed as:

[0127]

[0128] in, for The second-order neighborhood of for The second-order neighborhood of and Both represent the neighborhood size.

[0129] By using the Laplace norm, the regularization process plays the role of mean filtering, so that the target atmospheric light layer and the target transmittance layer are smooth.

[0130] Based on the above embodiment, the initial clean layer generation network is a fully convolutional neural network, and the initial transmittance layer generation network is a U-net architecture.

[0131] Specifically, if Figure 4 As shown, the initial Clear-Net can be a fully convolutional network, including multiple convolutional activation blocks connected in sequence, a 3×3 convolutional layer and an activation layer (Sigmoid), each convolutional activation layer includes a 3×3 convolutional layer and an activation layer (Relu),

[0132] like Figure 5 As shown, the initial T-Net can be a U-net architecture, including a compression path, an expansion path and a 1×1 convolution layer. The compression path includes five 3×3 first-class convolution blocks connected in sequence. The expansion path includes five 3×3 second-class convolution blocks connected in sequence. The first-class convolution block includes a padding layer (Padding), a 3×3 convolution layer, a batch normalization layer (BatchNorm), an activation layer (Leaky ReLU), a padding layer (Padding), a 3×3 convolution layer, a batch normalization layer (BatchNorm) and an activation layer (Leaky ReLU) connected in sequence.

[0133] The second type of convolutional block includes an upsampling layer, a batch normalization layer, a padding layer, a 3×3 convolution layer, a batch normalization layer, an activation layer (Leaky ReLU), a padding layer, a 1×1 convolution layer, a batch normalization layer, and an activation layer (Leaky ReLU), which are connected in sequence.

[0134] In summary, the zero-shot learning dehazing image enhancement method based on image decomposition provided in the embodiments of the present invention uses a three-branch network to decompose the input haze image into three parts, namely, the clean image layer, the atmospheric light layer, and the transmittance layer. At the same time, the three-branch network is learned and decomposed based on only one input image to be dehazed, avoiding the difficulties of sample collection and high network training time cost, which is very important for subsequent image processing steps.

[0135] The image is decomposed through the network, and the dark channel prior loss is used to constrain the output of the initial Clear-Net to have statistical characteristics. The KL divergence loss and dark channel prior loss are used to constrain the output of the initial A-Net. The output of the initial T-Net is constrained by a regularization loss. The three target layers are synthesized using the atmospheric scattering model to synthesize the haze image and the input image to be dehazed. The reconstruction loss guides the learning to generate accurate target atmospheric light layer and target transmittance layer. Finally, the dehazed image is obtained based on the target atmospheric light layer and target transmittance layer. The three initial networks used in this method do not require large datasets for training. They can directly estimate the final dehazed image based on the input image to be dehazed, preserving the content details of the clear image as much as possible from various interferences. This improves efficiency and robustness for subsequent work, avoids intensive data collection and the use of synthetic foggy images to handle real-world domain transfer problems, and provides more reliable data and a higher level of automation for industrial production.

[0136] like Figure 6 As shown, based on the above embodiment, an embodiment of the present invention provides a zero-shot learning defogging image enhancement device based on image decomposition, comprising:

[0137] The layer generation module 61 is used to obtain the image to be defogged and input the image to be defogged into the initial clean layer generation network, the initial atmospheric light layer generation network, and the initial transmittance layer generation network, respectively, to obtain the initial clean layer output by the initial clean layer generation network, the initial atmospheric light layer output by the initial atmospheric light layer generation network, and the initial transmittance layer output by the initial transmittance layer generation network;

[0138] a loss calculation module 62 for applying an atmospheric scattering model to reconstruct a haze image based on the initial clean layer, the initial atmospheric light layer, and the initial transmittance layer, and calculating a reconstruction loss based on the image to be dehazed and the reconstructed haze image;

[0139] The loss calculation module 62 is further configured to determine a dark channel value of each pixel in the image to be defogged, determine a real atmospheric light layer corresponding to the image to be defogged based on the dark channel value, and calculate the atmospheric light loss based on the real atmospheric light layer and the initial atmospheric light layer;

[0140] a network optimization module 63 for calculating a total loss based on the reconstruction loss and the atmospheric light loss, and jointly optimizing the initial clean layer generation network, the initial atmospheric light layer generation network, and the initial transmittance layer generation network based on the total loss to obtain a target clean layer generation network, a target atmospheric light layer generation network, and a target transmittance layer generation network;

[0141] The image defogging module 64 is used to input the image to be defogged into the target atmospheric light layer generation network and the target transmittance layer generation network to obtain the target atmospheric light layer and the target transmittance layer respectively, and based on the target atmospheric light layer and the target transmittance layer, apply the atmospheric scattering model to solve the defogged image corresponding to the image to be defogged.

[0142] Based on the above embodiment, the initial atmospheric light layer generation network is used to encode the image to be dehazed to obtain latent variables, determine the normal distribution satisfied by the latent variables, reconstruct the latent variables based on the normal distribution to obtain reconstructed variables, and decode the reconstructed variables to obtain the initial atmospheric light layer;

[0143] The network optimization module is also used to:

[0144] Calculate the KL divergence loss based on the normal distribution and the standard normal distribution;

[0145] The calculating of the total loss based on the reconstruction loss and the atmospheric light loss specifically includes:

[0146] The total loss is calculated based on the reconstruction loss, the atmospheric light loss, and the KL divergence loss.

[0147] Based on the above embodiment, the initial atmospheric light layer generation network includes an encoder, an intermediate block and a decoder, and the encoder and the decoder both include the same number of convolution blocks;

[0148] Each convolution block in the encoder includes a convolution layer, an activation layer, and a maximum pooling layer connected in sequence, and the encoder is used to encode the image to be dehazed to obtain a latent variable;

[0149] The intermediate block is used to determine the normal distribution satisfied by the latent variable, and reconstruct the latent variable based on the normal distribution to obtain a reconstructed variable;

[0150] Each convolution block in the decoder includes an upsampling layer, a convolution layer, a batch normalization layer and an activation layer connected in sequence, and the decoder is used to decode the reconstructed variables to obtain the initial atmospheric light layer.

[0151] Based on the above embodiment, the network optimization module is further configured to:

[0152] Calculating a dark channel prior loss based on the L1 norm of the dark channel value;

[0153] The total loss is calculated based on the reconstruction loss, the atmospheric light loss and the dark channel prior loss.

[0154] Based on the above embodiment, the network optimization module is further configured to:

[0155] Regularizing the initial atmospheric light layer and the initial transmittance layer respectively to obtain a regularized result of the atmospheric light layer and a regularized result of the transmittance layer;

[0156] Calculating a regularization loss based on the regularization result of the atmospheric light layer and the regularization result of the transmittance layer;

[0157] A total loss is calculated based on the reconstruction loss, the atmospheric light loss, and the regularization loss.

[0158] Based on the above embodiment, the network optimization module is further configured to:

[0159] Based on the Laplace norm, the initial atmospheric light layer and the initial transmittance layer are regularized respectively to obtain a regularized result of the atmospheric light layer and a regularized result of the transmittance layer.

[0160] Based on the above embodiment, the initial clean layer generation network is a fully convolutional neural network, and the initial transmittance layer generation network is a U-net architecture.

[0161] Specifically, the functions of each module in the zero-sample learning dehazing image enhancement device based on image decomposition provided in the embodiment of the present invention correspond one-to-one to the operating procedures of each step in the above-mentioned method embodiment, and the effects achieved are also consistent. Please refer to the above-mentioned embodiment for details, and no further details will be given in the embodiment of the present invention.

[0162] Figure 7 An example of a physical structure diagram of an electronic device is shown below. Figure 7 As shown, the electronic device may include: a processor (Processor) 710, a communication interface (Communications Interface) 720, a memory (Memory) 730 and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 can call the logic instructions in the memory 730 to execute the zero-shot learning dehazing image enhancement method based on image decomposition provided in the above embodiments.

[0163] Furthermore, the logic instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0164] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the zero-sample learning dehazing image enhancement method based on image decomposition provided in the above embodiments.

[0165] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the zero-sample learning dehazing image enhancement method based on image decomposition provided in the above-mentioned embodiments.

[0166] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0167] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A zero-shot learning dehazing image enhancement method based on image decomposition, characterized in that: include: Acquire an image to be defogged, and input the image to be defogged into an initial clean layer generation network, an initial atmospheric light layer generation network, and an initial transmittance layer generation network, respectively, to obtain an initial clean layer output by the initial clean layer generation network, an initial atmospheric light layer output by the initial atmospheric light layer generation network, and an initial transmittance layer output by the initial transmittance layer generation network; Based on the initial clean layer, the initial atmospheric light layer, and the initial transmittance layer, an atmospheric scattering model is applied to reconstruct a haze image, and a reconstruction loss is calculated based on the image to be dehazed and the reconstructed haze image; The calculation formula of the reconstruction loss is as follows: ; Among them, n is the number of pixels in the image to be dehazed x and the haze image f, is the pixel value of pixel i in the haze image f, is the pixel value of pixel i in the image x to be dehazed, is the reconstruction loss; Determining a dark channel value of each pixel in the image to be defogged, determining a real atmospheric light layer corresponding to the image to be defogged based on the dark channel value, and calculating atmospheric light loss based on the real atmospheric light layer and the initial atmospheric light layer; The formula for determining the dark channel value is as follows: ; in, is the dark channel value of pixel i in the image to be dehazed, is a local image block centered on pixel i in the image to be dehazed, that is, a small patch around pixel i, c is the color channel, r is the red channel, g is the green channel, b is the blue channel, is the pixel value of pixel y in the color channel c of the image to be dehazed, y is Pixels in ; The formula for determining the real atmospheric light layer is as follows: ; in, is the real atmospheric light layer corresponding to the image to be dehazed, is a set of preset number of pixels with high dark channel values, is the pixel value of pixel j in the image to be defogged; The calculation formula of the atmospheric light loss is as follows: ; in, is the atmospheric light loss, is the initial atmospheric light layer; Based on the reconstruction loss and the atmospheric light loss, a total loss is calculated, and based on the total loss, the initial clean layer generation network, the initial atmospheric light layer generation network, and the initial transmittance layer generation network are jointly optimized to obtain a target clean layer generation network, a target atmospheric light layer generation network, and a target transmittance layer generation network; The image to be defogged is input into the target atmospheric light layer generation network and the target transmittance layer generation network to obtain a target atmospheric light layer and a target transmittance layer respectively. Based on the target atmospheric light layer and the target transmittance layer, an atmospheric scattering model is applied to solve the defogged image corresponding to the image to be defogged.

2. The zero-shot learning dehazing image enhancement method based on image decomposition according to claim 1 is characterized in that: The initial atmospheric light layer generation network is used to encode the image to be dehazed to obtain latent variables, determine a normal distribution satisfied by the latent variables, reconstruct the latent variables based on the normal distribution to obtain reconstructed variables, and decode the reconstructed variables to obtain the initial atmospheric light layer; The method further comprises: Calculate the KL divergence loss based on the normal distribution and the standard normal distribution; The calculating of the total loss based on the reconstruction loss and the atmospheric light loss specifically includes: The total loss is calculated based on the reconstruction loss, the atmospheric light loss, and the KL divergence loss.

3. The zero-shot learning dehazing image enhancement method based on image decomposition according to claim 2, characterized in that: The initial atmospheric light layer generation network includes an encoder, an intermediate block and a decoder, and the encoder and the decoder both include the same number of convolution blocks; Each convolution block in the encoder includes a convolution layer, an activation layer, and a maximum pooling layer connected in sequence, and the encoder is used to encode the image to be dehazed to obtain a latent variable; The intermediate block is used to determine the normal distribution satisfied by the latent variable, and reconstruct the latent variable based on the normal distribution to obtain a reconstructed variable; Each convolution block in the decoder includes an upsampling layer, a convolution layer, a batch normalization layer and an activation layer connected in sequence, and the decoder is used to decode the reconstructed variables to obtain the initial atmospheric light layer.

4. The zero-shot learning dehazing image enhancement method based on image decomposition according to claim 1, characterized in that: The calculating of the total loss based on the reconstruction loss and the atmospheric light loss further includes: Calculating a dark channel prior loss based on the L1 norm of the dark channel value; The total loss is calculated based on the reconstruction loss, the atmospheric light loss and the dark channel prior loss.

5. The zero-shot learning dehazing image enhancement method based on image decomposition according to claim 1, characterized in that: The calculating of the total loss based on the reconstruction loss and the atmospheric light loss further includes: Regularizing the initial atmospheric light layer and the initial transmittance layer respectively to obtain a regularized result of the atmospheric light layer and a regularized result of the transmittance layer; Calculating a regularization loss based on the regularization result of the atmospheric light layer and the regularization result of the transmittance layer; A total loss is calculated based on the reconstruction loss, the atmospheric light loss, and the regularization loss.

6. The zero-shot learning dehazing image enhancement method based on image decomposition according to claim 5, characterized in that: Regularizing the initial atmospheric light layer and the initial transmittance layer respectively to obtain the atmospheric light layer regularization result and the transmittance layer regularization result includes: Based on the Laplace norm, the initial atmospheric light layer and the initial transmittance layer are regularized respectively to obtain a regularized result of the atmospheric light layer and a regularized result of the transmittance layer.

7. The zero-shot learning dehazing image enhancement method based on image decomposition according to any one of claims 1 to 6, characterized in that: The initial clean layer generation network is a fully convolutional neural network, and the initial transmittance layer generation network is a U-net architecture.

8. A zero-shot learning dehazing image enhancement device based on image decomposition, characterized in that: include: A layer generation module is used to obtain an image to be defogged, and input the image to be defogged into an initial clean layer generation network, an initial atmospheric light layer generation network, and an initial transmittance layer generation network, respectively, to obtain an initial clean layer output by the initial clean layer generation network, an initial atmospheric light layer output by the initial atmospheric light layer generation network, and an initial transmittance layer output by the initial transmittance layer generation network; a loss calculation module, configured to apply an atmospheric scattering model to reconstruct a haze image based on the initial clean layer, the initial atmospheric light layer, and the initial transmittance layer, and calculate a reconstruction loss based on the image to be dehazed and the reconstructed haze image; The calculation formula of the reconstruction loss is as follows: ; Among them, n is the number of pixels in the image to be dehazed x and the haze image f, is the pixel value of pixel i in the haze image f, is the pixel value of pixel i in the image x to be dehazed, is the reconstruction loss; The loss calculation module is further configured to determine a dark channel value of each pixel in the image to be defogged, determine a real atmospheric light layer corresponding to the image to be defogged based on the dark channel value, and calculate the atmospheric light loss based on the real atmospheric light layer and the initial atmospheric light layer; The formula for determining the dark channel value is as follows: ; in, is the dark channel value of pixel i in the image to be dehazed, is a local image block centered on pixel i in the image to be dehazed, that is, a small patch around pixel i, c is the color channel, r is the red channel, g is the green channel, b is the blue channel, is the pixel value of pixel y in the color channel c of the image to be dehazed, y is Pixels in ; The formula for determining the real atmospheric light layer is as follows: ; in, is the real atmospheric light layer corresponding to the image to be dehazed, is a set of preset number of pixels with high dark channel values, is the pixel value of pixel j in the image to be defogged; The calculation formula of the atmospheric light loss is as follows: ; in, is the atmospheric light loss, is the initial atmospheric light layer; a network optimization module, configured to calculate a total loss based on the reconstruction loss and the atmospheric light loss, and jointly optimize the initial clean layer generation network, the initial atmospheric light layer generation network, and the initial transmittance layer generation network based on the total loss to obtain a target clean layer generation network, a target atmospheric light layer generation network, and a target transmittance layer generation network; The image defogging module is used to input the image to be defogged into the target atmospheric light layer generation network and the target transmittance layer generation network to obtain the target atmospheric light layer and the target transmittance layer respectively, and based on the target atmospheric light layer and the target transmittance layer, apply the atmospheric scattering model to solve the defogged image corresponding to the image to be defogged.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the zero-shot learning dehazing image enhancement method based on image decomposition is implemented as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the zero-shot learning dehazing image enhancement method based on image decomposition is implemented as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Traffic image defogging method based on improved generative adversarial network

    CN112801902A

  • Image defogging method based on generative adversarial network fused with feature pyramid

    WO2021248938A1