Non-paired image defogging method based on cyclic diffusion model
Through the unpaired image defogging method based on the cyclic diffusion model, the non-pairing training is used to use real fogging images and clear images to solve the problem of poor image defogging effect in the prior art, and better visual quality and generalization ability are achieved.
Patent Information
- Application Number
- CN202411975235.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing image defogging method is not effective in real fogging image processing, especially the deep learning-based methods lack the ability to generalize on real data.
The unpaired image defogging method based on the cyclic diffusion model is adopted. By constructing a defogging branch network and a fogging branch network, the real fogging image and clear images are used for unpaired training, and the prior information in the cyclic diffusion model is fully utilized.
The image defogging effect is improved, so that the generated defogging images have better visual quality, and the generalization ability of the defogging network is enhanced, so that real fogging images can be processed more effectively.
Smart Images

Figure CN119991497A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image defogging processing, and in particular relates to an unpaired image defogging method based on a cyclic diffusion model. Background Art
[0002] Affected by haze weather, the quality of images collected by intelligent video surveillance systems is seriously degraded, with reduced clarity, reduced contrast, blurred details, and other degradation phenomena, which in turn affects subsequent image processing tasks such as object recognition and scene understanding. Image dehazing aims to restore images obtained under haze weather conditions to haze-free images, thereby improving the visual quality of images.
[0003] Existing image defogging methods are mainly divided into defogging methods based on image enhancement, image defogging methods based on prior information, and defogging methods based on deep learning. Image defogging methods based on prior information only process the contrast and color information of the image, and have a certain defogging effect. However, the defogging effect of this type of method is limited. Image defogging methods based on prior information take the image itself as the research object, use the information contained in the image itself as prior knowledge to obtain the transmission map and atmospheric light intensity, and then substitute it into the atmospheric scattering model to obtain the defogging image. This type of method takes into account the reasons for the formation of haze images in real scenes and can effectively restore the texture information of haze-free images. However, since the prior information artificially set for certain specific scenes is not applicable to all scenes, this type of method often over-defogs, resulting in a decrease in the quality of the defogged image. Image defogging methods based on deep learning have become the mainstream at this stage. Image defogging is achieved by learning the mapping between foggy images and clear images in a data-driven way. Such methods require a large amount of paired real image data for training. However, in the real world, it is difficult to collect paired real foggy training data sets and it takes a lot of manpower and material resources. Therefore, most existing studies use synthetic foggy data for training and can achieve good defogging effects on synthetic test data. However, due to the difference between synthetic foggy data and real foggy data, such methods are often unable to effectively defog real foggy images. To solve this problem, they rely on foggy images and clear images of real scenes and use an unpaired method for training. Such methods rely on real data for training and have good generalization capabilities on real image data.
[0004] As the cyclic diffusion model has demonstrated strong generative capabilities in the field of image generation, fine-tuning it can not only take advantage of its powerful feature representation capabilities, but also greatly reduce training costs.
[0005] Therefore, a well-designed unpaired image dehazing method based on a cyclic diffusion model is needed. The image dehazing method based on unpaired training can effectively utilize the training data of real scenes, so that the trained dehazing model has better generalization ability and the generated dehazed image has better visual quality. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide an unpaired image defogging method based on a cyclic diffusion model in view of the deficiencies in the above-mentioned prior art. The method has simple steps and reasonable design. A defogging network based on a cyclic diffusion model is trained based on unpaired real foggy images and real clear images. While ensuring the defogging capability, the prior information contained in the stable cyclic diffusion model in the defogging network based on the cyclic diffusion model is fully utilized, so that the generated defogging image has better visual quality and the defogging network has better generalization ability.
[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is: a non-paired image defogging method based on a cyclic diffusion model, characterized in that the method comprises the following steps:
[0008] Step 1: Obtaining training set images:
[0009] Collect real foggy images from the RTTS dataset and the URHI dataset in the foggy image database RESIDE, and collect real clear images from the ADE20K database and the OTS dataset in the foggy image database RESIDE as training data; among them, one real foggy image and one real clear image are used as a set of training data;
[0010] Step 2: Build a defogging network based on the cyclic diffusion model:
[0011] The defogging network based on the cyclic diffusion model includes a defogging branch network and a fogging branch network, the defogging branch network includes a first defogging network module, a first fogging network module and a defogging discriminator, and the fogging branch network includes a second fogging network module, a second defogging network module and a fogging discriminator; wherein the first defogging network module, the second defogging network module, the first fogging network module and the second fogging network module have the same structure, and the defogging discriminator and the fogging discriminator have the same structure;
[0012] The first defogging network module and the second defogging network module both include a defogging backbone network, a text perception guidance module, and a physical perception guidance module;
[0013] The first fogging network module and the second fogging network module both include a fogging backbone network and a text perception guidance module; the defogging backbone network and the fogging backbone network are both stable cyclic diffusion models;
[0014] The defogging backbone network and the fogging backbone network have the same structure and both include a VAE encoder, a UNet and a VAE decoder, and four zero convolution layers are added between the VAE encoder and the VAE decoder;
[0015] Step 3: Feature extraction of real foggy images and real clear images:
[0016] Step 301: Process the real foggy image x through the text perception guidance module in the first defogging network module to obtain a foggy image with text control conditions added, and pass the foggy image with text control conditions through the defogging backbone network in the first defogging network module to obtain a clear image after defogging.
[0017] The real foggy image x is processed by the text perception guidance module in the second fogging network module to obtain an input foggy image with text control conditions added, and the input foggy image with text control conditions is processed by the fogging backbone network in the second fogging network module to obtain an output fogged image F(x);
[0018] The real foggy image x passes through the physical perception guidance module in the first defogging network module to obtain the first reconstructed foggy image I phy ;
[0019] The clear image after defogging After the text perception guidance module in the first fogging network module, the defogging image with the text control condition is obtained. The defogging image with the text control condition is processed by the fogging backbone network in the first fogging network module to obtain a synthetic foggy image.
[0020] Step 302: Process the real clear image y through the text perception guidance module in the second fogging network module to obtain a clear image with text control conditions added, and pass the clear image with text control conditions through the fogging backbone network in the second fogging network module to obtain a fogged image.
[0021]
[0022] The real clear image y is processed by the text perception guidance module in the first defogging network module to obtain an input clear image with text control conditions added, and the input clear image with text control conditions is passed through the defogging backbone network of the first defogging network module to obtain an output clear image G(y);
[0023] The fogged image After the text perception guidance module in the second defogging network module, the fogged image with the text control condition is obtained, and the fogged image with the text control condition is processed by the defogging backbone network in the second defogging network module to obtain a synthetic defogging image
[0024] The fogged image After the physical perception guidance module in the second defogging network module, the second reconstructed foggy image I′ is obtained phy ;
[0025] Step 4: Establishment of total loss function:
[0026] Step 401: According to Get the generated adversarial loss L GAN Among them, D Y (y) represents the value output by the dehazing discriminator for the real clear image y. Represents a clear image after dehazing The value output by the dehazing discriminator, D X (x) represents the value output by the fogging discriminator for the real foggy image x. Represents the image after adding fog The value output by the fogging discriminator;
[0027] Step 402: According to Get the cycle consistency loss L cyc ;in, Represents the real foggy image x and the synthetic foggy image The Manhattan distance between Represents the real foggy image x and the synthetic foggy image The LPIPS distance between Represents the real clear image y and the synthetic dehazed image The Manhattan distance between Represents the real clear image y and the synthetic dehazed image The LPIPS distance between them;
[0028] Step 403: According to L idt =L rec1 (G(y),y)+L rec2 (G(y),y)+L rec1 (F(x),x)+L rec2 (F(x),x), we get the identity loss L idt Among them, L rec1 (G(y), y) represents the Manhattan distance between the real clear image y and the output clear image G(y), L rec2(G(y),y) represents the LPIPS distance between the real clear image y and the output clear image G(y), L rec1 (F(x),x) represents the Manhattan distance between the real foggy image x and the output fogged image F(x), L rec2 (F(x),x) represents the LPIPS distance between the real foggy image x and the output fogged image F(x);
[0029] Step 404: According to L phy (1) = L phy +L p ' hy , and get the reconstruction loss L phy (1); where L phy represents the first branch reconstruction loss, and L phy =L rec1 (x,I phy )+L rec2 (x,I phy ), L rec1 (x,I phy ) represents the real foggy image x and the first reconstructed foggy image I phy The Manhattan distance between rec2 (x,I phy ) represents the real foggy image x and the first reconstructed foggy image I phy LPIPS distance between p ' hy represents the second branch reconstruction loss, and Represents the image after adding fog and the second reconstructed foggy image I′ phy The Manhattan distance between Represents the image after adding fog and the second reconstructed foggy image I′ phy The LPIPS distance between them;
[0030] Step 405: According to L loss =L GAN +L cyc +L idt +0.5L phy (1), we get the total loss function L loss ;
[0031] Step 5: Training of defogging network based on cyclic diffusion model:
[0032] Step 501: The computer uses the Adam optimization algorithm to input a set of training data and first uses the generated adversarial loss L GAN The fogging discriminator and the defogging discriminator are trained until the generation adversarial loss L is obtained. GANFogging discriminator and defogging discriminator at maximum value;
[0033] Step 502: The computer uses the Adam optimization algorithm to input the set of training data, and uses the total loss function L under the fogging discriminator and defogging discriminator determined in step 501. loss The dehazing network based on the cyclic diffusion model is trained until the total loss function L loss Minimum, complete the training of this set of training data;
[0034] Step 503, follow the method from step 501 to step 502 until all training sets are trained and one iteration training is completed;
[0035] Step 504, repeating steps 501 to 503, iterative training until the preset number of iterative training times is met, and a trained defogging network based on the cyclic diffusion model is obtained;
[0036] Step 6: Use the trained defogging network based on the cyclic diffusion model to defog a single image:
[0037] A computer is used to input any foggy image into the first defogging network module in the defogging branch network of the trained cyclic diffusion model defogging network to perform defogging processing to obtain a clear image after defogging.
[0038] The above-mentioned unpaired image dehazing method based on the cyclic diffusion model is characterized in that: in step 2, the four zero convolution layers added between the VAE encoder and the VAE decoder are respectively the first zero convolution layer, the second zero convolution layer, the third zero convolution layer and the fourth zero convolution layer;
[0039] The VAE encoder includes an input convolutional layer, a first DownEncoderBlock2D, a second DownEncoderBlock2D, a third DownEncoderBlock2D, and a fourth DownEncoderBlock2D;
[0040] The VAE decoder includes an output convolutional layer plus UNetMidBlock2D, a first UpDecoderBlock2D, a second UpDecoderBlock2D, a third UpDecoderBlock2D, and a fourth UpDecoderBlock2D;
[0041] A first zero convolution layer is set between the output of the VAE encoder input convolution layer and the input of the fourth UpDecoderBlock2D, a second zero convolution layer is set between the output of the first DownEncoderBlock2D and the input of the third UpDecoderBlock2D, a third zero convolution layer is set between the output of the second DownEncoderBlock2D and the input of the second UpDecoderBlock2D, and a fourth zero convolution layer is set between the output of the third DownEncoderBlock2D and the input of the first UpDecoderBlock2D.
[0042] The above-mentioned unpaired image defogging method based on the cyclic diffusion model is characterized in that: the foggy image with the text control condition is passed through the defogging backbone network in the first defogging network module, the fogged image with the text control condition is passed through the defogging backbone network in the second defogging network module, the input clear image with the text control condition is passed through the defogging backbone network of the first defogging network module, the defogging image with the text control condition is passed through the fogging backbone network in the first fogging network module, the clear image with the text control condition is passed through the fogging backbone network in the second fogging network module, and the input foggy image with the text control condition is passed through the fogging backbone network in the second fogging network module. The processing method is the same, specifically as follows:
[0043] Step 3011: the foggy image with text control condition added, the fogged image with text control condition added, the input clear image with text control condition added, the defogged image with text control condition added, the clear image with text control condition added, and the input foggy image with text control condition added are recorded as input image and control text;
[0044] Step 3012: Pass the input image through the VAE encoder to obtain an encoded output feature map; wherein the input image is input into the input convolution layer of the VAE encoder to obtain a first intermediate feature map, and the first intermediate feature map is sequentially passed through the first DownEncoderBlock2D, the second DownEncoderBlock2D, the third DownEncoderBlock2D and the fourth DownEncoderBlock2D to obtain a second intermediate feature map, a third intermediate feature map, a fourth intermediate feature map and a fifth intermediate feature map;
[0045] Step 3013: The encoded output feature map and the control text are passed through UNet to obtain a sixth intermediate feature map;
[0046] Step 3014: the first intermediate feature map is passed through the first zero convolution layer to output a first transition feature map, the second intermediate feature map is passed through the second zero convolution layer to output a second transition feature map, the third intermediate feature map is passed through the third zero convolution layer to output a third transition feature map, and the fourth intermediate feature map is passed through the fourth zero convolution layer to output a fourth transition feature map;
[0047] Step 3015, the sixth intermediate feature map, the first transition feature map, the second transition feature map, the third transition feature map and the fourth transition feature map are passed through the VAE decoder to obtain a decoded output feature map; wherein, the sixth intermediate feature map is passed through the output convolution layer of the VAE decoder and UNetMidBlock2D is added to obtain a seventh intermediate feature map; the seventh intermediate feature map and the fourth transition feature map are added and input into the first UpDecoderBlock2D for processing to obtain an eighth intermediate feature map; the eighth intermediate feature map and the third transition feature map are added and input into the second UpDecoderBlock2D for processing to obtain a ninth intermediate feature map; the ninth intermediate feature map and the second transition feature map are added and input into the third UpDecoderBlock2D for processing to obtain a tenth intermediate feature map; the tenth intermediate feature map and the first transition feature map are added and input into the fourth UpDecoderBlock2D for processing to obtain an eleventh intermediate feature map;
[0048] Step 3016: record the decoded output feature map corresponding to the foggy image with the text control condition added as the clear image after defogging The decoded output feature map corresponding to the output of the fogged image with the text control condition added is recorded as the synthetic defogging image The decoded output feature map corresponding to the input clear image with the text control condition added is recorded as the output clear image G(y);
[0049] The decoded output feature map corresponding to the defogging image with text control condition added is recorded as the synthetic foggy image The decoded output feature map corresponding to the clear image with text control condition is recorded as the fogged image The decoded output feature map corresponding to the output of the input foggy image with the text control condition added is recorded as the output foggy image F(x).
[0050] The above-mentioned unpaired image defogging method based on the cyclic diffusion model is characterized in that: the real foggy image x is processed by the text perception guidance module in the first defogging network module, the real clear image y is processed by the text perception guidance module in the first defogging network module, and the fogged image The methods of the text perception guidance module in the second defogging network module are the same, as follows:
[0051] Step 3021: The real foggy image x, the real clear image y and the fogged image Both are recorded as the first initial image;
[0052] Step 3022: Input the first initial image into the BLIP2 module to obtain a first text title;
[0053] Step 3023: Process the real foggy image in the training data by using the text inversion method to obtain negative prompt features representing fog;
[0054] Step 3024: Delete haze-related words from the first text title and send it to the CLIP model text encoder to obtain text embedding representation features;
[0055] Step 3025, embedding the text representation feature and the negative prompt feature representing fog through the classifier-free guide to obtain the control of this paper;
[0056] Step 3026: The real foggy image x is processed by the text perception guidance module in the first defogging network module to obtain the control image and the real foggy image x as the foggy image with the text control condition added;
[0057] The real clear image y is processed by the text perception guidance module in the first defogging network module to obtain the control image and the real clear image y of this paper, which are recorded as the input clear image with the text control condition added;
[0058] Image after fogging The control and fogged images obtained by the text-aware guidance module in the second defogging network module Denote as fogged image with text control condition added;
[0059] Clear image after dehazing The method of processing the real foggy image x through the text perception guidance module in the first fogging network module, the real foggy image x through the text perception guidance module in the second fogging network module, and the real clear image y through the text perception guidance module in the second fogging network module is the same as follows:
[0060] Step 30A: Defogging the clear image The real foggy image x and the real clear image y are recorded as the second initial image;
[0061] Step 30B, input the second initial image into the BLIP2 module to obtain a second text title;
[0062] Step 30C, adding haze-related words to the second text title and sending it to the CLIP model text encoder to obtain a second text embedding representation feature;
[0063] Step 30D, embedding the second text representation feature through the non-classifier guide to obtain the second text control;
[0064] Step 30E: Defogging the clear image The second text control and the clear image after defogging are obtained by the text perception guidance module in the first fogging network module Denoted as the dehazed image with text control condition added;
[0065] The second text control and the real clear image y obtained by the text perception guidance module in the second fogging network module are recorded as the input clear image with the text control condition added;
[0066] The second text control obtained by passing the real foggy image x through the text perception guidance module in the second fogging network module and the real foggy image x are recorded as the input foggy image with the text control condition added.
[0067] The above-mentioned unpaired image defogging method based on the cyclic diffusion model is characterized in that: in step 301, the real foggy image x is passed through the physical perception guidance module in the first defogging network module to obtain a first reconstructed foggy image I phy and the image after adding fog in step 302 After the physical perception guidance module in the second defogging network module, the second reconstructed foggy image I′ is obtained phy The method is the same as the following:
[0068] Step A: Use the dark channel prior defogging DCP algorithm to process the real foggy image x to obtain the first defogging feature map J dcp , the first transmission image t dcp and atmospheric light A;
[0069] The boundary constraint prior BCCR algorithm is used to process the real foggy image x to obtain the second defogging feature map J bccr and the second transmission image t bccr ;
[0070] Step B: Use a computer to call the concatenate function to combine the real foggy image x and the first transmission image t dcp and the second transmission image t bccr After splicing, input the transmission map optimization network to obtain the optimized transmission map t ref ;
[0071] Step C: The real foggy image x and the first defogging feature map J dcp and the second dehazing feature map J bccr After the physical fusion module PFM, the fusion feature map F is obtained fuse ;
[0072] Step D: Fuse the feature map F fuse Input the dehazing backbone network to obtain the optimized dehazing image J ref ;
[0073] Step F: Using the atmospheric scattering model ASM to optimize the defogging image J ref , optimized transmission map t ref and atmospheric light A to obtain the reconstructed foggy image I phy ;
[0074] Step G: According to the method of step A to step F, the fogged image After processing, the second reconstructed foggy image I′ is obtained phy .
[0075] The above-mentioned unpaired image defogging method based on the cyclic diffusion model is characterized in that: step C, the specific process is as follows:
[0076] Step C01: The first defogging feature map J dcp After the first convolution + ReLU activation function layer, the first feature map f1 is obtained;
[0077] The real foggy image x passes through the first double-layer convolution module to obtain the second feature map f2;
[0078] The second dehazing feature map J bccr After the second convolution + ReLU activation function layer, the third feature map f3 is obtained;
[0079] Step C02, adding the second characteristic map f2 to the first characteristic map f1 to obtain a fourth characteristic map f4, and adding the second characteristic map f2 to the third characteristic map f3 to obtain a fifth characteristic map f5;
[0080] Step C03, the fourth feature map f4 is passed through the third convolution + ReLU activation function layer to obtain the sixth feature map f6;
[0081] The fifth feature map f5 passes through the fourth convolution + ReLU activation function layer to obtain the seventh feature map f7;
[0082] Step C04, performing a Hadamard product operation on the output of the first Sigmoid function layer of the sixth feature map f6 and the fourth feature map f4 to obtain an eighth feature map f8;
[0083] The output of the seventh feature map f7 after the second Sigmoid function layer is subjected to a Hadamard product operation with the fifth feature map f5 to obtain a ninth feature map f9;
[0084] Step C05: Use the concatenate function to concatenate the eighth feature map f8 and the ninth feature map f9 to obtain the tenth feature map f 10 , and the tenth feature map f 10 After the global adaptive pooling layer, the second double-layer convolution module and the Softmax layer, the eleventh feature map f is obtained. 11 ;
[0085] Step C06: The eleventh feature map f 11 With the tenth characteristic graph f 10 The features after the Hadamard product operation and the tenth feature map f 10 Add together to get the twelfth feature map f 12 ;
[0086] Step C07: Use the concatenate function to concatenate the first defogging feature map J dcp , the second defogging feature map J bccr and the twelfth characteristic graph f 12 After splicing, it is sent to the gated convolutional network to obtain the fused feature weight f w ;
[0087] Step C08: Use a computer to calculate the Get the fusion feature map F fuse ;in, represents the Hadamard product operation between feature map matrices, represents the addition operation between feature map matrices; f w1 Represents the feature weight f after fusion w The eigenvalue of the first channel in, f w2 Represents the feature weight f after fusion w The eigenvalue of the second channel in w3 Represents the feature weight f after fusion w The eigenvalue of the third channel in w4 Represents the feature weight f after fusion w The eigenvalue of the fourth channel in; f 12-1 represents the twelfth feature map f 12 Feature maps of the first three channels in f 12-2 represents the twelfth feature map f 12 Feature maps of the last three channels.
[0088] Compared with the prior art, the present invention has the following advantages:
[0089] 1. The method of the present invention has simple steps and reasonable design. First, the training set images are obtained; second, a defogging network based on a cyclic diffusion model is constructed, followed by feature extraction of real foggy images and feature extraction of real clear images, then a total loss function is established, and then the defogging network based on the cyclic diffusion model is trained. Finally, a single image is defogged using the trained defogging network based on the cyclic diffusion model to improve the image defogging effect.
[0090] 2. The defogging network module of the present invention includes a defogging backbone network, a text perception guidance module, and a physical perception guidance module. The fogging network module includes a fogging backbone network, a text perception guidance module, and a physical perception guidance module. The text perception guidance module is introduced in order to improve the quality of the defogging image by using text modality defogging; the physical perception guidance module is introduced in order to comprehensively utilize the advantages of the image defogging algorithm based on prior information and the atmospheric scattering model, so as to make the defogging more effective, especially for images with high haze concentration; the defogging backbone network and the fogging backbone network are both stable cyclic diffusion models, thereby making full use of the prior information contained in the stable cyclic diffusion model in the defogging network based on the cyclic diffusion model, so that the generated defogging image has better visual quality, and the defogging network has better generalization ability.
[0091] 3. Four zero convolution layers are added between the VAE encoder and the VAE decoder of the present invention to ensure that the generated image retains the basic content and details of the original image, ensuring that the details of the image after defogging or fogging are the same as the original image without distortion.
[0092] 4. The present invention utilizes unpaired real foggy images and real clear images to train the defogging network based on the cyclic diffusion model, so that the unpaired method can be processed on real foggy and real clear images, thereby enhancing the method's processing capability for real foggy images, solving the current difficulties in collecting large-scale paired real foggy images and clear images, and the problem that the defogging of synthetic paired foggy images and clear images cannot effectively process real foggy images due to the differences between real foggy images and synthetic foggy images.
[0093] In summary, the method of the present invention has simple steps and reasonable design. The defogging network based on the cyclic diffusion model is trained based on unpaired real foggy images and real clear images. While ensuring the defogging ability, the prior information contained in the stable cyclic diffusion model in the defogging network based on the cyclic diffusion model is fully utilized, so that the generated defogging image has better visual quality, and the defogging network has better generalization ability.
[0094] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 The figure is a flowchart of the method of the present invention.
[0096] Figure 2 It is a schematic diagram of the structure of the defogging network based on the cyclic diffusion model of the present invention.
[0097] Figure 3 It is a structural schematic diagram of the circulating diffusion model of the present invention.
[0098] Figure 4 It is a structural schematic diagram of the physical perception guidance module of the present invention.
[0099] Figure 5 It is a structural schematic diagram of the physical fusion module of the present invention.
[0100] Figure 6 This is an image tested using the method of the present invention. DETAILED DESCRIPTION
[0101] like Figures 1 to 5 As shown, the present invention provides an unpaired image defogging method based on a cyclic diffusion model, comprising the following steps:
[0102] Step 1: Obtaining training set images:
[0103] Collect real foggy images from the RTTS dataset and the URHI dataset in the foggy image database RESIDE, and collect real clear images from the ADE20K database and the OTS dataset in the foggy image database RESIDE as training data; among them, one real foggy image and one real clear image are used as a set of training data;
[0104] Step 2: Build a defogging network based on the cyclic diffusion model:
[0105] The defogging network based on the cyclic diffusion model includes a defogging branch network and a fogging branch network, the defogging branch network includes a first defogging network module, a first fogging network module and a defogging discriminator, and the fogging branch network includes a second fogging network module, a second defogging network module and a fogging discriminator; wherein the first defogging network module, the second defogging network module, the first fogging network module and the second fogging network module have the same structure, and the defogging discriminator and the fogging discriminator have the same structure;
[0106] The first defogging network module and the second defogging network module both include a defogging backbone network, a text perception guidance module, and a physical perception guidance module;
[0107] The first fogging network module and the second fogging network module both include a fogging backbone network and a text perception guidance module; the defogging backbone network and the fogging backbone network are both stable cyclic diffusion models;
[0108] The defogging backbone network and the fogging backbone network have the same structure and both include a VAE encoder, a UNet and a VAE decoder, and four zero convolution layers are added between the VAE encoder and the VAE decoder;
[0109] Step 3: Feature extraction of real foggy images and real clear images:
[0110] Step 301: Process the real foggy image x through the text perception guidance module in the first defogging network module to obtain a foggy image with text control conditions added, and pass the foggy image with text control conditions through the defogging backbone network in the first defogging network module to obtain a clear image after defogging.
[0111] The real foggy image x is processed by the text perception guidance module in the second fogging network module to obtain an input foggy image with text control conditions added, and the input foggy image with text control conditions is processed by the fogging backbone network in the second fogging network module to obtain an output fogged image F(x);
[0112] The real foggy image x passes through the physical perception guidance module in the first defogging network module to obtain the first reconstructed foggy image I phy ;
[0113] The clear image after defogging After the text perception guidance module in the first fogging network module, the defogging image with the text control condition is obtained. The defogging image with the text control condition is processed by the fogging backbone network in the first fogging network module to obtain a synthetic foggy image.
[0114] Step 302: Process the real clear image y through the text perception guidance module in the second fogging network module to obtain a clear image with text control conditions added, and pass the clear image with text control conditions through the fogging backbone network in the second fogging network module to obtain a fogged image.
[0115]
[0116] The real clear image y is processed by the text perception guidance module in the first defogging network module to obtain an input clear image with text control conditions added, and the input clear image with text control conditions is passed through the defogging backbone network of the first defogging network module to obtain an output clear image G(y);
[0117] The fogged image After the text perception guidance module in the second defogging network module, the fogged image with the text control condition is obtained, and the fogged image with the text control condition is processed by the defogging backbone network in the second defogging network module to obtain a synthetic defogging image
[0118] The fogged image After the physical perception guidance module in the second defogging network module, the second reconstructed foggy image I′ is obtained phy ;
[0119] Step 4: Establishment of total loss function:
[0120] Step 401: According to Get the generated adversarial loss L GAN Among them, D Y (y) represents the value output by the dehazing discriminator for the real clear image y. Represents a clear image after dehazing The value output by the dehazing discriminator, D X (x) represents the value output by the fogging discriminator for the real foggy image x. Represents the image after adding fog The value output by the fogging discriminator;
[0121] Step 402: According to Get the cycle consistency loss L cyc ;in, Represents the real foggy image x and the synthetic foggy image The Manhattan distance between Represents the real foggy image x and the synthetic foggy image The LPIPS distance between Represents the real clear image y and the synthetic dehazed image The Manhattan distance between Represents the real clear image y and the synthetic dehazed image The LPIPS distance between them;
[0122] Step 403: According to L idt =L rec1 (G(y),y)+L rec2 (G(y),y)+L rec1 (F(x),x)+L rec2 (F(x),x), we get the identity loss L idt Among them, L rec1 (G(y), y) represents the Manhattan distance between the real clear image y and the output clear image G(y), L rec2(G(y),y) represents the LPIPS distance between the real clear image y and the output clear image G(y), L rec1 (F(x),x) represents the Manhattan distance between the real foggy image x and the output fogged image F(x), L rec2 (F(x),x) represents the LPIPS distance between the real foggy image x and the output fogged image F(x);
[0123] Step 404: According to L phy (1) = L phy +L p ' hy , and get the reconstruction loss L phy (1); where L phy represents the first branch reconstruction loss, and L phy =L rec1 (x,I phy )+L rec2 (x,I phy ), L rec1 (x,I phy ) represents the real foggy image x and the first reconstructed foggy image I phy The Manhattan distance between rec2 (x,I phy ) represents the real foggy image x and the first reconstructed foggy image I phy LPIPS distance between p ' hy represents the second branch reconstruction loss, and Represents the image after adding fog and the second reconstructed foggy image I′ phy The Manhattan distance between Represents the image after adding fog and the second reconstructed foggy image I′ phy The LPIPS distance between them;
[0124] Step 405: According to L loss =L GAN +L cyc +L idt +0.5L phy (1), we get the total loss function L loss ;
[0125] Step 5: Training of defogging network based on cyclic diffusion model:
[0126] Step 501: The computer uses the Adam optimization algorithm to input a set of training data and first uses the generated adversarial loss L GAN The fogging discriminator and the defogging discriminator are trained until the generation adversarial loss L is obtained. GANFogging discriminator and defogging discriminator at maximum value;
[0127] Step 502: The computer uses the Adam optimization algorithm to input the set of training data, and uses the total loss function L under the fogging discriminator and defogging discriminator determined in step 501. loss The dehazing network based on the cyclic diffusion model is trained until the total loss function L loss Minimum, complete the training of this set of training data;
[0128] Step 503, follow the method from step 501 to step 502 until all training sets are trained and one iteration training is completed;
[0129] Step 504, repeating steps 501 to 503, iterative training until the preset number of iterative training times is met, and a trained defogging network based on the cyclic diffusion model is obtained;
[0130] Step 6: Use the trained defogging network based on the cyclic diffusion model to defog a single image:
[0131] A computer is used to input any foggy image into the first defogging network module in the defogging branch network of the trained cyclic diffusion model defogging network to perform defogging processing to obtain a clear image after defogging.
[0132] In this embodiment, in step 2, the four zero convolution layers added between the VAE encoder and the VAE decoder are the first zero convolution layer, the second zero convolution layer, the third zero convolution layer and the fourth zero convolution layer;
[0133] The VAE encoder includes an input convolutional layer, a first DownEncoderBlock2D, a second DownEncoderBlock2D, a third DownEncoderBlock2D, and a fourth DownEncoderBlock2D;
[0134] The VAE decoder includes an output convolutional layer plus UNetMidBlock2D, a first UpDecoderBlock2D, a second UpDecoderBlock2D, a third UpDecoderBlock2D, and a fourth UpDecoderBlock2D;
[0135] A first zero convolution layer is set between the output of the VAE encoder input convolution layer and the input of the fourth UpDecoderBlock2D, a second zero convolution layer is set between the output of the first DownEncoderBlock2D and the input of the third UpDecoderBlock2D, a third zero convolution layer is set between the output of the second DownEncoderBlock2D and the input of the second UpDecoderBlock2D, and a fourth zero convolution layer is set between the output of the third DownEncoderBlock2D and the input of the first UpDecoderBlock2D.
[0136] In this embodiment, the foggy image with the text control condition is passed through the defogging backbone network in the first defogging network module, the fogged image with the text control condition is passed through the defogging backbone network in the second defogging network module, the input clear image with the text control condition is passed through the defogging backbone network of the first defogging network module, the defogging image with the text control condition is passed through the fogging backbone network in the first fogging network module, the clear image with the text control condition is passed through the fogging backbone network in the second fogging network module, and the input foggy image with the text control condition is passed through the fogging backbone network in the second fogging network module. The processing method is the same, as follows:
[0137] Step 3011: the foggy image with text control condition added, the fogged image with text control condition added, the input clear image with text control condition added, the defogged image with text control condition added, the clear image with text control condition added, and the input foggy image with text control condition added are recorded as input image and control text;
[0138] Step 3012: Pass the input image through the VAE encoder to obtain an encoded output feature map; wherein the input image is input into the input convolution layer of the VAE encoder to obtain a first intermediate feature map, and the first intermediate feature map is sequentially passed through the first DownEncoderBlock2D, the second DownEncoderBlock2D, the third DownEncoderBlock2D and the fourth DownEncoderBlock2D to obtain a second intermediate feature map, a third intermediate feature map, a fourth intermediate feature map and a fifth intermediate feature map;
[0139] Step 3013: The encoded output feature map and the control text are passed through UNet to obtain a sixth intermediate feature map;
[0140] Step 3014: the first intermediate feature map is passed through the first zero convolution layer to output a first transition feature map, the second intermediate feature map is passed through the second zero convolution layer to output a second transition feature map, the third intermediate feature map is passed through the third zero convolution layer to output a third transition feature map, and the fourth intermediate feature map is passed through the fourth zero convolution layer to output a fourth transition feature map;
[0141] Step 3015, the sixth intermediate feature map, the first transition feature map, the second transition feature map, the third transition feature map and the fourth transition feature map are passed through the VAE decoder to obtain a decoded output feature map; wherein, the sixth intermediate feature map is passed through the output convolution layer of the VAE decoder and UNetMidBlock2D is added to obtain a seventh intermediate feature map; the seventh intermediate feature map and the fourth transition feature map are added and input into the first UpDecoderBlock2D for processing to obtain an eighth intermediate feature map; the eighth intermediate feature map and the third transition feature map are added and input into the second UpDecoderBlock2D for processing to obtain a ninth intermediate feature map; the ninth intermediate feature map and the second transition feature map are added and input into the third UpDecoderBlock2D for processing to obtain a tenth intermediate feature map; the tenth intermediate feature map and the first transition feature map are added and input into the fourth UpDecoderBlock2D for processing to obtain an eleventh intermediate feature map;
[0142] Step 3016: record the decoded output feature map corresponding to the foggy image with the text control condition added as the clear image after defogging The decoded output feature map corresponding to the output of the fogged image with the text control condition added is recorded as the synthetic defogging image The decoded output feature map corresponding to the input clear image with the text control condition added is recorded as the output clear image G(y);
[0143] The decoded output feature map corresponding to the defogging image with text control condition added is recorded as the synthetic foggy image The decoded output feature map corresponding to the clear image with text control condition is recorded as the fogged image The decoded output feature map corresponding to the output of the input foggy image with the text control condition added is recorded as the output foggy image F(x).
[0144] In this embodiment, the real foggy image x is processed by the text perception guidance module in the first defogging network module, the real clear image y is processed by the text perception guidance module in the first defogging network module, and the fogged image The methods of the text perception guidance module in the second defogging network module are the same, as follows:
[0145] Step 3021: The real foggy image x, the real clear image y and the fogged image Both are recorded as the first initial image;
[0146] Step 3022: Input the first initial image into the BLIP2 module to obtain a first text title;
[0147] Step 3023: Process the real foggy image in the training data by using the text inversion method to obtain negative prompt features representing fog;
[0148] Step 3024: Delete haze-related words from the first text title and send it to the CLIP model text encoder to obtain text embedding representation features;
[0149] Step 3025, embedding the text representation feature and the negative prompt feature representing fog through the classifier-free guide to obtain the control of this paper;
[0150] Step 3026: The real foggy image x is processed by the text perception guidance module in the first defogging network module to obtain the control image and the real foggy image x as the foggy image with the text control condition added;
[0151] The real clear image y is processed by the text perception guidance module in the first defogging network module to obtain the control image and the real clear image y of this paper, which are recorded as the input clear image with the text control condition added;
[0152] Image after fogging The control and fogged images obtained by the text-aware guidance module in the second defogging network module Denote as fogged image with text control condition added;
[0153] Clear image after dehazing The method of processing the real foggy image x through the text perception guidance module in the first fogging network module, the real foggy image x through the text perception guidance module in the second fogging network module, and the real clear image y through the text perception guidance module in the second fogging network module is the same as follows:
[0154] Step 30A: Defogging the clear image The real foggy image x and the real clear image y are recorded as the second initial image;
[0155] Step 30B, input the second initial image into the BLIP2 module to obtain a second text title;
[0156] Step 30C, adding haze-related words to the second text title and sending it to the CLIP model text encoder to obtain a second text embedding representation feature;
[0157] Step 30D, embedding the second text representation feature through the non-classifier guide to obtain the second text control;
[0158] Step 30E: Defogging the clear image The second text control and the clear image after defogging are obtained by the text perception guidance module in the first fogging network module Denoted as the dehazed image with text control condition added;
[0159] The second text control and the real clear image y obtained by the text perception guidance module in the second fogging network module are recorded as the input clear image with the text control condition added;
[0160] The second text control obtained by passing the real foggy image x through the text perception guidance module in the second fogging network module and the real foggy image x are recorded as the input foggy image with the text control condition added.
[0161] In this embodiment, in step 301, the real foggy image x is passed through the physical perception guidance module in the first defogging network module to obtain a first reconstructed foggy image I phy and the image after adding fog in step 302 After the physical perception guidance module in the second defogging network module, the second reconstructed foggy image I′ is obtained phy The method is the same as the following:
[0162] Step A: Use the dark channel prior defogging DCP algorithm to process the real foggy image x to obtain the first defogging feature map J dcp , the first transmission image t dcp and atmospheric light A;
[0163] The boundary constraint prior BCCR algorithm is used to process the real foggy image x to obtain the second defogging feature map J bccr and the second transmission image t bccr ;
[0164] Step B: Use a computer to call the concatenate function to combine the real foggy image x and the first transmission image t dcp and the second transmission image t bccr After splicing, input the transmission map optimization network to obtain the optimized transmission map t ref ;
[0165] Step C: The real foggy image x and the first defogging feature map J dcp and the second dehazing feature map J bccr After the physical fusion module PFM, the fusion feature map F is obtained fuse ;
[0166] Step D: Fuse the feature map F fuse Input the dehazing backbone network to obtain the optimized dehazing image J ref ;
[0167] Step F: Using the atmospheric scattering model ASM to optimize the defogging image J ref , optimized transmission map t ref and atmospheric light A to obtain the reconstructed foggy image I phy ;
[0168] Step G: According to the method of step A to step F, the fogged image is After processing, the second reconstructed foggy image I′ is obtained. phy .
[0169] In this embodiment, step C, the specific process is as follows:
[0170] Step C01: The first defogging feature map J dcp After the first convolution + ReLU activation function layer, the first feature map f1 is obtained;
[0171] The real foggy image x passes through the first double-layer convolution module to obtain the second feature map f2;
[0172] The second dehazing feature map J bccr After the second convolution + ReLU activation function layer, the third feature map f3 is obtained;
[0173] Step C02, adding the second characteristic map f2 and the first characteristic map f1 to obtain a fourth characteristic map f4, and adding the second characteristic map f2 and the third characteristic map f3 to obtain a fifth characteristic map f5;
[0174] Step C03, the fourth feature map f4 is passed through the third convolution + ReLU activation function layer to obtain the sixth feature map f6;
[0175] The fifth feature map f5 passes through the fourth convolution + ReLU activation function layer to obtain the seventh feature map f7;
[0176] Step C04, performing a Hadamard product operation on the output of the first Sigmoid function layer of the sixth feature map f6 and the fourth feature map f4 to obtain an eighth feature map f8;
[0177] The output of the seventh feature map f7 after the second Sigmoid function layer is subjected to a Hadamard product operation with the fifth feature map f5 to obtain a ninth feature map f9;
[0178] Step C05: Use the concatenate function to concatenate the eighth feature map f8 and the ninth feature map f9 to obtain the tenth feature map f 10 , and the tenth feature map f 10 After the global adaptive pooling layer, the second double-layer convolution module and the Softmax layer, the eleventh feature map f is obtained.11 ;
[0179] Step C06: The eleventh feature map f 11 With the tenth characteristic graph f 10 The features after the Hadamard product operation and the tenth feature map f 10 Add together to get the twelfth feature map f 12 ;
[0180] Step C07: Use the concatenate function to concatenate the first defogging feature map J dcp , the second defogging feature map J bccr and the twelfth characteristic graph f 12 After splicing, it is sent to the gated convolutional network to obtain the fused feature weight f w ;
[0181] Step C08: Use a computer to calculate the Get the fusion feature map F fuse ;in, represents the Hadamard product operation between feature map matrices, represents the addition operation between feature map matrices; f w1 Represents the feature weight f after fusion w The eigenvalue of the first channel in w2 Represents the feature weight f after fusion w The eigenvalue of the second channel in w3 Represents the feature weight after fusion f w The eigenvalue of the third channel in w4 Represents the feature weight after fusion f w The eigenvalue of the fourth channel in; f 12-1 represents the twelfth feature map f 12 Feature maps of the first three channels in f 12-2 represents the twelfth feature map f 12 Feature maps of the last three channels.
[0182] In this embodiment, more than 7,000 real foggy images are collected from the (Real-world Task-Driven Testing Set, RTTS) dataset and the (Unannotated Real-world Hazy Images, URHI) dataset in the foggy image database RESIDE as training data. At the same time, more than 10,000 real clear images are collected from the (Outdoor Training Set, OTS) dataset in the ADE20K database and the foggy image database RESIDE as training data.
[0183] In this embodiment, Figure 2The TAG on the left is the text perception guidance module, and the PAG is the physical perception guidance module.
[0184] In this embodiment, it should be noted that the defogging discriminator and the fogging discriminator have the same structure and both refer to the Vision-aided-gan network in the paper Ensembling Off-the-shelf Models for GAN Training, where cv_type is set to clip.
[0185] In this embodiment, it should be noted that the VAE encoder, UNet and VAE decoder in the defogging backbone network and the fogging backbone network can refer to the conventional stable diffusion model 2.1Turbo version.
[0186] In this embodiment, it should be noted that four zero convolution layers are added between the VAE encoder and the VAE decoder in order to ensure that the generated image retains the basic content and details of the original image, ensuring that the details of the image after defogging or fogging are the same as the original image without distortion.
[0187] In this embodiment, the size of the convolution kernel in the first zero convolution layer, the second zero convolution layer, the third zero convolution layer, and the fourth zero convolution layer is 1×1, the step size is 1, the padding is 0, and the zero initialization strategy is adopted; the number of convolution kernels in the first zero convolution layer is 256, and the number of convolution kernels in the second zero convolution layer, the third zero convolution layer, and the fourth zero convolution layer is 512;
[0188] The number of convolution kernels in the input convolution layer is 128, the size of the convolution kernel is 3×3, the stride is 1, and the padding is 1;
[0189] The number of convolution kernels in the output convolution layer is 512, the size of the convolution kernel is 3×3, the stride is 1, and the padding is 1.
[0190] In this embodiment, it should be noted that when training is performed in step five, the fine-tuned layers in the VAE encoder and the VAE decoder include "conv1", "conv2", "conv_in", "conv_shortcut", "conv", "conv_out", "skip_conv_1", "skip_conv_2", "skip_conv_3", "skip_conv_4", "to_k", "to_q", "to_v", and "to_out.0", their rank is set to 4, and the initialization weight parameter is set to "gaussian". For the UNet part, the fine-tuned layers include "to_k", "to_q", "to_v", "to_out.0", "conv", "conv1", "conv2", "conv_in", "conv_shortcut", "conv_out", "proj_out", "proj_in", "ff.net.2", "ff.net.0.proj", their rank is set to 8, and the initialization weight is set to "gaussian". Among them, "skip_conv_1", "skip_conv_2", "skip_conv_3", "skip_conv_4" are the first zero convolution layer, the second zero convolution layer, the third zero convolution layer and the fourth zero convolution layer, respectively.
[0191] In this embodiment, it should be noted that the BLIP2 module can refer to the BLIP2 module in the paper Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models;
[0192] For the text inversion method, please refer to the paper An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion;
[0193] The CLIP model text encoder can refer to the CLIP model in the paper Learning Transferable Visual Models From Natural Language Supervision;
[0194] For information without classifier guidance, please refer to the paper Classifier-Free Diffusion Guidance.
[0195] In this embodiment, the specific implementation is that the size of the feature map is expressed by the number of channels × length × width, the size of the input image in step 3011 is 3×256×256, the size of the first intermediate feature map is 128×256×256, the size of the second intermediate feature map is 128×128×128, the size of the third intermediate feature map is 256×64×64, the size of the fourth intermediate feature map is 512×32×32, and the size of the fifth intermediate feature map is 512×32×32; the encoded output feature map of the fifth intermediate feature map output by the subsequent module of the VAE encoder is 4×32×32;
[0196] The size of the sixth intermediate feature map is 4×32×32, the size of the seventh intermediate feature map is 512×32×32, the size of the eighth intermediate feature map is 512×64×64, the size of the ninth intermediate feature map is 512×128×128, the size of the tenth intermediate feature map is 256×256×256, the size of the eleventh intermediate feature map is 128×256×256, and the size of the decoded output feature map output by the subsequent module of the VAE decoder of the eleventh intermediate feature map is 3×256×256.
[0197] The size of the first transition feature map is 256×256×256, the size of the second transition feature map is 512×128×128, the size of the third transition feature map is 512×64×64, and the size of the fourth transition feature map is 512×32×32.
[0198] In this embodiment, it should be noted that the transmission map optimization network has the same structure as the transmission map optimization network in the paper RefineDNet: A Weakly Supervised Refinement Framework for Single Image Dehazing, except that the input channel of the present invention is 5.
[0199] In this embodiment, it should be noted that in step D, the fusion feature map F fuse The method for inputting the defogging backbone network refers to the method from step 3012 to step 3017, and there is no need to input the control text.
[0200] In this embodiment, the size of the feature map is expressed by the number of channels × length × width. The size of the real foggy image x is 3 × 256 × 256. The first defogging feature map J dcp and the second dehazing feature map J bccr The size of is 3×256×256; the first transmission image t dcp and the second transmission image t bccr The size of is 1×256×256; the optimized transmission map t ref The size is 1×256×256;
[0201] Optimized dehazed image J ref and reconstruct the foggy image I phy The size is 3×256×256;
[0202] In this embodiment, the sizes of the first feature map f1 to the ninth feature map f9 are all 3×256×256, and the sizes of the tenth feature map f 10 The size of is 6×256×256, and the eleventh feature map f 11 The size of the twelfth feature map f 12 The size of is 6×256×256, and the feature weight after fusion is f w The size of each is 4×256×256, and the fusion feature map F fuse The size is 3×256×256.
[0203] In this embodiment, the convolutions in the first convolution + ReLU activation function layer, the second convolution + ReLU activation function layer, the third convolution + ReLU activation function layer, and the fourth convolution + ReLU activation function layer are point convolutions, the convolution kernel size is 1×1, the number of convolution kernels is 3, the step size is 1, and the padding is 0;
[0204] The first two-layer convolution module includes convolution layer 1, ReLU activation function and convolution layer 2. The convolution kernel size of convolution layer 1 is 1×1, the number of convolution kernels is 6, the stride is 1, and the padding is 0.
[0205] The convolution kernel size of convolution layer 2 is 1×1, the number of convolution kernels is 3, the stride is 1, and the padding is 0;
[0206] The second double-layer convolution module includes convolution layer-1, ReLU activation function and convolution layer-2. The convolution kernel size of convolution layer-1 is 1×1, the number of convolution kernels is 6, the stride is 1, and the padding is 0.
[0207] The convolution kernel size of convolution layer-2 is 1×1, the number of convolution kernels is 6, the stride is 1, and the padding is 0;
[0208] The convolution kernel size in the gated convolutional network is 3×3, the number of convolution kernels is 4, the stride is 1, and the padding is 1.
[0209] In this embodiment, it should be noted that the preset number of iterative training in step 502 is 10.
[0210] In this embodiment, Figure 6The upper row in the middle is a foggy image. By using the method of the present invention, a clear image after defogging in the lower row is obtained. Therefore, the present invention can greatly improve the defogging effect of the foggy image and also has a high defogging quality in the real foggy image.
[0211] In summary, the method of the present invention has simple steps and reasonable design. The image defogging method based on unpaired training can effectively utilize the training data of real scenes, so that the trained defogging model has better generalization ability, and the generated defogging image has better visual quality.
[0212] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any simple modification, change and equivalent structural change made to the above embodiment based on the technical essence of the present invention still falls within the protection scope of the technical solution of the present invention.
Claims
1. An unpaired image dehazing method based on a cyclic diffusion model, characterized in that: The method comprises the following steps: Step 1: Obtaining training set images: Collect real foggy images from the RTTS dataset and the URHI dataset in the foggy image database RESIDE, and collect real clear images from the ADE20K database and the OTS dataset in the foggy image database RESIDE as training data; among them, one real foggy image and one real clear image are used as a set of training data; Step 2: Build a defogging network based on the cyclic diffusion model: The defogging network based on the cyclic diffusion model includes a defogging branch network and a fogging branch network, the defogging branch network includes a first defogging network module, a first fogging network module and a defogging discriminator, and the fogging branch network includes a second fogging network module, a second defogging network module and a fogging discriminator; wherein the first defogging network module, the second defogging network module, the first fogging network module and the second fogging network module have the same structure, and the defogging discriminator and the fogging discriminator have the same structure; The first defogging network module and the second defogging network module both include a defogging backbone network, a text perception guidance module, and a physical perception guidance module; The first fogging network module and the second fogging network module both include a fogging backbone network and a text perception guidance module; the defogging backbone network and the fogging backbone network are both stable cyclic diffusion models; The defogging backbone network and the fogging backbone network have the same structure and both include a VAE encoder, a UNet and a VAE decoder, and four zero convolution layers are added between the VAE encoder and the VAE decoder; Step 3: Feature extraction of real foggy images and real clear images: Step 301: Process the real foggy image x through the text perception guidance module in the first defogging network module to obtain a foggy image with text control conditions added, and pass the foggy image with text control conditions through the defogging backbone network in the first defogging network module to obtain a clear image after defogging. The real foggy image x is processed by the text perception guidance module in the second fogging network module to obtain an input foggy image with text control conditions added, and the input foggy image with text control conditions is processed by the fogging backbone network in the second fogging network module to obtain an output fogged image F(x); The real foggy image x passes through the physical perception guidance module in the first defogging network module to obtain the first reconstructed foggy image I phy ; The clear image after defogging After the text perception guidance module in the first fogging network module, the defogging image with the text control condition is obtained. The defogging image with the text control condition is processed by the fogging backbone network in the first fogging network module to obtain a synthetic foggy image. Step 302: Process the real clear image y through the text perception guidance module in the second fogging network module to obtain a clear image with text control conditions added, and pass the clear image with text control conditions through the fogging backbone network in the second fogging network module to obtain a fogged image. The real clear image y is processed by the text perception guidance module in the first defogging network module to obtain an input clear image with text control conditions added, and the input clear image with text control conditions is passed through the defogging backbone network of the first defogging network module to obtain an output clear image G(y); The fogged image After the text perception guidance module in the second defogging network module, the fogged image with the text control condition is obtained, and the fogged image with the text control condition is processed by the defogging backbone network in the second defogging network module to obtain a synthetic defogging image The fogged image After the physical perception guidance module in the second defogging network module, the second reconstructed foggy image I′ is obtained phy ; Step 4: Establishment of total loss function: Step 401: According to Get the generated adversarial loss L GAN ; Among them, D Y (y) represents the value output by the dehazing discriminator for the real clear image y. Represents a clear image after dehazing The value output by the dehazing discriminator, D X (x) represents the value output by the fogging discriminator for the real foggy image x. Represents the image after adding fog The value output by the fogging discriminator; Step 402: According to Get the cycle consistency loss L cyc ;in, Represents the real foggy image x and the synthetic foggy image The Manhattan distance between Represents the real foggy image x and the synthetic foggy image The LPIPS distance between Represents the real clear image y and the synthetic dehazed image The Manhattan distance between Represents the real clear image y and the synthetic dehazed image The LPIPS distance between them; Step 403: According to L idt =L rec1 (G(y),y)+L rec2 (G(y),y)+L rec1 (F(x),x)+L rec2 (F(x),x), we get the identity loss L idt Among them, L rec1 (G(y), y) represents the Manhattan distance between the real clear image y and the output clear image G(y), L rec2 (G(y),y) represents the LPIPS distance between the real clear image y and the output clear image G(y), L rec1 (F(x),x) represents the Manhattan distance between the real foggy image x and the output fogged image F(x), L rec2 (F(x),x) represents the LPIPS distance between the real foggy image x and the output fogged image F(x); Step 404: According to L phy (1) = L phy +L′ phy , and get the reconstruction loss L phy (1); where L phy represents the first branch reconstruction loss, and L phy =L rec1 (x,I phy )+L rec2 (x,I phy ), L rec1 (x,I phy ) represents the real foggy image x and the first reconstructed foggy image I phy The Manhattan distance between rec2 (x,I phy ) represents the real foggy image x and the first reconstructed foggy image I phy LPIPS distance between p ' hy represents the second branch reconstruction loss, and Represents the image after adding fog and the second reconstructed foggy image I′ phy The Manhattan distance between Represents the image after adding fog and the second reconstructed foggy image I′ phy The LPIPS distance between them; Step 405: According to L loss =L GAN +L cyc +L idt +0.5L phy (1), we get the total loss function L loss ; Step 5: Training of defogging network based on cyclic diffusion model: Step 501: The computer uses the Adam optimization algorithm to input a set of training data and first uses the generated adversarial loss L GAN The fogging discriminator and the defogging discriminator are trained until the generation adversarial loss L is obtained. GAN Fogging discriminator and defogging discriminator at maximum value; Step 502: The computer uses the Adam optimization algorithm to input the set of training data, and uses the total loss function L under the fogging discriminator and defogging discriminator determined in step 501. loss The dehazing network based on the cyclic diffusion model is trained until the total loss function L loss Minimum, complete the training of this set of training data; Step 503, follow the method from step 501 to step 502 until all training sets are trained and one iteration training is completed; Step 504, repeating steps 501 to 503, iterative training until the preset number of iterative training times is met, and a trained defogging network based on the cyclic diffusion model is obtained; Step 6: Use the trained defogging network based on the cyclic diffusion model to defog a single image: A computer is used to input any foggy image into the first defogging network module in the defogging branch network of the trained cyclic diffusion model defogging network to perform defogging processing to obtain a clear image after defogging.
2. The unpaired image defogging method based on a cyclic diffusion model according to claim 1, characterized in that: In step 2, the four zero-convolutional layers added between the VAE encoder and the VAE decoder are the first zero-convolutional layer, the second zero-convolutional layer, the third zero-convolutional layer, and the fourth zero-convolutional layer; The VAE encoder includes an input convolutional layer, a first DownEncoderBlock2D, a second DownEncoderBlock2D, a third DownEncoderBlock2D, and a fourth DownEncoderBlock2D; The VAE decoder includes an output convolutional layer plus UNetMidBlock2D, a first UpDecoderBlock2D, a second UpDecoderBlock2D, a third UpDecoderBlock2D, and a fourth UpDecoderBlock2D; A first zero convolution layer is set between the output of the VAE encoder input convolution layer and the input of the fourth UpDecoderBlock2D, a second zero convolution layer is set between the output of the first DownEncoderBlock2D and the input of the third UpDecoderBlock2D, a third zero convolution layer is set between the output of the second DownEncoderBlock2D and the input of the second UpDecoderBlock2D, and a fourth zero convolution layer is set between the output of the third DownEncoderBlock2D and the input of the first UpDecoderBlock2D.
3. The unpaired image defogging method based on a cyclic diffusion model according to claim 1, characterized in that: The foggy image with the text control condition is passed through the defogging backbone network in the first defogging network module, the fogged image with the text control condition is passed through the defogging backbone network in the second defogging network module, the input clear image with the text control condition is passed through the defogging backbone network of the first defogging network module, the defogging image with the text control condition is passed through the fogging backbone network in the first fogging network module, the clear image with the text control condition is passed through the fogging backbone network in the second fogging network module, and the input foggy image with the text control condition is passed through the fogging backbone network in the second fogging network module. The processing method is the same, as follows: Step 3011: the foggy image with text control condition added, the fogged image with text control condition added, the input clear image with text control condition added, the defogged image with text control condition added, the clear image with text control condition added, and the input foggy image with text control condition added are recorded as input image and control text; Step 3012: Pass the input image through the VAE encoder to obtain an encoded output feature map; wherein the input image is input into the input convolution layer of the VAE encoder to obtain a first intermediate feature map, and the first intermediate feature map is sequentially passed through the first DownEncoderBlock2D, the second DownEncoderBlock2D, the third DownEncoderBlock2D and the fourth DownEncoderBlock2D to obtain a second intermediate feature map, a third intermediate feature map, a fourth intermediate feature map and a fifth intermediate feature map; Step 3013: The encoded output feature map and the control text are passed through UNet to obtain a sixth intermediate feature map; Step 3014: the first intermediate feature map is passed through the first zero convolution layer to output a first transition feature map, the second intermediate feature map is passed through the second zero convolution layer to output a second transition feature map, the third intermediate feature map is passed through the third zero convolution layer to output a third transition feature map, and the fourth intermediate feature map is passed through the fourth zero convolution layer to output a fourth transition feature map; Step 3015, the sixth intermediate feature map, the first transition feature map, the second transition feature map, the third transition feature map and the fourth transition feature map are passed through the VAE decoder to obtain a decoded output feature map; wherein, the sixth intermediate feature map is passed through the output convolution layer of the VAE decoder and UNetMidBlock2D is added to obtain a seventh intermediate feature map; the seventh intermediate feature map and the fourth transition feature map are added and input into the first UpDecoderBlock2D for processing to obtain an eighth intermediate feature map; the eighth intermediate feature map and the third transition feature map are added and input into the second UpDecoderBlock2D for processing to obtain a ninth intermediate feature map; the ninth intermediate feature map and the second transition feature map are added and input into the third UpDecoderBlock2D for processing to obtain a tenth intermediate feature map; the tenth intermediate feature map and the first transition feature map are added and input into the fourth UpDecoderBlock2D for processing to obtain an eleventh intermediate feature map; Step 3016: record the decoded output feature map corresponding to the foggy image with the text control condition added as the clear image after defogging The decoded output feature map corresponding to the output of the fogged image with the text control condition added is recorded as the synthetic defogging image The decoded output feature map corresponding to the input clear image with the text control condition added is recorded as the output clear image G(y); The decoded output feature map corresponding to the defogging image with text control condition added is recorded as the synthetic foggy image The decoded output feature map corresponding to the clear image with text control condition is recorded as the fogged image The decoded output feature map corresponding to the output of the input foggy image with the text control condition added is recorded as the output foggy image F(x).
4. The unpaired image defogging method based on a cyclic diffusion model according to claim 1, characterized in that: The real foggy image x is processed by the text perception guidance module in the first defogging network module, the real clear image y is processed by the text perception guidance module in the first defogging network module, and the fogged image The methods of the text perception guidance module in the second defogging network module are the same, as follows: Step 3021: The real foggy image x, the real clear image y and the fogged image Both are recorded as the first initial image; Step 3022: Input the first initial image into the BLIP2 module to obtain a first text title; Step 3023: Process the real foggy image in the training data by using the text inversion method to obtain negative prompt features representing fog; Step 3024: Delete haze-related words from the first text title and send it to the CLIP model text encoder to obtain text embedding representation features; Step 3025, embedding the text representation feature and the negative prompt feature representing fog through the classifier-free guide to obtain the control of this paper; Step 3026: The real foggy image x is processed by the text perception guidance module in the first defogging network module to obtain the control image and the real foggy image x as the foggy image with the text control condition added; The real clear image y is processed by the text perception guidance module in the first defogging network module to obtain the control image and the real clear image y of this paper, which are recorded as the input clear image with the text control condition added; Image after fogging The control and fogged images obtained by the text-aware guidance module in the second defogging network module Denote as fogged image with text control condition added; Clear image after dehazing The method of processing the real foggy image x through the text perception guidance module in the first fogging network module, the real foggy image x through the text perception guidance module in the second fogging network module, and the real clear image y through the text perception guidance module in the second fogging network module is the same as follows: Step 30A: Defogging the clear image The real foggy image x and the real clear image y are recorded as the second initial image; Step 30B, input the second initial image into the BLIP2 module to obtain a second text title; Step 30C, adding haze-related words to the second text title and sending it to the CLIP model text encoder to obtain a second text embedding representation feature; Step 30D, embedding the second text representation feature through the non-classifier guide to obtain the second text control; Step 30E: Defogging the clear image The second text control and the clear image after defogging are obtained by the text perception guidance module in the first fogging network module Denoted as the dehazed image with text control condition added; The second text control and the real clear image y obtained by the text perception guidance module in the second fogging network module are recorded as the input clear image with the text control condition added; The second text control obtained by passing the real foggy image x through the text perception guidance module in the second fogging network module and the real foggy image x are recorded as the input foggy image with the text control condition added.
5. The unpaired image defogging method based on a cyclic diffusion model according to claim 1, characterized in that: In step 301, the real foggy image x is passed through the physical perception guidance module in the first defogging network module to obtain the first reconstructed foggy image I phy and the image after adding fog in step 302 After the physical perception guidance module in the second defogging network module, the second reconstructed foggy image I′ is obtained phy The method is the same as the following: Step A: Use the dark channel prior defogging DCP algorithm to process the real foggy image x to obtain the first defogging feature map J dcp , the first transmission image t dcp and atmospheric light A; The boundary constraint prior BCCR algorithm is used to process the real foggy image x to obtain the second defogging feature map J bccr and the second transmission image t bccr ; Step B: Use a computer to call the concatenate function to combine the real foggy image x and the first transmission image t dcp and the second transmission image t bccr After splicing, input the transmission map optimization network to obtain the optimized transmission map t ref ; Step C: The real foggy image x and the first defogging feature map J dcp and the second dehazing feature map J bccr After the physical fusion module PFM, the fusion feature map F is obtained fuse ; Step D: Fuse the feature map F fuse Input the dehazing backbone network to obtain the optimized dehazing image J ref ; Step F: Using the atmospheric scattering model ASM to optimize the defogging image J ref , optimized transmission map t ref and atmospheric light A to obtain the reconstructed foggy image I phy ; Step G: According to the method of step A to step F, the fogged image After processing, the second reconstructed foggy image I′ is obtained phy .
6. The unpaired image defogging method based on a cyclic diffusion model according to claim 5, characterized in that: Step C, the specific process is as follows: Step C01: The first defogging feature map J dcp After the first convolution + ReLU activation function layer, the first feature map f1 is obtained; The real foggy image x passes through the first double-layer convolution module to obtain the second feature map f2; The second dehazing feature map J bccr After the second convolution + ReLU activation function layer, the third feature map f3 is obtained; Step C02, adding the second characteristic map f2 to the first characteristic map f1 to obtain a fourth characteristic map f4, and adding the second characteristic map f2 to the third characteristic map f3 to obtain a fifth characteristic map f5; Step C03, the fourth feature map f4 is passed through the third convolution + ReLU activation function layer to obtain the sixth feature map f6; The fifth feature map f5 passes through the fourth convolution + ReLU activation function layer to obtain the seventh feature map f7; Step C04, performing a Hadamard product operation on the output of the first Sigmoid function layer of the sixth feature map f6 and the fourth feature map f4 to obtain an eighth feature map f8; The output of the seventh feature map f7 after the second Sigmoid function layer is subjected to a Hadamard product operation with the fifth feature map f5 to obtain a ninth feature map f9; Step C05: Use the concatenate function to concatenate the eighth feature map f8 and the ninth feature map f9 to obtain the tenth feature map f 10 , and the tenth feature map f 10 After the global adaptive pooling layer, the second double-layer convolution module and the Softmax layer, the eleventh feature map f is obtained. 11 ; Step C06: The eleventh feature map f 11 With the tenth characteristic graph f 10 The features after the Hadamard product operation and the tenth feature map f 10 Add together to get the twelfth feature map f 12 ; Step C07: Use the concatenate function to concatenate the first defogging feature map J dcp , the second defogging feature map J bccr and the twelfth characteristic graph f 12 After splicing, it is sent to the gated convolutional network to obtain the fused feature weight f w ; Step C08: Use a computer to calculate the Get the fusion feature map F fuse ;in, represents the Hadamard product operation between feature map matrices, represents the addition operation between feature map matrices; f w1 Represents the feature weight after fusion f w The eigenvalue of the first channel in, f w2 Represents the feature weight f after fusion w The eigenvalue of the second channel in w3 Represents the feature weight f after fusion w The eigenvalue of the third channel in w4 Represents the feature weight after fusion f w The eigenvalue of the fourth channel in; f 12-1 represents the twelfth feature map f 12 Feature maps of the first three channels in f 12-2 represents the twelfth feature map f 12 Feature maps of the last three channels.
Citation Information
Patent Citations
Image defogging method and system based on edge attention and multi-order differential loss
CN116703750A
Mask segmentation map guided Chinese landscape painting generation model construction method
CN118262195A
Information filling defogging method and system for fog-containing image
CN118279187A
Image dehazing method and system based on cyclegan
US20220414838A1
Cited By
Non-paired low-light real image enhancement method based on multi-mode guidance
CN121860873A