An AI Smoke Removal Method for Real Fog Scenes

By introducing an unaligned supervised learning framework and newly defined atmospheric light and transmission medium diagrams into the defog algorithm, the problem of the existing defog algorithm being poor in real scenes is solved, achieving better defog removal effects and defog removal results that are closer to real scenes.

CN114913093BActive Publication Date: 2025-06-13NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210565008.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-23
Publication Date
2025-06-13
Estimated Expiration
2042-05-23

AI Technical Summary

Technical Problem

The existing fog removal algorithm is not effective in real scenes, making it difficult to completely remove fog, and the restored scene map has problems such as color distortion and artifacts.

Method used

A non-aligned supervised defog learning framework is proposed. By redefining the mean-variance description method of atmospheric light and the three-channel method of transmission medium diagram, the corresponding neural network structure is designed, and the atmospheric light and transmission medium diagram are learned to effectively train the defog network.

Benefits of technology

A better smoke removal effect is achieved in real scenarios, reducing the need for data alignment, allowing the model to be trained in real scenarios, and avoiding the bad effects caused by inconsistency between synthetic data training and real data testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114913093B_ABST
    Figure CN114913093B_ABST
Patent Text Reader

Abstract

The present invention discloses an AI de-smoking method for real scenes, providing a clearer vision for real smoke scenes. This method uses unaligned clear images to design a loss function to supervise the training of the de-smoking network, and redefines the mean-variance description method of atmospheric light and the three-channel method of the transmission medium map, and proposes a corresponding neural network structure to better learn the atmospheric light and the transmission medium map, so that the reconstruction loss function can effectively train the de-smoking network. The present invention reduces the strict alignment requirement of traditional supervised models for data, enables the model to complete training in real scenes, and avoids the bad effects caused by the inconsistency between synthetic data training and real data testing; redefines the atmospheric light and the transmission medium map, making it more in line with real scenes, thereby improving the de-smoking effect in real scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image restoration and scene reconstruction, and particularly to an image dehazing algorithm under real scenes. Background Art

[0002] All along, adverse weather conditions have been an important cause of traffic accidents. Especially in foggy days, due to the limited visibility in foggy days, it is extremely easy to cause vehicle collisions and rear-end collisions. Therefore, image dehazing is also an important task in machine vision research. Currently, both in-vehicle driving systems and driverless assistance systems are very interested in the "dehazing" technology. Although the task of dehazing has been studied in the visual field for decades, there are still many problems with existing dehazing methods. For example, it is difficult to completely remove the fog, and problems such as color distortion and artifacts appear in the restored clear scene images. The existence of these problems makes it difficult for dehazing algorithms to be truly applied in driving systems. Existing dehazing algorithms are mainly divided into two categories: one is dehazing based on physical prior knowledge (statistical attributes of fog itself), and the other is dehazing based on learning methods (such as deep learning, adversarial learning, etc.).

[0003] In terms of dehazing based on physical prior knowledge, these methods mainly use the statistical features or assumptions of foggy images for dehazing. The most influential dehazing method based on dark channel prior is to improve the accuracy of estimating the atmospheric light A ∞ and the transmission medium map T by statistically analyzing the physical properties of foggy images, so as to obtain a high-quality dehazing effect. However, this method still has defects. For example, artifacts are likely to appear in the sky area, and the color of the dehazed scene is dark. Similarly, there are still the above problems with many prior-based dehazing methods. These design model ideas based on physical priors are limited to dehazing problems in scenarios that conform to prior assumptions, and it is almost impossible to achieve full coverage.

[0004] In terms of dehazing based on learning methods, it is roughly divided into three categories of methods. The first category is to use a convolutional network to learn the atmospheric light A ∞ and the transmission medium map T from foggy images, and then obtain J (i.e., the clear image) according to the atmospheric scattering model. The second category is to directly train an end-to-end network framework to directly map from foggy images to clear images. The third category is to directly generate clear images through adversarial learning. These methods rely heavily on aligned fog / clear sample pairs of data. However, it is very difficult to obtain aligned clear images in real scenes. Therefore, these methods can only train models on synthetic datasets. However, the models trained on synthetic datasets cannot achieve satisfactory results in real scenes. In addition, these methods still use the dark channel prior method to define the atmospheric light A ∞and the transmission medium image T. This definition is too idealistic for atmospheric light, completely ignoring the influence of the wavelength of light, not being close enough to the real scene, and it can hardly distinguish white objects and fog regions for a single-channel transmission medium image. Summary of the Invention

[0005] The object of the present invention is to provide an AI de-smoking method for real fog scenes. On the one hand, a loss function is designed using unaligned clear images to supervise the training of the de-smoking network; on the other hand, the mean-variance description method of atmospheric light and the three-channel method of the transmission medium image are redefined, and a corresponding neural network structure is proposed to better learn atmospheric light and the transmission medium image, so that the reconstruction loss function effectively trains the de-smoking network.

[0006] The technical solution to implement the present invention is as follows: In the first aspect, the present invention provides an AI de-smoking method for real scenes, including:

[0007] Step 1: Input the smoke image I, and initialize to obtain a preliminary de-smoked image J DCP and the dark channel image D, and input J DCP into the de-smoking network to generate a de-smoking result J;

[0008] Step 2: Input the smoke image I and the dark channel image D into a convolutional network to extract the spatial convolutional feature map F I and F D , calculate the self-attention feature F I and F D of F using the self-attention mechanism, and screen the mean value of the 1% brightest pixel points in each channel as the relative mean feature A att , then take the difference between F m and F I as the relative difference feature A att , and finally calculate the atmospheric light A d of the real scene; ∞ ;

[0009] Step 3: Input the smoke image I into a convolutional network with channel attention to generate a rough three-channel transmission medium map and then use the guided filtering module to calculate a more accurate and refined transmission medium map T;

[0010] Step 4: According to J, A ∞ and T generated in Steps 1-3, calculate the reconstructed smoke image I rec using the atmospheric scattering model, and then define the reconstruction loss function of I rec and I, which is the sum of three loss functions, namely the first-order norm loss Perceptual loss and structural similarity loss wherein VGG16 is used to calculate the perceptual content;

[0011] Step 5: Use the multi-scale non-aligned reference loss to perform supervised learning on the color and texture of the de-smoking network;

[0012] Step 6: According to the two loss functions in Step 4 and Step 5 and optimize the network parameters of the entire NSDNet framework to obtain the de-smoking network; finally, input the test real RGB smoke image I t , and generate the clear scene image result J according to the process in Step 1 t .

[0013] In one embodiment, in Step 1, the dark channel prior method is used to calculate the preliminary de-smoking image J DCP and the dark channel image D, which respectively correspond to the inputs of the de-smoking network and the atmospheric light network, and generate the de-smoking result J

[0014] In one embodiment, in Step 2, the convolutional network is used to extract the features F I and F D of the smoke image I and the dark channel image D respectively, calculate their self-attention features F att , and screen the mean value of the 1% brightest pixel points in each channel as the relative mean feature A m , then take the difference between F I and F att as the relative difference feature A d , and finally calculate the atmospheric light A of the real scene through linear combination ∞ =αA m +βA d , and A ∞ is non-uniform

[0015] In one embodiment, in Step 4, the predicted J, A ∞ and T are brought into the atmospheric scattering model I(x)=J(x)t(x)+A ∞ (1 - t(x)), and the reconstructed smoke image I rec is calculated. The reconstruction loss function performs unsupervised learning on the texture and color of I rec , wherein VGG16 is used to calculate the perceptual content

[0016] In one of the embodiments, the multi-scale misaligned reference loss in step 5 Perform supervised learning on the de-smoking network for color and texture, where Use the VGG16 network to calculate the multi-scale content loss between J and J ref and Use the discriminator of patch-GAN to calculate the multi-scale adversarial loss between J and J ref between.

[0017] In a second aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in the first aspect are implemented.

[0018] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0019] In a fourth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0020] Compared with the prior art, the present invention has the following remarkable advantages: (1) The present invention proposes a misaligned supervised de-hazing learning framework, which reduces the strict alignment requirement of traditional supervised models for data, enables the model to complete training in real scenarios, and avoids the adverse effects caused by the inconsistency between synthetic data training and real data testing. (2) The present invention redefines the atmospheric light and transmission medium map, making it more in line with real scenarios, thereby improving the de-smoking effect in real scenarios.

[0021] The following further describes the present invention in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is the network architecture diagram of the method of the present invention.

[0023] Figure 2 is the channel attention structure diagram.

[0024] Figure 3 are the de-smoking (left) and de-hazing (right) effects of the method of the present invention in real scenarios.

[0025] Figure 4 are the visualization effect diagrams of the method of the present invention for different variables, including the predicted atmospheric light A ∞ , the transmission medium map T, the redefined A ∞ and T, and the reconstructed clear scene image J.

[0026] Figure 5 This is the comparison of the effects of the method of the present invention with other methods in a real smoke scenario.

[0027] Figure 6 This is the comparison of the effects of the method of the present invention with other methods in a real fog scenario.

[0028] Figure 7 This is the comparison of the effects of the method of the present invention with other methods on real fog scenarios collected from a third-party network. Detailed implementation manners

[0029] Existing de-fogging methods mostly rely on aligned smoke / clear sample pairs of data. However, it is very difficult to obtain aligned clear images in real scenarios. Therefore, these methods can only train models on synthetic data sets. However, the models trained on synthetic data sets cannot achieve satisfactory effects in real scenarios. In addition, these methods still use the dark channel prior method to define the atmospheric light and the transmission medium image. This definition is not only too idealistic for the atmospheric light, completely ignoring the influence of the wavelength of light and not being close enough to the real scenario, but also can hardly distinguish white objects from fog regions for the single-channel transmission medium map.

[0030] In view of this, the present invention proposes to train a model supervised by non-aligned clear images in the same scenario, thereby reducing the data acquisition requirements for real scenarios and at the same time changing the definition of the atmospheric light A ∞ and the transmission medium map T, and designing corresponding network modules to better learn A ∞With respect to T, ultimately, the method of the present invention can complete training in a real scenario and achieve a better de-smoking effect. For example: driving scenarios and mobile phone photography in smoky weather, etc. This method is named NSDNet (Non-aligned Supervised De-smoking Network), which includes three sub-networks, namely, a de-smoking network, an atmospheric light network, and a transmission medium map network. The core of the present invention includes three parts: a non-aligned supervision framework, a mean-variance simulation of atmospheric light, and a three-channel transmission medium map. In the non-aligned supervision framework, we introduce non-aligned and clear reference images to supervise the training of the de-smoking network, and propose a multi-scale non-aligned reference loss function, which includes multi-scale adversarial and content loss functions. This framework reduces the strict requirements for data in traditional methods (fully aligned smoke / clear image sample pairs), enabling the present invention to easily collect a large amount of data to train the model in a real smoke scenario, thereby getting rid of the limitation that the current work can only be trained on synthetic smoke data and the defect of poor de-smoking effect. In the mean-variance simulation of atmospheric light, we propose a self-attention mechanism for mean and variance to discover the smoke areas in the image, and construct a novel network structure to adaptively learn more realistic atmospheric light. In addition, we modify the traditional single-channel transmission medium map to a three-channel transmission medium map to make it closer to the real scenario. In summary, the method of the present invention directly extends the de-smoking technology to training on real smoke data for the first time and quickly repairs smoke images. The de-smoking effect of NSDNet in a real scenario is shown in the accompanying drawings and reaches the current best effect.

[0031] As Figure 1 shown, the present invention provides a calculation process of an AI de-smoking method for a real scenario. First, a smoky scene image is given, and the initial de-smoking result and the dark channel map are obtained by using the Dark Channel Prior (DCP) algorithm. Then, this initial de-smoking result is input into a de-smoking network to generate a clear scene image J. Secondly, we use a shared convolutional network to extract the features of the RGB smoky image and the corresponding dark channel map respectively, and then perform self-attention operations on these two features to obtain self-attention features (the purpose is to highlight the features of the smoke area), and select the mean value of the 1% brightest pixel points from the self-attention features as the relative mean A m , take the difference between the smoky image features and the self-attention features as the relative difference A d , and use linear combination to approximate the atmospheric light A ∞ = αA m + βA d (A ∞ is non-uniform, and α and β are combination coefficients). Then, we use a channel attention network to predict a three-channel rough transmission medium map And perform guided filtering operation on it to obtain a more refined and accurate transmission medium map T. Finally, based on the A predicted by the three sub-networks respectively in the present invention ∞ , T and J, reconstruct the smoke or fog image I according to the atmospheric scattering model rec , and establish a supervised training signal for the de-smoke network through the reconstruction loss . More importantly, the model of the present invention also establishes a multi-scale misaligned reference loss function regarding J and J ref as a further supervised training signal for the de-smoke network, and jointly and perform network parameter optimization to obtain high-quality de-smoke images (in terms of color, texture, and brightness). Note that the size representation of the image is for the case where the batch size is 1. The input sample pair of this method is the real RGB smoke image I and the misaligned and clear reference image J , and the specific steps are as follows: ref

[0032] Step 1, as Figure 1 shown, given the image I ∈ R 3×H×W (where H and W are the height and width of the image), use the existing Dark Channel Prior (DCP) algorithm to remove fog and obtain an initial de-smoke result J DCP ∈ R 3×H×W and the dark channel map D ∈ R 1×H×W , and then input J DCP into the de-smoke generation network, which is in the form of encoder + resblock + decoder structure, where there are 9 resblocks (residual blocks). Through the DCP algorithm and the de-smoke generation network, the model of the present invention can generate a clear scene image J ∈ R 3×H×W .

[0033] Step 2, as Figure 1 shown, for the smoke image I and the dark channel map D obtained in Step 1, use the shared network (basic Unet structure) to extract the feature maps F I and F D ∈ R 64×256×256 , and then perform self-attention operation on the feature maps F I and F D . Here, let F D be the K of the self-attention operation, and F I be the Q and V of the self-attention operation. Perform convolution operations on K, Q, and V respectively to adjust the channels to K D , Q I and V I ∈ R 8×256×256 , and then perform operations on K​D and V I Perform 4×4 downsampling respectively to reduce the computational complexity. Then K D and Q I are matrix-multiplied and then enter the softmax function to obtain the attention weight W att ∈R 65536×4096 . Then, multiply W att by V I to obtain the attention feature map F att ∈R 8×256×256 . Finally, select the mean of the brightest 1% pixel points from F att as the relative mean A m . At the same time, take the absolute value of the difference between F att and V I as the relative difference A d . Then, perform a weighted linear combination on A m and A d to approximate the A ∞ ∈R 3×256×256 (non-uniform) of the real scene. This can better predict an A ∞ that is closer to the real scene. The specific formula description is as follows:

[0034] A ∞ =αA m +βA d

[0035] A d =|V I -F att |,

[0036]

[0037] K D =C 1×1 (F D ), Q I =C 1×1 (F I ), V I =C 1×1 (F I ),

[0038] where C 1×1 represents the convolution operation with a kernel size of 1×1, represents matrix multiplication, represents the operation of selecting the brightest 1% pixel points.

[0039] Step 3, as Figure 1As shown, the RGB smoke or fog image is input into a channel attention Unet network (Note: Here, in the Unet network, the downsampled features in the encoder are concatenated with the corresponding scale features in the upsampling along the channel dimension after passing through the channel attention and then input into the next layer), and a rough three-channel is predicted. Then, guided filtering (filtering radius = 40, parameter ∈ = 0.001) is used for fine-tuning to obtain T ∈ R 3×H×W , making the predicted T have clear boundaries and sharp shapes. Among them, the structure of the channel attention network is as Figure 2 shown.

[0040] Step 4: As Figure 1 shown, in Step 4, the predicted J, A ∞ and T in the above steps are substituted into the atmospheric scattering model I(x) = J(x)t(x) + A ∞ (1 - t(x)) to obtain the reconstructed smoke image I rec . Here, three loss functions ( and ) are used to constrain the texture and color of the reconstructed I rec to obtain a realistic reconstruction effect. The content loss is calculated using VGG16 (selecting the feature calculation of the content loss after the Relu layer, specific layers: 3, 8, 15, 22, 29). Note: The representation of the loss formula is for a pair of samples.

[0041]

[0042]

[0043] Among them, the parameters θ, λ, and η are default set to 1 to maintain the stable variable C 1 =(k 1 h) 2 , C 2 =(k 2 h) 2 . The parameter h is the pixel dynamic range value, and k 1 and k 2 are default set to 0.01 and 0.03. Φ l (I rec ) and Φ l (I) respectively represent the feature maps corresponding to I rec and I in the l-th layer extracted by the VGG16 network. N represents the number of extracted feature layers. and represent the mean value of the image I rec . and represents the variance of image I, represents image I rec and the covariance of I.

[0044] Step 5, as Figure 1 shown, the multi-scale non-aligned reference loss proposed by the present invention constrains the color and texture of the generated J to be closer to the reference image, and the content loss between the generated clear scene graph J and the non-aligned reference graph J is also calculated using the VGG16 network. ref The discriminator structure based on patch-GAN adds multi-scale discrimination to calculate the adversarial loss. The discriminator of this adversarial generation network consists of 5 layers of Conv+BatchNorm2d+leakyRelu, where the first layer does not contain BatchNorm2d.

[0045]

[0046] Among them, J ref is the non-aligned reference image, and J is the clear scene graph generated by the model of the present invention, represents J ref at different scales (i = 1, 2, 3 represent 0.5, 1, 1.5 times the scale of J ref respectively), J i represents J at different scales (i = 1, 2, 3 represent 0.5, 1, 1.5 times the scale of J), and S is the content similarity between image features. represents taking the expectation based on the sampled samples, D() represents the discriminator network of patch-GAN, log() represents taking the logarithm to the base 10, and Φ l (J) and Φ l (J ref ) respectively represent the feature maps corresponding to the l-th layer of the images J and J ref extracted by the VGG16 network. ω 1 and ω 2 are respectively the weights for balancing and the two losses.

[0047] Step 6, as Figure 1 shown, the network parameters of the entire NSDNet framework are optimized (model training) according to the loss constraints in Steps 4 and 5 to obtain the trained clear scene graph J t generation network.

[0048] ​The present invention constrains the generation of clear scene images by establishing a multi-scale misaligned reference loss function for the same scene, enabling the model to be trained in real smoke scenes and achieving the best current de-smoking effect, thus avoiding the adverse effects caused by the inconsistency between synthetic data training and real data testing. Secondly, this method redefines the atmospheric light and transmission medium map to make it more conform to the real scene, and improves the constraint accuracy of the de-smoking network through the reconstruction loss.

[0049] The present invention compares several existing state-of-the-art (SOTA) dehazing methods, namely DCP, Cycle-Dehaze, DAD, RefineNet, and PSD. DCP uses the statistical properties of fog as prior knowledge for dehazing, that is, in the RGB three channels of the fog region of a foggy image, there is always one channel value tending to 0.

[0050] Cycle-Dehaze uses CycleGAN and perceptual loss to perform adversarial generation for a dehazing method. DAD uses domain adaptation to reduce the domain difference between synthetic data and real scenes for dehazing. RefineNet is based on unpaired scenes and uses a network to learn the atmospheric scattering model parameters J and T, and then performs perceptual fusion on the obtained reconstructed J rec and the learned J and the fog image. PSD is a dehazing framework that fuses various physical prior constraints, and the dehazing network can be of any structure.

[0051] In addition, we also compare several SOTA methods on synthetic datasets, namely MSBDN, FFA, UHD, IPUDN, and these methods are trained based on groundtruth, that is, aligned references.

[0052] Table 1. Quantitative comparison of different dehazing algorithms on two real scene datasets

[0053] Defogging method Real smoke dataset Real fog dataset DCP 4.1303 7.4029 Cycle-Dehaze 3.8018 5.1236 DAD 5.5886 5.5886 RefineNet 4.1840 7.0540 PSD 4.4485 6.5610 Ours 3.7686 4.9803

[0054] Table 1 is a quantitative comparison of the de-smoking results of the present invention and existing SOTA methods in real smoke scenes. The evaluation index is NIQE (Natural Image Quality Evaluation), which can be used to evaluate the naturalness of an image, and the lower the index value, the better. The data in the table show that the method of the present invention can achieve the best de-smoking effect, fully demonstrating the effectiveness of our model in de-smoking in real scenes.

[0055] Table 2. Ablation experiments on ALA and on the real fog dataset, ↓ indicates that the lower the index value, the better.

[0056]

[0057]

[0058] Table 2 is for verifying the proposed Atmospheric Light Attention (ALA) and multi-scale reference loss of the present invention We conducted ablation experiments on the method. We designed three methods, namely baseline, baseline + ALA, and The baseline method here only includes: the transmission medium map T prediction network and the clear scene map J generation network, and the atmospheric light A (uniform) is obtained by using DCP. From the results in Table 2, we can find that when we separately add our designed ALA module and loss to the baseline, the quantitative metrics FADE and NIQE of the method gradually decrease. Therefore, the ablation experiment shows the effectiveness of ALA and for dehazing in real scenes.

[0059] Table 3. Ablation experiment on the scale of on the real fog dataset. ↓ indicates that the lower the metric value, the better.

[0060]

[0061] Table 3 is for verifying the effectiveness of the multi-scale in

[0062] Table 4. Ablation experiment on the misaligned scale on the real smoke dataset. ↓ indicates that the lower the metric value, the better.

[0063] Non-alignedscale 30 pixels 60 pixels 90 pixels 120 pixels FADE↓ 0.3031 0.3225 0.3107 0.3438 NIQE↓ 3.7686 3.8819 3.9891 4.1252

[0064] Table 4 is to explore the influence of the misaligned scale on the performance of the method of the present invention. We conducted ablation experiments on four different misaligned scales (30 pixels, 60 pixels, 90 pixels, and 120 pixels) on the real smoke dataset. The experimental results show that the more aligned the reference map is with the input foggy image, the better the dehazing effect.

[0065] Figure 3 The dehazing effect diagrams of the method of the present invention in real smoke scenes are shown. The left figure shows the de-smoking effects under uniform, non-uniform, and heavy smoke conditions; the right figure shows the de-fogging effects under light fog and heavy fog conditions. This figure fully demonstrates the effectiveness of the method of the present invention in dehazing real smoke scenes.

[0066] Figure 4To demonstrate the effectiveness of the method of the present invention by showing the intermediate process of the model, taking an image (a) in a real non-uniform smoke scene as an example, (b) and (c) are A learned by our network ∞ (non-uniform) and T (three channels), while (f) and (g) are A of the dark channel dehazing operation ∞ (uniform) and T (single channel), and then we obtain the clear scene image J according to the atmospheric scattering model I(x) = J(x)t(x) + A ∞ (1 - t(x)). From the figure, we can find that the reconstructed J by the method of the present invention has a better effect. Therefore, through the display of this intermediate process, it shows that A predicted by the method of the present invention ∞ and T are more accurate, thus verifying that our definitions of A ∞ and T are effective. In addition, the self-attention mechanism proposed by the present invention has the effect shown in (e). We can see that the attention area just covers the thick smoke area, indicating the effectiveness of the design of the atmospheric light network module

[0067] Figure 5 Shows the de-smoking visualization comparison between the method of the present invention and several current SOTA methods in real smoke (two states of uniform and non-uniform) scenes. We can find that the de-smoking effect of the method of the present invention is closer to the reference real scene (including color, scene brightness and texture details). In contrast, the SOTA methods cannot remove smoke well. For example, although the DCP and RefineNet methods can remove some smoke, they reduce the brightness of the scene and still retain a small part of the smoke, and the texture details of the restored clear scene image are not perfect enough. While DAD and PSD can basically not complete the de-smoking task in real scenes. In addition, Cycle-Dehaze is a structure model based on CycleGAN. Due to the existence of two different distributions of smoke in the smoke dataset, there is domain inconsistency, resulting in a large number of artifacts. Therefore, through the comparison of the experimental effects, it fully shows the de-smoking effectiveness of the model of the present invention in real smoke scenes

[0068] Figure 6 Shows the de-hazing visualization comparison between the method of the present invention and several current SOTA methods in real fog scenes. Similarly, we can find that the method of the present invention is also effective in real fog scenes. Different from smoke, in real fog scenes, the farther the distance, the thicker the fog (denser), so it is easier to remove near-view fog and more difficult to remove far-view fog. Therefore, compared with other SOTA methods, the model of the present invention can well remove far-view fog, while the comparison methods have poor effects on removing far-view fog. This fully shows the de-hazing effectiveness of the method of the present invention in real fog scenes

[0069] To verify the generalization performance of the method of the present invention, we use a third-party dataset RTTS (real fog images collected from the Internet) to verify the defogging effects of the method of the present invention and the current existing SOTA models in real scenarios, as Figure 7 shown. It can be easily found from the figure that our method is still very effective in thick fog scenarios and can well remove the fog in the distance and restore the target objects in the scene.

[0070] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An AI smoke removal method for real scenarios, characterized in that, comprising: Step 1: Input the smoke image I, and initialize the preliminary de-smoked image J using the dark channel prior method DCP and the dark channel image D, and input J DCP into the de-smoking network to generate the de-smoking result J; Step 2: Input the smoke image I and the dark channel image D into the convolutional network to extract the spatial convolutional feature map F I and F D , use the self-attention mechanism to calculate F I and F D 's self-attention feature F att , and screen the mean value of the 1% brightest pixel points in each channel as the relative mean feature A m , then take the difference between F I and F att as the relative difference feature A d , and finally calculate the atmospheric light A of the real scene ∞ ; Step 3: Input the smoke image I into the convolutional network with channel attention to generate a rough three-channel transmission medium map Then use the guided filtering module to calculate a more accurate and refined transmission medium map T; Step 4. Calculate and reconstruct the smoke image I using the atmospheric scattering model based on J, A ∞ and T generated in Steps 1 - 3, and then define the reconstruction loss function of I rec and I rec which consists of the sum of three loss functions, namely the first-order norm loss the perceptual loss and the structural similarity loss where the perceptual content is calculated using VGG16; ​ Step 5: Use a multi-scale misaligned reference loss Perform supervised learning on the de-smoking network for color and texture; Step 6: Based on the two loss functions in Step 4 and Step 5 and optimize the network parameters of the entire NSDNet framework to obtain a de-smoking network; finally, input the test real RGB smoke image I t and generate a clear scene image result J according to the process in Step 1 t .

2. The AI smoke removal method for real scenarios according to claim 1, characterized in that, In step 1, the initial dehazed image J is calculated using the dark channel prior method DCP and the dark channel image D, which are the inputs to the dehazing network and the atmospheric light network respectively.

3. The AI smoke removal method for real scenarios according to claim 1, characterized in that, Step 2 uses a convolutional network to extract the features F of the smoke image I and the dark channel image D respectively I and F D , calculate its self-attention feature F att , screen the mean value of the 1% brightest pixel points in each channel as the relative mean feature A m , then take the difference between F I and F att as the relative difference feature A d , and finally calculate the atmospheric light A of the real scene through linear combination ∞ =αA m +βA d , A ∞ is non-uniform, and α, β are combination coefficients.

4. The AI smoke removal method for real scenarios according to claim 1, characterized in that, In step 4, the predicted J, A ∞ and T are substituted into the atmospheric scattering model I(x) = J(x)t(x) + A ∞ (1 - t(x)) to calculate the reconstructed smoke image I rec , and the reconstruction loss function performs unsupervised learning on the texture and color of I rec , where VGG16 is used to calculate the perceptual content.

5. The AI smoke removal method for real scenarios according to claim 1, characterized in that, Multi-scale misaligned reference loss in step 5 Perform supervised learning on the de-smoking network for color and texture, where Use the VGG16 network to calculate the multi-scale content loss between J and J ref and Use the discriminator of patch-GAN to calculate the multi-scale adversarial loss between J and J ref between them.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the program, it implements the AI smoke removal method for real scenarios according to any one of claims 1-5.

7. A computer-readable storage medium, having a computer program stored thereon, characterized in that, when the program is executed by a processor, it implements the AI smoke removal method for real scenarios according to any one of claims 1-5.

8. A computer program product, comprising a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Light-weight and high-efficiency single-image smoke removing method

    CN112381723A

  • Multi-Feature Image Haze Removal

    US20160005152A1