An image dehazing method based on a deep neural network

By constructing a deep neural network that includes multi-scale feature fusion of airspace and frequency domain feature extraction, combined with hollow convolution and attention mechanism, the problems of incomplete fog removal and loss of details in the existing technology are solved, and better image recovery effect in haze environments is achieved.

CN115689932BActive Publication Date: 2025-07-01CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211398228.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2025-07-01
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

The existing image defogging method is not effective in dealing with images in complex haze environments, and traditional methods are difficult to effectively embed in end-to-end networks, resulting in incomplete image defogging, color distortion and loss of details.

Method used

A deep neural network including a multi-scale feature fusion module in the airspace and a frequency domain feature extraction module is adopted, combining the hollow convolution Res-Block and CBAM attention mechanism to restore clear images through end-to-end training.

Benefits of technology

Effectively extract image airspace and frequency domain features, improve the fog removal effect in complex environments, retain image details and edge information, and improve image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115689932B_ABST
    Figure CN115689932B_ABST
Patent Text Reader

Abstract

The present invention relates to an image defogging method based on a deep neural network and belongs to the field of image processing. Image data of a foggy image and a fog-free image are acquired; a deep neural network including a dual-domain multi-scale feature fusion module is created so that the model can better extract various detailed features in the image; the neural network is trained and tested to obtain an image defogging network model; the foggy image is defogged by using the trained deep neural image defogging network model, and a fog-free image is output. The deep neural defogging network provided by the present invention can more effectively extract features from the spatial domain and the frequency domain of the image, and also achieves good results in defogging images in complex environments and low visibility. The present invention can be used for tasks such as target detection in haze weather and traffic sign detection in autonomous driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and relates to an image dehazing method based on a deep neural network. Background Art

[0002] Currently, image processing has become an important field in deep learning. In image processing, image dehazing also plays an important role. In the environment we live in, due to the existence of atmospheric particles in the air, the quality of the pictures we take in outdoor scenes is often unsatisfactory, and the visual effect will decline significantly. Fog, smoke, etc. are the phenomena caused by the absorption and refraction of atmospheric light by these particles. Images in these scenes usually show increased brightness, decreased contrast, etc., which are not conducive to the extraction and identification of image features, resulting in difficulties in the use of various devices relying on the imaging system, and further increasing the difficulty of subsequent image processing. In the fields of image processing and computer vision, such as image recognition, target tracking, image classification, target detection, etc., people usually assume that the images to be processed are clear. Hazy images will make these problems more difficult to handle.

[0003] On the other hand, with environmental pollution, there are more and more hazy days, which not only bring troubles to people's daily lives, but also cause great difficulties to video surveillance. As mentioned before, video images taken in foggy weather have high brightness, low contrast, and low visibility, resulting in difficulties in distinguishing some illegal license plates and the like that appear in surveillance videos. Image dehazing is a key technology for further processing hazy images, and its expected result is to recover the potential clear image from the hazy image taken in foggy weather, and improve the quality and visual effect of the hazy image. It can be seen that, whether from the research perspective or the application perspective, there is an urgent need for dehazing processing of hazy images.

[0004] In recent years, many image dehazing methods have been proposed at home and abroad, but there are still some problems in image dehazing:

[0005] (1) When using the atmospheric light scattering model for dehazing, due to the inaccuracy of the estimation of the atmospheric light intensity value and the transmittance, it often leads to further amplification of the error when calculating the haze-free image, and finally results in problems such as incomplete image dehazing and color distortion.

[0006] (2) Using the atmospheric light scattering model to dehaze images instead of dehazing images in an end-to-end manner, which will make it difficult for the network model to be well embedded into the network models of other visual tasks for image preprocessing.

[0007] (3) In real environments, foggy images can be complex and diverse. For thick fog, images with cluttered backgrounds, indoor foggy images, foggy images under multiple light sources at night, etc., the defogging effect is still not ideal enough. A simple deep neural network that only processes images from top to bottom will lose some low-level features during repeated convolution and pooling processes, resulting in incomplete defogging and loss of details in the backgrounds of some complex images. Summary of the Invention

[0008] In view of the above, the purpose of the present invention is to provide an image defogging method based on a deep neural network.

[0009] To achieve the above purpose, the present invention provides the following technical solutions:

[0010] An image defogging method based on a deep neural network, comprising the following steps:

[0011] S1: Obtain foggy images and corresponding fog-free image data, and divide them into a training set and a test set;

[0012] S2: Create a deep neural network model including a spatial multi-scale feature fusion module and a frequency domain feature extraction module;

[0013] S3: Use the created training set and test set to train and test the deep neural network model;

[0014] S4: Use the trained image defogging network model to defog the foggy images.

[0015] Optionally, the specific steps of S1 are as follows:

[0016] S11: Obtain fog-free images in various environments, including indoor and outdoor environments;

[0017] S12: Calculate the corresponding foggy images according to the atmospheric light scattering model; the atmospheric light scattering model is as follows:

[0018] I(x) = J(x)t(x) + A(1 - t(x)) (1)

[0019] In the formula, I(x) and J(x) respectively represent the pixel values of the observed foggy image and the clear image at x, A is the global atmospheric light intensity, which is set as a constant, and t(x) represents the transmittance of the medium, as follows:

[0020] t(x) = e -βd(x)

[0021] The physical meaning of t(x) is the proportion of light that can reach the detection system after attenuation by particles; where β is the refractive coefficient and d(x) represents the depth value of the image at x;

[0022] By setting the atmospheric light intensity A value between [0.7, 1.0] and uniformly selecting the β value between [0.1, 1.8], the atmospheric light scattering model is used to synthesize corresponding foggy pictures for the collected fog-free pictures. Each clear picture generates 10 different foggy pictures, and finally a dataset for image defogging is obtained;

[0023] S13: Divide the training set and the test set in a ratio of 7:3.

[0024] Optionally, the specific S2 is as follows:

[0025] S21: Determine the image defogging deep neural network model. Using the encoder-decoder network as the basic structure of the spatial domain feature fusion network, it is corrected. Three residual blocks with dilated convolutions are added to the feature extraction network, and the features extracted by each residual block are fused to ensure that while extracting deep features, shallow features are also retained, and all detailed features of the defogged image are restored as much as possible; at the same time, a frequency domain feature extraction sub-network is introduced into the network, and the frequency domain feature extraction sub-network is added to the backbone network to ensure multi-level and multi-scale extraction of the features of the image, ensuring the thoroughness of image de-fogging and the restoration of image details and edges;

[0026] S22: The overall network structure is divided into a spatial domain feature extraction sub-network and a frequency domain feature extraction sub-network. Then, through splicing, the feature maps extracted by the spatial domain feature sub-network and the frequency domain feature extraction sub-network are added and fused. Finally, deconvolution and convolution operations are performed to restore the image size and output a clear image;

[0027] The spatial domain feature extraction sub-network is a spatial domain feature extraction sub-network with two convolutional operations, one downsampling convolutional operation, and three dilated convolution Res-Block modules. The outputs of each Res-Block dilated convolution are concatenated to obtain the final output of the sub-network; and the dilated convolution Res-Block module includes a smooth dilated convolution block, a residual block, a BN layer, and a Relu layer; the smooth dilated convolution block uses dilated convolution, sets the dilation rates to 2 and 3 respectively, performs dilated convolution respectively and then stacks them to eliminate the grid artifact problem; and the actual receptive field size of the dilated convolution is:

[0028] K = k+(k - 1)(r - 1) (2)

[0029] k is the size of the original convolution kernel, set to 3×3, and r is the dilation rate; the BN layer and the Relu layer are adopted, and the spatial-channel CBAM attention mechanism module is added at the end of each Res-Block, combining the spatial and channel attention mechanism modules; this module first passes the input through the channel attention module to focus on which channels in the image have more important features, and obtains the corresponding channel attention map M C (F):

[0030] M C (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) (3)

[0031] Where σ is the Sigmoid function, that is, the input feature map is respectively passed through max pooling and average pooling, then through MLP, and the features output by MLP are added, and finally through the Sigmoid function to generate the final channel attention feature map; then the channel attention feature map is used as the input of the spatial attention module to make it focus on which features on the channel are more important, and obtain the further spatial attention map M S (F):

[0032] M S (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])) (4)

[0033] The spatial attention module first performs channel-based max pooling and average pooling, reducing it to one channel, then performs a concat splicing operation on the two results, then passes through a convolution operation f, and finally passes through the Sigmoid function to obtain the spatial attention feature map; finally, the spatial attention feature map is multiplied by the channel attention feature map to obtain the finally generated feature map, so that the dehazing network model can have better performance; there are two outputs in each Res-Block, one is used as a part of the subsequent multi-scale information fusion, and the other is used as the input of the next Res-Block to further extract deep features;

[0034] S24: In the frequency-domain feature extraction sub-network, first, the input image is transformed by Fourier transform to obtain the frequency-domain map of the image. Then, the Laplacian operator and the Gaussian filter are used to filter out the low-frequency and high-frequency components respectively, and finally, the high-frequency and low-frequency images corresponding to the original image are obtained. The high- and low-frequency sub-images are convolved with the dilated convolution Res-Block to extract the corresponding feature maps in the frequency domain. Then, the inverse Fourier transform is performed to transform back from the frequency domain to the spatial domain. Next, this feature map is concatenated with the feature map obtained from the previous spatial-domain feature extraction sub-network, that is, the feature maps are added. Finally, it is input into the spatial-channel attention module, and finally, through deconvolution and two convolutions, the final clear image is obtained.

[0035] S3: The loss function of the network is divided into two parts: L1 loss and perceptual loss.

[0036] Training with L1 loss is more stable and is widely used in image restoration tasks. The function is as follows:

[0037]

[0038] where I i (x) represents the pixel value of the foggy image in the i-th channel, and J i (x) represents the pixel value of the fog-free image in the i-th channel.

[0039] The pre-trained VGG is used as the loss network, and the feature maps are extracted from the last layer of the first three stages. Therefore, the perceptual loss function is as follows:

[0040]

[0041] where and are the feature maps of the fog-free image after defogging and the real fog-free image in the VGG network respectively. C, H, and W represent the number of channels, height, and width of the feature map. The total loss function is as follows:

[0042] L = L1 + λL p (7).

[0043] Optionally, the S3 is specifically: According to the constructed deep neural network, the training set is used to train the deep neural network to make the network reach the optimal state, realizing end-to-end training.

[0044] Optionally, the S4 is specifically: Using the trained image defogging deep neural network model to defog the foggy image to obtain the corresponding clear image.

[0045] The beneficial effects of the present invention are as follows: The deep neural dehazing network provided by the present invention can more effectively extract features from the spatial and frequency domains of images, and also achieves good results in dehazing images in complex environments and low visibility conditions. The present invention can be used for tasks such as target detection in haze weather and traffic sign detection in autonomous driving.

[0046] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0048] Figure 1 is a flowchart of an image dehazing method based on a deep neural network according to the present invention;

[0049] Figure 2 is a sub-network diagram of spatial feature extraction according to the present invention;

[0050] Figure 3 is a sub-network diagram of frequency domain feature extraction according to the present invention;

[0051] Figure 4 is a module structure diagram of a feature extraction residual block combined with a spatial-channel attention mechanism in the present invention;

[0052] Figure 5 is an overall flowchart of the deep neural network model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically, and the following embodiments and the features in the embodiments can be combined with each other without conflict.

[0054] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams rather than physical diagrams, and should not be construed as limiting the present invention; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.

[0055] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the attached drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0056] As Figure 1 shown, an image dehazing method based on a deep neural network includes the following steps:

[0057] S1. Obtain foggy images and corresponding fog-free image data, and divide them into a training set and a test set;

[0058] The specific method of step S1 is as follows:

[0059] S11. Obtain fog-free images in various environments, including fog-free images in indoor and outdoor environments, a total of 1500 images.

[0060] S12. Calculate the corresponding foggy images according to the atmospheric light scattering model. The atmospheric light scattering model is as follows:

[0061] I(x) = J(x)t(x) + A(1 - t(x)) (1)

[0062] In the formula, I(x) and J(x) respectively represent the pixel values of the observed foggy image and the clear image at x, A is the global atmospheric light intensity, usually assumed to be a constant, and t(x) represents the transmittance of the medium, as follows:

[0063] t(x) = e -βd(x)

[0064] The physical meaning of t(x) is the proportion of light that can reach the detection system after attenuation by particles. Among them, is the refractive index, and d(x) represents the depth value of the image at x.

[0065] By setting the atmospheric light intensity A value between [0.7, 1.0] and the value between [0.1, 1.8], the collected fog-free images are synthesized into corresponding foggy images using the atmospheric light scattering model. That is, for each clear image, A and the value are randomly selected within the above ranges, and 10 different foggy images are generated. Finally, an image defogging dataset is obtained, and this dataset has a total of 15,000 blurred foggy images.

[0066] S13. Divide the training set and the test set in a ratio of 7:3.

[0067] S2. Create a deep neural network model that includes a spatial domain multi-scale feature fusion module and a frequency domain feature extraction module;

[0068] The specific operation of step S2 is as follows:

[0069] S21. Determine the image defogging deep neural network model. Using the encoder-decoder network as the basic structure of the spatial domain feature fusion network, the present invention further modifies it. Three residual blocks with dilated convolutions are added to the feature extraction network, and the features extracted by each residual block are fused to ensure that while extracting deep features, shallow features are also retained, and all detailed features of the defogged image are restored as much as possible. At the same time, a frequency domain feature extraction sub-network is introduced into the network. Many feature information of the image is also contained in the frequency domain of the image, such as information about the edges, textures, brightness, and contrast of the image. Therefore, the frequency domain feature extraction sub-network is added to the backbone network to ensure multi-level and multi-scale extraction of the features of the image, and to ensure the thoroughness of image defogging and the restoration of image details and edges.

[0070] S22. Figure 5 This is the overall network structure diagram of the present invention, which is divided into a spatial domain feature extraction sub-network and a frequency domain feature extraction sub-network. Then, through splicing, the feature maps extracted by the spatial domain feature sub-network and the frequency domain feature extraction sub-network are fused. Finally, deconvolution and convolution operations are performed to restore the image size and output a clear image. The following is a detailed explanation of the two sub-networks respectively.

[0071] S23. Figure 2 This is the spatial domain feature extraction sub-network of the present invention, which is a spatial domain feature extraction sub-network with two convolutional operations, one downsampling convolutional operation, and three dilated convolution Res-Block modules. The outputs of each Res-Block dilated convolution are concatenated to obtain the final output of the sub-network. And the dilated convolution Res-Block module is as Figure 4 shown, including a smooth dilated convolution block, a residual block, a BN layer, and a Relu layer. The dilation rates are set to 2 and 3 respectively, and then dilated convolutions are performed respectively and stacked to eliminate the grid artifact problem. And the actual receptive field size of the dilated convolution is:

[0072] K = k+(k - 1)(r - 1) (2)

[0073] k is the size of the original convolution kernel, set to 3×3, and r is the dilation rate. Then, a BN layer and a Relu layer are adopted, mainly to solve the problem of gradient disappearance and reduce the computational complexity during backpropagation. And a CBAM (spatial-channel) attention mechanism module is added at the end of each Res-Block. It is a module that combines spatial and channel attention mechanisms to obtain the corresponding channel attention map M C (F):

[0074] M C (F)=σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) (3) where σ is the Sigmoid function and F is the feature map obtained by the encoder. That is, the input feature map is respectively passed through max pooling and average pooling, then through an MLP (Multi-layer Perceptron), and the features output by the MLP are added, and finally through the Sigmoid function to generate the final channel attention feature map; then the channel attention feature map is used as the input of the spatial attention module to make it focus on which features on the channel are more important to obtain a further spatial attention map M S (F):

[0075] M S (F)=σ(f 7×7 ([AvgPool(F); MaxPool(F)])) (4)

[0076] The spatial attention module first performs a channel-based max pooling and average pooling to reduce to 1 channel, then performs a concat splicing operation on the two results, then passes through a convolution operation with a convolution kernel of 7×7, and finally passes through the Sigmoid function to obtain the spatial attention feature map. Finally, the spatial attention feature map and the channel attention feature map are multiplied element by element to obtain the finally generated feature map, so that the defogging network model can have better performance. There are two outputs in each Res-Block. One is part of the fusion of subsequent multi-scale information, and the other is the input of the next Res-Block to further extract deep features.

[0077] S24、 Figure 3This is the frequency-domain feature extraction subnet of the present invention. First, the input image is subjected to Fourier transform to obtain the frequency-domain map of the image. Then, the Laplacian operator and Gaussian filter are used to filter out the low-frequency and high-frequency components respectively, and finally, the high-frequency and low-frequency images corresponding to the original image are obtained respectively. Since a lot of information such as details, edges, and contrast of the image is contained in the high-frequency and low-frequency components of the image, the high-frequency and low-frequency sub-images are also subjected to convolution operations and dilated convolution Res-Blocks to extract the corresponding feature maps in the frequency domain. Then, inverse Fourier transform is performed to transform back from the frequency domain to the spatial domain. Next, this feature map is concatenated with the feature map obtained from the previous spatial-domain feature extraction sub-network, that is, the feature maps are added. Finally, it is input into the spatial-channel attention module, and finally, after deconvolution and two convolutions, the final clear image is obtained.

[0078] S25. The loss function of the network is divided into two parts: L1 loss and perceptual loss.

[0079] Training with L1 loss is more stable and is widely used in image restoration tasks. The function is as follows:

[0080]

[0081] where I i (x) represents the pixel value of the foggy image in the i-th channel, and J i (x) represents the pixel value of the fog-free image in the i-th channel.

[0082] In this paper, a pre-trained VGG is used as the loss network, and feature maps are extracted from the last layer of the first three stages. Therefore, the perceptual loss function is as follows:

[0083]

[0084] where and are the three feature maps of the fog-free image after dehazing and the real fog-free image passing through the loss network respectively. C, H, and W represent the number of channels, height, and width of the feature map respectively. Therefore, the total loss function is as follows:

[0085] L = L1 + λL p (7)

[0086] λ is the weight of L p and takes 0.05.

[0087] S3. Use the created training set and test set to train and test the deep neural network model;

[0088] The specific operation of step S3 is as follows:

[0089] According to the steps S1 - S2, configure the network environment, select the Windows operating system and the Pytorch deep learning framework for training. Use the Relu function as the activation function, and at the same time add the Batch Normalization layer to avoid problems such as gradient disappearance. Then use the training set to train the deep neural network. The network model is trained using the SGDM (SGD with Momentum) optimizer, where the β1 parameter is set to 0.9, the batch_size is set to 10, the initial learning rate is set to 0.005, and train for 2×10 5 iterations.

[0090] Then load the test set data, input it into the trained deep neural network, and then calculate its Structural Similarity (SSIM) and Peak Signal-to-Noise Ratio (PSNR) as the network model evaluation metrics.

[0091] Among them, SSIM is an index to measure the similarity between two images, and it is used to measure the similarity between the real haze-free image and the dehazed image after passing through the dehazing network. The SSIM is as follows:

[0092]

[0093] where μ x is the mean of x, μ y is the mean of y, is the variance of x, is the variance of y, σ xy is the covariance of x and y, c1 = (k1L) 2 and c2 = (k2L) 2 , which are two constants, L is the pixel value range, and k1 = 0.01, k2 = 0.03 are the default values.

[0094] And PSNR is an index to measure the image quality, and it is used to measure the image quality of the dehazed image after passing through the network. The PSNR is as follows:

[0095]

[0096] where MSE is the mean square error between the real haze-free image and the dehazed image after passing through the network, and MAX I represents the maximum value of the image point color.

[0097] The experimental situation of the present invention and some mainstream dehazing methods on the dataset is shown in Table 1:

[0098] Table 1 Comparison table of experimental results

[0099] Methods PSNR(dB) SSIM DehazeNet 22.46 0.8472 AOD-Net 20.29 0.8504 FFA-Net 31.04 0.9602 Ours 31.22 0.9762

[0100] Compared with other methods, there is a further improvement in both peak signal-to-noise ratio (PSNR) and structural similarity (SSIM).

[0101] S4. Use the trained image dehazing network model to dehaze the foggy image.

[0102] Furthermore, the step S4 specifically includes:

[0103] According to the above steps, use the trained image dehazing deep neural network model to dehaze the foggy image to obtain the corresponding clear image. According to the actual dehazing effect of this method, it can be known that whether it is an indoor or outdoor environment, the fog in the picture can be better removed.

[0104] In summary, the present invention can improve the dehazing effect, and due to the introduction of the frequency domain and attention mechanism, the restoration effect of the detailed features of the picture has been further improved, and it can ensure that the information such as the contrast and brightness of the picture is closer to the real clear picture.

[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. An image defogging method based on a deep neural network, characterized in that: The method includes the following steps: S1: Obtain the foggy image and the corresponding fog-free image data, and divide them into a training set and a test set; S2: Create a deep neural network model including a spatial multi-scale feature fusion module and a frequency domain feature extraction module; S21: Determine the image dehazing deep neural network model. Use the encoder-decoder network as the basic structure of the spatial feature fusion network and modify it. Add three residual blocks with dilated convolutions to the feature extraction network, and fuse the features extracted by each residual block to ensure that while extracting deep features, shallow features are also retained, and all detailed features of the dehazed image are restored as much as possible. At the same time, introduce a frequency domain feature extraction sub-network into the network, and add the frequency domain feature extraction sub-network to the backbone network to ensure multi-level and multi-scale extraction of image features, and ensure the thoroughness of image haze removal and the restoration of image details and edges; S22: The overall network structure is divided into a spatial feature extraction sub-network and a frequency domain feature extraction sub-network. Then, through splicing, the feature maps extracted by the spatial feature sub-network and the frequency domain feature extraction sub-network are added and fused. Finally, deconvolution and convolution operations are performed to restore the image size and output a clear image; S23: The spatial feature extraction sub-network is a spatial feature extraction sub-network with two convolutional operations, one downsampling convolutional operation, and three dilated convolution Res-Block modules. The outputs of each Res-Block dilated convolution are concatenated to obtain the final output of the sub-network. The dilated convolution Res-Block module includes a smooth dilated convolution block, a residual block, a BN layer, and a Relu layer. The smooth dilated convolution block uses dilated convolution, sets the dilation rates to 2 and 3 respectively, performs dilated convolution respectively and then stacks them to eliminate the grid artifact problem. The actual receptive field size of the dilated convolution is: K = k+(k - 1)(r - 1) (2) k is the size of the original convolution kernel, set to 3×3, and r is the dilation rate; the BN layer and the Relu layer are adopted, and the spatial-channel CBAM attention mechanism module is added at the end of each Res-Block block, combining the spatial and channel attention mechanism modules; this module first passes the input through the channel attention module to focus on which channels in the image have more important features and obtains the corresponding channel attention map M C (F): M C (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) (3) where σ is the Sigmoid function, that is, the input feature map is respectively passed through max pooling and average pooling, then through the MLP, and the features output by the MLP are summed, and finally passed through the Sigmoid function to generate the final channel attention feature map; then the channel attention feature map is used as the input of the spatial attention module to make it focus on which features on the channel are more important, and a further spatial attention map M is obtained S (F): M S (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])) (4) The spatial attention module first performs channel-based max pooling and average pooling, reduces to one channel, then concatenates the two results through a concat operation, then passes through a convolutional operation f, and finally passes through the Sigmoid function to obtain the spatial attention feature map. Finally, multiply the spatial attention feature map by the channel attention feature map to obtain the finally generated feature map, so that the dehazing network model can have better performance. There are two outputs in each Res-Block. One is used as part of the subsequent multi-scale information fusion, and the other is used as the input of the next Res-Block to further extract deep features; S24: In the frequency-domain feature extraction sub-network, first, the input image is transformed by Fourier transform to obtain the frequency-domain map of the image. Then, the Laplacian operator and the Gaussian filter are used to filter out the low-frequency and high-frequency components respectively, and finally, the high-frequency and low-frequency images corresponding to the original image are obtained. The high-frequency and low-frequency sub-images are convolved with the dilated convolution Res-Block to extract the corresponding feature maps in the frequency domain. Then, the inverse Fourier transform is performed to transform back from the frequency domain to the spatial domain. Next, this feature map is concatenated with the feature map obtained from the previous spatial-domain feature extraction sub-network, that is, the feature maps are added. Finally, it is input into the spatial-channel attention module, and finally, deconvolution and two convolutions are performed to obtain the final clear image. S3: Use the created training set and test set to train and test the deep neural network model. The loss function of the network is divided into two parts: L1 loss and perceptual loss. Training with L1 loss is more stable and is widely used in image restoration tasks. The function is as follows: Among which I i (x) represents the pixel value of the foggy image in the i-th channel, and J i (x) represents the pixel value of the fog-free image in the i-th channel; Use the pre-trained VGG as the loss network, and extract the feature maps from the last layer of the first three stages. The perceptual loss function is as follows: where and j = 1, 2, 3 represent the feature maps of the haze-free image after haze removal and the true haze-free image in the VGG network, respectively, and C, H, and W represent the number of channels, height, and width of the feature map; the total loss function is as follows: L = L1 + λL p (7) S4: Use the trained image dehazing network model to dehaze the hazy image.

2. The image defogging method based on a deep neural network according to claim 1, characterized in that: The specific content of S1 is as follows: S11: Obtain fog-free images in various environments, including fog-free images in indoor and outdoor environments. S12: Calculate the corresponding hazy images according to the atmospheric light scattering model. The atmospheric light scattering model is as follows: I(x) = J(x)t(x) + A(1 - t(x)) (1) In the formula, I(x) and J(x) represent the pixel values of the observed hazy image and the clear image at x respectively. A is the global atmospheric light intensity, which is set as a constant. t(x) represents the transmittance of the medium, as follows: t(x) = e -βd(x) The physical meaning of t(x) is the proportion of light that can reach the detection system after attenuation by particles. Among them, β is the refractive index, and d(x) represents the depth value of the image at x. By setting the value of the atmospheric light intensity A between [0.7, 1.0] and uniformly selecting the value of β between [0.1, 1.8], the collected fog-free images are synthesized into corresponding foggy pictures using the atmospheric light scattering model. Ten different foggy images are generated for each clear image, and finally, a dataset for image dehazing is obtained. S13: Divide the training set and test set in a ratio of 7:

3.

3. A method for image dehazing based on a deep neural network according to claim 1, characterized in that: The specific content of S3 is as follows: According to the constructed deep neural network, use the training set to train the deep neural network to make the network reach the optimal state and achieve end-to-end training.

4. The image defogging method based on a deep neural network according to claim 3, characterized in that: The specific content of S4 is as follows: Use the trained image dehazing deep neural network model to dehaze the hazy image to obtain the corresponding clear image.