End-to-end image defogging processing method based on deep neural network

Through the end-to-end image defog treatment method based on deep neural network, the full-dimensional dynamic convolution and pyramid pooling module extract features, and combined with the atmospheric scattering model, the problem of poor defog removal effect in complex scenes is solved, and a better image defog removal effect is achieved.

CN119941569APending Publication Date: 2025-05-06FUJIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510015168.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing image defogging technology is difficult to restore foggy-free images consistent with real scenes when processing complex scenes, resulting in problems such as artifacts and excessive enhancement.

Method used

The end-to-end image defogging treatment method based on deep neural network is adopted. By constructing a deep neural network defogging model, the features are extracted using full-dimensional dynamic convolution and pyramid pooling modules, and combined with the atmospheric scattering model, the fogging-free image is estimated.

Benefits of technology

It effectively improves the defog performance, and the image details retention and image contrast after defog are better, closer to the real scene, and conform to the visual experience of the human eye.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941569A_ABST
    Figure CN119941569A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to an end-to-end image defogging processing method based on a deep neural network, and the method comprises the following steps: S1, constructing a data set, and dividing the data set into a training set, a verification set and a test set; s2, constructing a deep neural network defogging model; s3, training a deep neural network defogging model by using the data set in the step S1; and S4, inputting a to-be-processed foggy image into the trained deep neural network defogging model for defogging processing to obtain a fogless image. According to the invention, the defogging effect of the foggy image can be effectively improved, so that the defogged image is closer to a real scene and accords with the visual perception of human eyes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an end-to-end image defogging method based on a deep neural network. Background Art

[0002] The introduction and development of convolutional neural networks (CNNs) have opened up new avenues for image dehazing research and further promoted the development of image dehazing research. Traditional methods usually have a strong dependence on the input image scene and are difficult to generalize to diverse real environments. When dealing with complex scenes, traditional methods may not be able to restore a haze-free image that is consistent with the real scene, resulting in artifacts, over-enhancement and other problems.

[0003] The atmospheric scattering model can better describe the scattering and attenuation of light by fog. This method based on the physical model can more accurately characterize the impact of fog on the image. By estimating the transmittance and atmospheric light value, the clear image of the scene can be better restored. Deep learning has superior learning ability and is better adapted to complex scenes. With the growth of data and computing power, it has great potential. Although the classic AOD-Net (All in One Dehazing Network) has achieved better results than traditional dehazing algorithms, it still lacks dehazing capabilities, resulting in dark images and unclear details in the dehazed images. Summary of the invention

[0004] To solve the above problems, the present invention provides an end-to-end image dehazing processing method based on a deep neural network.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: an end-to-end image defogging processing method based on a deep neural network, comprising the following steps:

[0006] S1: Construct a dataset and divide it into training set, validation set and test set;

[0007] S2: Build a deep neural network dehazing model;

[0008] S3: Use the data set in step S1 to train the deep neural network dehazing model;

[0009] S4: Input the foggy image to be processed into the trained deep neural network defogging model for defogging to obtain a fog-free image.

[0010] As a possible implementation, further, in step S1, the data set includes the NYU2 data set and the RESIDE data set;

[0011] The NYU2 dataset contains foggy images and corresponding fog-free images, and the NYU2 dataset is divided into a training set and a validation set in a ratio of 9:1;

[0012] The test set uses the SOTS outdoor synthetic fog dataset and the RTTS real fog dataset in the RESIDE dataset.

[0013] As a possible implementation, further, the deep neural network dehazing model in step S2 includes a first convolution, a second convolution, a third convolution, a fourth convolution, a fifth convolution, a pyramid pooling module and a sixth convolution.

[0014] As a possible implementation, further, the image processing method of the deep neural network defogging model includes the following steps:

[0015] 1) Feature extraction is performed through the first to fifth convolutions to obtain feature map F8;

[0016] 2) The global features extracted from the feature map F8 are input into the improved pyramid pooling module for multi-scale pooling, and then 1×1 depth convolution is applied to each layer of the pooled feature map to obtain a feature map with a channel number of 1; then the feature map of each layer is upsampled by bilinear interpolation filling to restore the feature map to the size of the original feature map, and then the generated feature map is fused with the original feature map in the channel dimension to obtain the feature map F9;

[0017] 3) The feature map F9 is input into the sixth convolution, and the sixth convolution uses a 3×3 two-dimensional convolution to convolve the feature map F9. The output result of the convolution is the K(x) value to be estimated;

[0018] 4) Use K(x) as its adaptive parameter to input into the atmospheric scattering model to estimate J(x), and finally obtain the corresponding fog-free image; where J(x) is expressed as follows:

[0019] J(x)=K(x)I(x)-K(x)+b

[0020] In the above formula, x is the pixel position, I(x) and J(x) are the target fog image and the fog-free image respectively, and b is a constant bias with a default value of 1.

[0021] As a possible implementation, further, in step 1), the first and second convolutions use full-dimensional dynamic convolution (Pyramid Pooling Module, ODConv); the third, fourth, and fifth convolutions use two-dimensional convolutions, which use 5×5, 7×7, and 3×3 convolution kernels respectively;

[0022] Step 1) specifically includes the following steps:

[0023] The input foggy image is passed through the first full-dimensional dynamic convolution to obtain the feature map F1;

[0024] The first feature map is subjected to a second full-dimensional dynamic convolution to obtain a feature map F2;

[0025] The feature map F1 and the feature map F2 are fused to obtain the feature map F3, and the feature map F3 is subjected to a third convolution to obtain the feature map F4;

[0026] The feature map F4 and the feature map F2 are fused to obtain a feature map F5, and the feature map F5 is subjected to the fourth convolution to obtain a feature map F6;

[0027] The feature map F1, the feature map F2, the feature map F4 and the feature map F6 are fused to obtain the feature map F7, and the feature map F7 is subjected to the fifth convolution to obtain the feature map F8.

[0028] As a possible implementation, further, the improved pyramid pooling module (PPM) in step 2) uses three pyramids of different scales for feature fusion; it performs three scale pooling on the feature map extracted by the network to obtain feature pyramids of 2×2, 4×4, and 8×8.

[0029] As a possible implementation, further, the model training method in step S3 is as follows:

[0030] The network hyperparameters are set, and the deep neural network defogging model is trained using the PyTorch framework. During the training process, a training model and defogging results are generated at each iteration until the maximum number of iterations is reached and the network training is completed.

[0031] As a possible implementation, further, the loss function in the model training uses the SSIM loss function; the maximum number of iterations is set to 10; the Bach size is set to 8; and the learning rate is set to 1×10 -4 .

[0032] As a possible implementation, further, the peak signal-to-noise ratio (PSNR) and structural self-similarity (SSIM) between the input foggy image and the corresponding output fog-free image are calculated, and the values ​​of SSIM and PSNR are combined to measure the defogging effects of the deep convolutional neural network and the optimal model.

[0033] As a possible implementation method, further, SOTS outdoor synthetic fog images and RTTS real fog images are used to test the performance of deep convolutional neural networks and optimal models.

[0034] The beneficial effects of the present invention are:

[0035] The deep convolutional neural network constructed by the present invention uses a full-dimensional dynamic convolution combined with a common two-dimensional convolution to effectively extract rich features to improve the defogging performance; by introducing an improved pyramid pooling module, the network's receptive field can be expanded and features of different scales can be integrated to better capture image information. In addition, the present invention uses the SSIM loss function to optimize the brightness and structural detail similarity of the image, so that the defogging image is closer to the real scene and conforms to the visual perception of the human eye. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a simplified flow chart of the present invention;

[0037] Figure 2 It is the structural diagram of the PPM module in the present invention;

[0038] Figure 3 This is a structural diagram of the deep neural network defogging model in the present invention;

[0039] Figure 4 This is a comparison chart of PSNR indicators;

[0040] Figure 5 This is a comparison chart of SSIM indicators;

[0041] Figure 6 This is the comparison of SOTS synthetic fog defogging effect;

[0042] Figure 7 Comparison of RTTS real fog dehazing effects. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] See attached Figure 1 As shown, this embodiment provides an end-to-end image dehazing processing method based on a deep neural network, comprising the following steps:

[0045] S1: Construct a dataset and divide it into a training set, a validation set, and a test set. In this embodiment, the dataset includes the NYU2 dataset and the RESIDE dataset. The NYU2 dataset contains 27,256 foggy images and 1,449 corresponding fog-free images, which are divided into a training set and a validation set at a ratio of 9:1. The test set uses the SOTS (outdoor synthetic fog) dataset in the RESIDE dataset, with a total of 500 synthetic fog images, and there is also the RTTS real fog dataset.

[0046] S2: Construct a deep neural network dehazing model; in this embodiment, the deep neural network dehazing model includes a first convolution, a second convolution, a third convolution, a fourth convolution, a fifth convolution, a pyramid pooling module and a sixth convolution.

[0047] See attached Figure 3 As shown, in this embodiment, the image processing method of the deep neural network defogging model includes the following steps:

[0048] 1) Feature extraction is performed through the first to fifth convolutions to obtain a feature map F8; wherein, the first and second convolutions use full-dimensional dynamic convolution (Pyramid Pooling Module, ODConv); the third, fourth, and fifth convolutions use two-dimensional convolutions, which use 5×5, 7×7, and 3×3 convolution kernels respectively; step 1) specifically includes the following steps:

[0049] The input foggy image is passed through the first full-dimensional dynamic convolution to obtain the feature map F1;

[0050] The first feature map is subjected to a second full-dimensional dynamic convolution to obtain a feature map F2;

[0051] The feature map F1 and the feature map F2 are fused to obtain the feature map F3, and the feature map F3 is subjected to a third convolution to obtain the feature map F4;

[0052] The feature map F4 and the feature map F2 are fused to obtain a feature map F5, and the feature map F5 is subjected to the fourth convolution to obtain a feature map F6;

[0053] The feature map F1, the feature map F2, the feature map F4 and the feature map F6 are fused to obtain the feature map F7, and the feature map F7 is subjected to the fifth convolution to obtain the feature map F8.

[0054] 2) As attached Figure 2As shown, the global features extracted by the obtained feature map F8 are input into the improved pyramid pooling module (Pyramid Pooling Module, PPM). In this embodiment, three pyramids of different scales are used for feature fusion. The specific process is: three scale pooling is performed on the feature map extracted by the network to obtain 2×2, 4×4, and 8×8 feature pyramids; then, a 1×1 depth convolution is used for each layer of the pooled feature map to obtain a feature map with a channel number of 1; then, the feature map of each layer is upsampled by bilinear interpolation filling, and the feature map is restored to the size of the original feature map, and then the generated feature map is fused with the original feature map in the channel dimension to obtain the feature map F9.

[0055] 3) The feature map F9 is input into the sixth convolution, and the sixth convolution uses a 3×3 two-dimensional convolution to convolve the feature map F9. The output result of the convolution is the K(x) value to be estimated;

[0056] 4) Use K(x) as its adaptive parameter to input into the atmospheric scattering model to estimate J(x), and finally obtain the corresponding fog-free image; where J(x) is expressed as follows:

[0057] J(x)=K(χ)I(χ)-K(x)+b

[0058] In the above formula, x is the pixel position, I(x) and J(x) are the target fog image and the fog-free image respectively, and b is a constant bias with a default value of 1.

[0059] S3: Use the data set in step S1 to train the deep neural network dehazing model; the specific model training method is as follows:

[0060] The network hyperparameters are set, and the deep neural network defogging model is trained using the PyTorch framework. During the training process, a training model and defogging results are generated at each iteration until the maximum number of iterations is reached and the network training is completed.

[0061] The loss function used in model training is the SSIM loss function; the maximum number of iterations is set to 10; the Bach size is set to 8; and the learning rate is set to 1×10 -4 .

[0062] In this embodiment, the peak signal-to-noise ratio (PSNR) and structural self-similarity (SSIM) between the input foggy image and the corresponding output fog-free image are calculated, and the values ​​of SSIM and PSNR are combined to measure the defogging effects of the deep convolutional neural network and the optimal model.

[0063] In this embodiment, SOTS outdoor synthetic fog images and RTTS real fog images are used to test the deep convolutional neural network and the optimal model performance.

[0064] S4: Input the foggy image to be processed into the trained deep neural network defogging model for defogging to obtain a fog-free image.

[0065] Performance Testing

[0066] The dehazing effect of the scheme of the present invention is compared with that of the classic AOD-Net (All in One Dehazing Network).

[0067] See attached Figure 4 and 5 As shown in the figure, it can be seen that the PSNR and PSNR of the scheme of the present invention are higher than those of the classic AOD-Net. Therefore, the dehazing effect of the scheme of the present invention is better than that of the classic AOD-Net.

[0068] Through experimental tests on the SOTS data set, the PSNR mean and PSNR mean of the present invention and AOD-Net are shown in Table 1 below:

[0069]

[0070] It can be seen from Table 1 that the SSIM (structural similarity) mean of the solution of the present invention reaches 0.88, and the PSNR mean (peak signal-to-noise ratio) is improved by 2.86 dB compared with AOD-Net; therefore, the present invention has a better performance in image defogging, and the image detail retention and image contrast after defogging are better (as shown in the attached figure). Figure 6 shown).

[0071] See attached Figure 7 As shown in the figure, it can be seen that the present invention also has better detail retention and image contrast than AOD-Net on the real fog image.

[0072] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An end-to-end image dehazing method based on a deep neural network, characterized in that: The steps include: S1: Construct a dataset and divide it into training set, validation set and test set; S2: Build a deep neural network dehazing model; S3: Use the data set in step S1 to train the deep neural network dehazing model; S4: Input the foggy image to be processed into the trained deep neural network defogging model for defogging to obtain a fog-free image.

2. The end-to-end image dehazing processing method based on deep neural network according to claim 1 is characterized in that: The data sets in step S1 include the NYU2 data set and the RESIDE data set; The NYU2 dataset contains foggy images and corresponding fog-free images, and the NYU2 dataset is divided into a training set and a validation set in a ratio of 9:1; The test set uses the SOTS outdoor synthetic fog dataset and the RTTS real fog dataset in the RESIDE dataset.

3. The end-to-end image dehazing processing method based on deep neural network according to claim 1, characterized in that: The deep neural network dehazing model in step S2 includes a first convolution, a second convolution, a third convolution, a fourth convolution, a fifth convolution, a pyramid pooling module and a sixth convolution.

4. The end-to-end image dehazing processing method based on deep neural network according to claim 3 is characterized in that: The image processing method of the deep neural network defogging model comprises the following steps: 1) Feature extraction is performed through the first to fifth convolutions to obtain feature map F8; 2) The global features extracted from the feature map F8 are input into the improved pyramid pooling module for multi-scale pooling, and then 1×1 depth convolution is applied to each layer of the pooled feature map to obtain a feature map with a channel number of 1; then the feature map of each layer is upsampled by bilinear interpolation filling to restore the feature map to the size of the original feature map, and then the generated feature map is fused with the original feature map in the channel dimension to obtain the feature map F9; 3) The feature map F9 is input into the sixth convolution, and the sixth convolution uses a 3×3 two-dimensional convolution to convolve the feature map F9. The output result of the convolution is the K(x) value to be estimated; 4) Use K(x) as its adaptive parameter to input into the atmospheric scattering model to estimate J(x), and finally obtain the corresponding fog-free image; where J(x) is expressed as follows: J(x)=K(x)I(x)-K(x)+b In the above formula, x is the pixel position, I(x) and J(x) are the target fog image and the fog-free image respectively, and b is a constant bias.

5. The end-to-end image dehazing method based on deep neural network according to claim 4, characterized in that: In step 1), the first and second convolutions use full-dimensional dynamic convolution; the third, fourth, and fifth convolutions use two-dimensional convolution, which use 5×5, 7×7, and 3×3 convolution kernels respectively; Step 1) specifically includes the following steps: The input foggy image is passed through the first full-dimensional dynamic convolution to obtain the feature map F1; The feature map F1 is subjected to a second full-dimensional dynamic convolution to obtain a feature map F2; The feature map F1 and the feature map F2 are fused to obtain the feature map F3, and the feature map F3 is subjected to a third convolution to obtain the feature map F4; The feature map F4 and the feature map F2 are fused to obtain a feature map F5, and the feature map F5 is subjected to the fourth convolution to obtain a feature map F6; The feature map F1, the feature map F2, the feature map F4 and the feature map F6 are fused to obtain the feature map F7, and the feature map F7 is subjected to the fifth convolution to obtain the feature map F8.

6. The end-to-end image dehazing processing method based on deep neural network according to claim 4, characterized in that: The improved pyramid pooling module in step 2) uses three pyramids of different scales for feature fusion; it performs three scale pooling on the feature map extracted by the network to obtain feature pyramids of 2×2, 4×4, and 8×8.

7. The end-to-end image dehazing processing method based on deep neural network according to claim 1, characterized in that: The model training method in step S3 is as follows: The network hyperparameters are set, and the deep neural network defogging model is trained using the PyTorch framework. During the training process, a training model and defogging results are generated at each iteration until the maximum number of iterations is reached and the network training is completed.

8. The end-to-end image dehazing processing method based on deep neural network according to claim 7, characterized in that: The loss function used in model training is the SSIM loss function; the maximum number of iterations is set to 10; the Bach size is set to 8; and the learning rate is set to 1×10 -4 .

9. The end-to-end image dehazing processing method based on deep neural network according to claim 7, characterized in that: The PSNR and SSIM between the input foggy image and the corresponding output fog-free image are calculated, and the values ​​of SSIM and PSNR are combined to measure the defogging effects of the deep convolutional neural network and the optimal model.

10. The end-to-end image dehazing processing method based on deep neural network according to claim 7, characterized in that: SOTS indoor and outdoor synthetic fog images and RTTS real fog images are used to test the performance of deep convolutional neural networks and optimal models.