An Image Dehazing Method Based on a Convolutional Neural Network Integrated with Transformer

By fusing Transformer and convolutional neural networks in image defog, the low-frequency and high-frequency information of the image are extracted, the problem of limited learning ability of convolutional neural networks in the prior art is solved, and a better defog removal effect is achieved.

CN116012253BActive Publication Date: 2025-06-24SHAANXI YUNYUE FUTURE ELECTRONIC TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310076270.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2025-06-24
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

The existing image defog removal method has the problem of limited learning ability of convolutional neural networks, resulting in poor defog removal effect.

Method used

The convolutional neural network based on fusion Transformer is adopted to extract the low-frequency information of the image through convolution operations. Transformer integrates the high-frequency information of the image to build an end-to-end defogging network model.

Benefits of technology

It achieves a smaller number of parameters and faster computing speed, which can better restore the overall information of the image, directly perform end-to-end defog removal, improving the defog removal effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116012253B_ABST
    Figure CN116012253B_ABST
Patent Text Reader

Abstract

The present invention relates to an image defogging method based on a convolutional neural network integrating Transformer, belonging to the technical fields of computer vision and image processing. The method includes: S1: Obtain the synthetic foggy dataset RESIDE as the training dataset and preprocess the dataset; S2: Construct a defogging network model: Based on the network structure of U-Net and the residual module, use the convolutional module and the Transformer module to construct an end-to-end defogging network model; S3: Input the preprocessed dataset into the constructed defogging network model, calculate the loss through the loss function during the training process, continuously iterate and update the model parameters, and finally obtain a trained defogging network model for image defogging. The present invention can better restore the overall information of the image and can directly perform end-to-end defogging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and image processing, and relates to an image dehazing method based on a convolutional neural network integrating Transformer. Background Art

[0002] In the context of the rapid development of information technology, images have become the main way for humans to express, obtain, and transmit information. Due to some external factors, such as bad weather, instability of imaging devices, etc., the quality of the acquired images is relatively low. The unclear images obtained will not only affect the subjective visual experience of the human eye, but also have an extremely serious impact on the performance of various image-based intelligent information processing systems. In order to improve the quality of images and enhance the clarity of images, it is very necessary to perform dehazing processing on the images.

[0003] Currently, image dehazing algorithms are mainly divided into three categories: dehazing methods based on image enhancement, dehazing methods based on prior knowledge, and dehazing methods based on deep learning. The first category, dehazing methods based on image enhancement, start from the foggy image itself, ignoring the formation mechanism of the foggy image, and make the visual effect of the picture clear by improving the contrast and color saturation of the foggy image. The second category, dehazing methods based on prior knowledge, based on the atmospheric scattering model, estimate the intermediate parameters of the atmospheric scattering model through various prior knowledge or theoretical assumptions, and then solve the atmospheric scattering model inversely to calculate the fog-free image. The third category, dehazing methods based on deep learning, this method no longer relies on manually extracting features, but constructs a neural network model to let the model learn how to restore clear images from the data. Early dehazing networks still relied on the atmospheric scattering model, such as: DehazeNet, by constructing a network model, learning the intermediate parameters, and then obtaining the fog-free image by inversely solving the atmospheric scattering model.

[0004] Although the deep learning method based on the atmospheric scattering model can achieve better dehazing effects than traditional methods, it also limits the learning ability of the convolutional neural network. Therefore, current dehazing methods directly use the convolutional neural network for end-to-end dehazing. The methods based on deep learning have developed rapidly in recent years, but there are still certain limitations and need to be further improved. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide an image dehazing method based on a convolutional neural network integrating Transformer, where the convolutional operation extracts the low-frequency information of the image, Transformer integrates the high-frequency information of the image, and uses fewer parameters and faster operation speed to achieve end-to-end dehazing.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] An image dehazing method based on a convolutional neural network integrating Transformer, specifically including the following steps:

[0008] S1: Obtain the synthetic foggy dataset RESIDE as the training dataset, and preprocess the dataset;

[0009] S2: Construct a dehazing network model: Based on the network structure of U-Net and residual modules, use convolutional modules and Transformer modules to construct an end-to-end dehazing network model;

[0010] S3: Input the preprocessed dataset into the constructed dehazing network model. During training, calculate the loss through the loss function, continuously iterate and update the model parameters, and finally obtain a trained dehazing network model for image dehazing.

[0011] Furthermore, step S1 specifically includes the following steps:

[0012] S11: Obtain the training dataset, including a pair of foggy images and clear images, denoted as (I, J), where I, J ∈ R C ×H×W , I represents the foggy image, J represents the clear image, C represents the number of channels of the image, and H and W respectively represent the height and width of the image;

[0013] S12: Randomly crop the images to obtain the images (I crop , J crop ), I crop represents the cropped foggy image, J crop represents the cropped clear image, I crop , J crop ∈ R C×h×w , h and w respectively represent the height and width of the cropped image, both set to 256;

[0014] S13: Randomly flip the images obtained in step S12 to enhance the dataset; and convert them into tensor form as the input of the dehazing network model.

[0015] Furthermore, in step S2, the architecture of the dehazing network model includes: an encoder, a decoder, and a feature fusion module; Based on the network structure of U-Net, use residual structures to construct an end-to-end dehazing network. The encoder consists of several dehazing modules, and each dehazing module consists of a Transformer and a convolutional neural network. The convolutional neural network extracts low-frequency information, and the Transformer extracts high-frequency information; The decoder consists of transposed convolution modules to complete the reconstruction of the fog-free image.

[0016] Further, in step S2, a defogging network model is constructed, which specifically includes the following steps:

[0017] S21: The encoder contains 3 defogging modules and 2 patch-merging modules. Each defogging module consists of several convolutional sub-modules in the first half and several Transformer sub-modules in the second half. The ratio of the convolutional module to the Transformer module in each defogging module decreases layer by layer. That is, in the shallow layer of the network, the convolutional module mainly extracts the detailed information of the image, while as the network depth increases, it is mainly the Transformer module that extracts the global information of the image;

[0018] S22: The role of the patch-merging module is to first partition the input into the patch partition module, that is, take every 4×4 adjacent pixels as a patch, and then flatten it in the channel direction. Then the height and width of the feature map are reduced to half of the original, and the number of channels becomes 4 times the original. After that, through a fully connected layer, the number of channels is mapped to 2 times the original number of channels, realizing the downsampling operation of the feature map;

[0019] S23: The defogging module is composed of 8 defogging sub-modules, and each defogging sub-module is a convolutional neural network module or a Transformer module;

[0020] S24: The decoder consists of a transposed convolution module and the convolutional module in the defogging module. The transposed convolution module consists of a convolutional layer and a PixelShuffle upsampling layer, and its role is to double the width and height of the feature map and halve the number of channels;

[0021] S25: The feature fusion module adopts SKNet and a residual connection network; SKNet adds an attention mechanism to different branches to enable the network to learn the importance of different branches, aiming to enable the feature fusion module to more effectively fuse residual features.

[0022] Further, in step S23, each defogging sub-module is a convolutional neural network module or a Transformer module;

[0023] (1) In the convolutional neural network module, in addition to ordinary convolution, it also includes SENet and PixNet. SENet is to enable the network to learn the importance of different channel features, while PixNet is to enable the network to learn the importance of features at different spatial points. The purpose of doing this is also to enable the network to better capture sufficient detailed information and improve the quality of the restored fog-free image;

[0024] (2) The Transformer module uses window self-attention to calculate self-attention, aiming to reduce the computational complexity. When using window self-attention, the feature map is first divided into windows of size 7×7, and then Self-Attention is performed separately inside each window, greatly reducing the computational complexity. At the same time, in order to enable information interaction between different windows, a Shifted Windows Multi-Head Self-Attention (SW-MS) follows a normal Windows Multi-head Self-Attention (W-MSA). The calculation formula is as follows:

[0025]

[0026] Among them, Q represents Query, K represents key, V represents the extracted information, and B is the relative position bias; Q, K, V ∈ R b×l×d , b is the Batch Size, d is the input dimension, and l is the number of divided windows;

[0027] The self-attention mechanism first performs a dot product operation on Query and Key, then uses the SoftMax function to extract the attention ratio and multiplies it by Value to obtain the specific feature map.

[0028] Further, in step S25, the specific operation process of effectively fusing residual features is: the feature map output by upsampling and the feature Figure 1 passed from the residual connection in the encoding stage are input into the feature fusion module, so that the network will not lose shallow features when restoring the foggy image; the formula is as follows:

[0029] [r1, r2] = SoftMax(F fc (Averpool(layer i + x)))

[0030] y = r1 × layer i + r2 × x + b

[0031] Among them, x is the output feature map of the previous layer, layer i is the feature map of the residual connection, Averpool is the global average pooling function, F fc (Linear-Relu-Linear) is the linear mapping layer, y is the final output, r1 and r2 are the weight values of the feature maps x and layer i output by the feature fusion module, and b is the bias value.

[0032] Furthermore, in step S2, in order to ensure that the restored image is more accurate, a soft constraint function derived from a variant of the atmospheric scattering model is introduced at the end of the network. The purpose of doing this is that the Transformer model itself lacks some inductive biases of the convolutional model and performs poorly on small-scale datasets. After introducing the soft constraint, the results of the network become more accurate. Its formula is:

[0033]

[0034] Among them, x represents the pixels in the image, K(x) represents the output feature map of the previous layer, I(x) is the input hazy image, J(x) represents the restored haze-free image, and t(x) represents the transmittance.

[0035] Furthermore, in step S3, the loss function consists of two parts: L1 loss and perceptual loss. Among them, the L1 loss is the average of the absolute errors between the predicted value and the true value, and its calculation formula is:

[0036]

[0037] Among them, N represents the number of pixels in the image, x represents the pixel value at a point in the dehazed image, and y represents the pixel value at a point in the clear image;

[0038] To ensure the details of the restored image, perceptual loss is introduced. The perceptual loss is to use other pre-trained feature extraction networks to extract the features of the output image and the target image to make them as similar as possible. The formula for the perceptual loss is:

[0039]

[0040] Among them, and Θ j (J) respectively represent the feature maps of the VGG16 network for the input hazy image and the true haze-free image, and C j H j W j represents the product of the number of channels output by the feature extraction network;

[0041] VGG16 trained on ImageNet is used as the feature extraction network, and its features of layers 3, 5, and 7 are used to generate the perceptual loss;

[0042] The final loss function is: L = L1 + 0.05L perc .

[0043] The beneficial effects of the present invention are as follows:

[0044] 1) The present invention proposes an image dehazing network that integrates Transformer and convolution. Its network structure is a residual network based on the U-Net structure, consisting of an encoding layer, a decoding layer, and bottleneck regions in two intermediate areas. This model extracts the detailed information of the image through a convolutional network and extracts the global information through Transformer. It can better restore the overall information of the image and can directly perform end-to-end dehazing.

[0045] 2) Most of the previous methods adjusted the feature extraction of the network by changing the scale and depth of the convolutional neural network. Due to the defects of convolution itself, stacking a sufficient network depth was required for the network to extract global features. The present invention proposes a network structure that integrates Transformer and convolution, which can extract sufficient local information and capture global dependency information without increasing the network depth.

[0046] 3) The present invention introduces SKNet and PixNet into the convolutional dehazing sub-module and adds an attention mechanism to both channel information and spatial information, enabling the model to learn the weight values of different information, focus on more important information, and improve the dehazing effect of the model.

[0047] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail and preferably with reference to the accompanying drawings, where:

[0049] Figure 1 is the overall flowchart of the image dehazing method based on the convolutional neural network integrating Transformer of the present invention;

[0050] Figure 2 is the structure diagram of the convolutional neural network model of the present invention;

[0051] Figure 3 is the structure diagram of the dehazing module of the present invention;

[0052] Figure 4 is the structure diagram of the convolutional dehazing sub-module of the present invention;

[0053] Figure 5 is the structure diagram of the Transformer dehazing sub-module of the present invention;

[0054] Figure 6It is the network diagram of the dehazing sub-module of the present invention;

[0055] Figure 7 It is the structural diagram of the SENet module of the present invention. Among them, Figure 7 (a) is the schematic diagram of the effect of the SENet module processing the image, Figure 7 (b) is the flowchart of the SENet module processing the image;

[0056] Figure 8 It is the structural diagram of the SKNet module of the present invention. Specific embodiments

[0057] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0058] Please refer to Figures 1 to 8 , the present invention provides an image dehazing method based on a convolutional neural network integrating Transformer, which specifically includes the following steps:

[0059] S1: Obtain the synthetic foggy dataset RESIDE as the training dataset, and preprocess the dataset: specifically include the following steps:

[0060] S11: A pair of foggy images and clear images in the training dataset, denoted as (I, J), where I, J ∈ R C×H×W , I represents the foggy image, J represents the clear image, C represents the number of channels of the image, and H and W respectively represent the height and width of the image.

[0061] S12: Randomly crop the image to obtain the images (I crop , J crop ), I crop represents the cropped foggy image, J crop represents the cropped clear image, I crop , J crop ∈ R C×h×w , where h and w respectively represent the height and width of the cropped image, both set to 256.

[0062] S13: Randomly flip the images obtained in step S12 to enhance the dataset. And convert it into a tensor form, which is used as the input of the dehazing network.

[0063] S2: Construct a defogging network model.

[0064] The architecture of the defogging model includes: an encoder, a decoder, and a feature fusion module. It is characterized in that based on the network structure of U-Net, a residual structure is used to construct an end-to-end defogging network. The encoder is composed of several defogging modules, and each module is composed of a Transformer and a convolutional neural network. The convolutional neural network extracts low-frequency information, and the Transformer extracts high-frequency information. The decoder is composed of deconvolution modules to complete the reconstruction of the fog-free image.

[0065] S21: The encoder contains 3 defogging modules and 2 patch-merging modules. Each defogging module is composed of several convolutional sub-modules in the first half and several Transformer sub-modules in the second half. The ratio of the convolutional module to the Transformer module in each defogging module decreases layer by layer. That is, in the shallow layer of the network, the convolutional module mainly extracts the detailed information of the image, and as the network depth increases, it is mainly the Transformer module that extracts the global information of the image.

[0066] The function of the patch-merging module is to first partition the input into the patch partition module, that is, take every 4×4 adjacent pixels as a patch, and then flatten it in the channel direction. Then the height and width of the feature map are reduced to half of the original, and the number of channels becomes 4 times the original. After that, through a fully connected layer, the number of channels is mapped to 2 times the original number of channels to achieve the downsampling operation of the feature map.

[0067] The defogging module is composed of 8 defogging sub-modules, and each sub-module is a convolutional neural network module or a Transformer module.

[0068] In the convolutional module, in addition to ordinary convolution, SENet and PixNet are also introduced. SENet is to enable the network to learn the importance of different channel features, and PixNet is to enable the network to learn the importance of features at different spatial points. The purpose of doing this is also to enable the network to better capture sufficient detailed information and improve the quality of the restored fog-free image.

[0069] S232: The Transformer module calculates self-attention in the way of window self-attention, aiming to reduce the computational complexity. When using window self-attention, the feature map is first divided into windows of size 7×7, and then Self-Attention is performed separately inside each window, greatly reducing the computational complexity. At the same time, in order to enable information interaction between different windows, a Shifted Windows Multi-Head Self-Attention (SW-MS) follows a normal Windows Multi-head Self-Attention (W-MSA). Its calculation formula is as follows:

[0070]

[0071] Among them, Q represents Query, K represents key, V represents the extracted information, B is the relative position bias; Q, K, V ∈ R b×l×d , b is the Batch Size, d is the input dimension, and l is the number of divided windows.

[0072] The self-attention mechanism first performs a dot product operation on Query and Key, then uses the SoftMax function to extract the attention ratio and multiply it by Value to obtain the specific feature map.

[0073] S24: The decoder consists of a transposed convolution module and the convolution module in the defogging module. The transposed convolution module consists of a convolution layer and a PixelShuffle upsampling layer, which doubles the width and height of the feature map and halves the number of channels.

[0074] S25: The feature fusion module uses SKNet and the residual connection network. By adding an attention mechanism to different branches, SKNet enables the network to learn the importance of different branches, aiming to make the feature fusion module more effectively fuse residual features. The specific operation process is to input the feature map output by upsampling and the features Figure 1 transmitted by the residual connection in the encoding stage into the feature fusion module, so that the network will not lose shallow features when restoring the foggy image. Its formula is as follows:

[0075] [r1, r2] = SoftMax(F fc (Averpool(layer i +x)))

[0076] y = r1×layer i +r2×x + b

[0077] Among them, x is the output feature map of the previous layer, layer i is the feature map of the residual connection, Averpool is the global average pooling function, F fc (Linear-Relu-Linear) is the linear mapping layer, y is the final output, r1 and r2 are the weight information of the feature maps x and layer i calculated by the feature fusion module, and b is the bias value.

[0078] S3: To ensure that the restored image is more accurate, a soft constraint function derived from a variant of the atmospheric scattering model is introduced at the end of the network. The purpose of doing this is because the Transformer model itself lacks some inductive biases of the convolutional model and performs poorly on small-scale datasets. After introducing the soft constraint, the results of the network are more accurate. Its formula is:

[0079]

[0080] Among them, x represents the pixels in the image, K(x) represents the output feature map of the previous layer, I(x) is the input foggy image, J(x) represents the restored fog-free image, and t(x) represents the transmittance.

[0081] Step 4: The loss function used in this method consists of two parts, L1 loss and perceptual loss. Among them, L1 loss is the average of the absolute errors between the predicted value and the true value. Calculation formula:

[0082]

[0083] Among them, N represents the number of pixels in the image, x represents the pixel value at a point in the defogged image, and y represents the pixel value at a point in the clear image.

[0084] To ensure the details of the restored image, perceptual loss is introduced. Perceptual loss is to use other pre-trained feature extraction networks to extract the features of the output image and the target image to make them as similar as possible. The formula for perceptual loss is:

[0085]

[0086] Among them, and Θ j (J) represent the feature maps of the VGG16 network for the input foggy image and the true fog-free image respectively, C j H j W j represents the product of the number of channels output by the feature extraction network.

[0087] VGG16 trained on ImageNet is used as the feature extraction network, and its 3rd, 5th, and 7th layer features are used to generate perceptual loss.

[0088] The final loss function is: L = L1 + 0.05L perc .

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An image dehazing method based on a convolutional neural network integrating Transformer, characterized in that, The method specifically includes the following steps: S1: Obtain the synthetic foggy dataset RESIDE as the training dataset and preprocess the dataset; S2: Construct a defogging network model: Based on the network structure of U-Net and the residual module, use the convolutional module and the Transformer module to construct an end-to-end defogging network model; The architecture of the defogging network model includes: an encoder, a decoder, and a feature fusion module; Based on the network structure of U-Net, use the residual structure to construct an end-to-end defogging network. The encoder consists of several defogging modules, and each defogging module consists of a Transformer and a convolutional neural network. The convolutional neural network extracts low-frequency information, and the Transformer extracts high-frequency information; The decoder consists of deconvolution modules to complete the reconstruction of the fog-free image; Constructing the defogging network model specifically includes the following steps: S21: The encoder contains 3 defogging modules and 2 patch-merging modules. Each defogging module consists of several convolutional sub-modules in the first half and several Transformer sub-modules in the second half. The ratio of the convolutional module to the Transformer module in each defogging module decreases layer by layer. That is, in the shallow layer of the network, the convolutional module extracts the detailed information of the image, and as the network depth increases, it is changed to the Transformer module to extract the global information of the image; S22: The role of the patch-merging module is to first partition the input into the patch partition module, that is, take every 4×4 adjacent pixels as a patch, and then flatten it in the channel direction, so that the height and width of the feature map are reduced to half of the original, and the number of channels becomes 4 times the original. Then, through a fully connected layer, the number of channels is mapped to 2 times the original number of channels to achieve the downsampling operation of the feature map; S23: The defogging module consists of 8 defogging sub-modules, and each defogging sub-module is a convolutional neural network module or a Transformer module; S24: The decoder consists of a deconvolution module and the convolutional module in the defogging module. The deconvolution module consists of a convolutional layer and a PixelShuffle upsampling layer, and its role is to double the width and height of the feature map and halve the number of channels; S25: The feature fusion module uses SKNet and the residual connection network; SKNet adds an attention mechanism to different branches, enabling the network to learn the importance of different branches, so that the feature fusion module can more effectively fuse the residual features; S3: Input the preprocessed dataset into the constructed defogging network model. During the training process, calculate the loss through the loss function, continuously iterate and update the model parameters, and finally obtain the trained defogging network model for image defogging.

2. The image defogging method based on the convolutional neural network integrating Transformer according to claim 1, characterized in that, Step S1 specifically includes the following steps: S11: Obtain a training data set, including a pair of a foggy image and a clear image, denoted as (I, J), where I, J ∈ R C×H×W , I represents the foggy image, J represents the clear image, C represents the number of channels of the image, and H and W respectively represent the height and width of the image; S12: Randomly crop the image, and the resulting images are (I crop , J crop ), where I crop represents the hazy image after cropping, and J crop represents the clear image after cropping, and I crop , J crop ∈ R C×h×w , and h and w respectively represent the height and width of the cropped image; S13: Randomly flip the image obtained in step S12 to enhance the dataset; and convert it into a tensor form as the input of the defogging network model.

3. The image dehazing method of the convolutional neural network based on the fusion Transformer according to claim 1, wherein, In step S23, each defogging sub-module is a convolutional neural network module or a Transformer module; (1) In the convolutional neural network module, in addition to ordinary convolution, it also includes SENet and PixNet. SENet is used to enable the network to learn the importance of different channel features, while PixNet is used to enable the network to learn the importance of features at different spatial points; (2) The Transformer module uses the window self-attention method to calculate self-attention; when using window self-attention, first divide the feature map into windows of size 7×7, and then perform Self-Attention on each window internally; at the same time, in order to enable information interaction between different windows, a Shifted Windows Multi-Head Self-Attention follows a normal Windows Multi-head Self-Attention; Its calculation formula is: Among them, Q represents Query, K represents key, V represents the extracted information, and B is the relative position bias; Q, K, V ∈ R b ×l×d , b is the Batch Size, d is the input dimension, and l is the number of divided windows; The self-attention mechanism first performs a dot product operation on Query and Key, then uses the SoftMax function to extract the attention ratio and multiplies it with Value to obtain the specific feature map.

4. The image dehazing method of the convolutional neural network based on the fusion Transformer according to claim 1, wherein, In step S25, the specific operation process of effectively fusing residual features is: by inputting the feature map output by upsampling and the feature map transmitted by the residual connection in the encoding stage into the feature fusion module together, so that the network will not lose shallow features when restoring the foggy image; Its formula is: [r1,r2] = SoftMax(F fc (Averpool(layer i +x))) y = r1×layer i +r2×x + b Among them, x is the output feature map of the previous layer, layer i is the feature map of the residual connection, Averpool is the global average pooling function, F fc is the linear mapping layer, y is the final output, r1 and r2 are the weight information of the feature maps x and layer i calculated by the feature fusion module, and b is the bias value.

5. The image dehazing method based on the convolutional neural network integrating Transformer according to claim 1, characterized in that, In step S2, in order to ensure that the restored image is more accurate, a soft constraint function derived from a variant of the atmospheric scattering model is introduced at the end of the network; Its formula is: Among them, x represents the pixel in the image, K(x) represents the output feature map of the previous layer, I(x) is the input foggy image, J(x) represents the restored fog-free image, and t(x) represents the transmittance.

6. The image defogging method of the convolutional neural network based on the fusion Transformer according to claim 1, wherein, In step S3, the loss function consists of two parts: L1 loss and perceptual loss. Among them, the L1 loss is the average of the absolute errors between the predicted value and the true value, and the calculation formula is: Among them, N represents the number of pixels in the image, x represents the pixel value at a point in the defogged image, and y represents the pixel value at a point in the clear image; The perceptual loss is to use other pre-trained feature extraction networks to extract the features of the output image and the target image to make them as similar as possible. The formula for the perceptual loss is: Among them, and Θ j (J) represent the feature maps of the VGG16 network for the input foggy image and the real fog-free image respectively, and C j H j W j represents the product of the number of channels output by the feature extraction network; VGG16 trained on ImageNet is used as the feature extraction network, and its 3rd, 5th, and 7th layer features are used to generate the perceptual loss; The final loss function is: L = L1 + 0.05L perc .

Citation Information

Patent Citations

  • Single image defogging network based on U-Net structure and residual network and defogging method thereof

    CN114881875A

  • Adaptive target detection method in strong / weak illumination and fog environment

    CN115375991A