An implementation method of a dual attention defogging network based on feature enhancement
Patent Information
- Application Number
- CN202410059099.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-01-15
AI Technical Summary
图像增强去雾一般不考虑雾天成像原理,通过增强图像的像素点达到去雾效果,但当雾浓度分布不均匀时,易出现光晕、噪声过大等现象,主要用于早期去雾;图像恢复去雾的方法通过分析雾天成像原理,建立雾天成像模型,利用先验知识和假设方法求解未知数,反推出无雾图像,例如何凯明等提出的基于暗通道先验的图像去雾技术,此类算法有一定的去雾效果,但过于依赖先验知识,导致模型泛化能力欠缺,且缺少多尺度结构,导致遇到复杂场景中的团雾时去雾不均匀;基于深度学习的去雾算法通过学习有参考图像,迭代更新参数,实现图像去雾,是近年来主流的去雾方法,例如基于CNN的一体化去雾网络(AOD-Net)等,此类方法去雾效果良好,但常出现网络模型复杂,参数量大的情况
[0030]1、网络模型优化,参数量显著减少:本发明运用Ghost模块代替传统卷积进行特征丰富,减少参数量,实现高效去雾;
Smart Images

Figure CN117852590B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method for implementing a dual attention dehazing network based on feature enhancement. Background Technology
[0002] In severe smog environments, atmospheric particles suspended in the air scatter light, causing both the reflected light from the target object and the ambient light incident on the imaging device to weaken simultaneously. This results in low contrast, blurred details, and image distortion in the acquired images, which degrades their quality and hinders image recognition and application. Furthermore, it negatively impacts subsequent computer vision work.
[0003] Currently, dehazing methods are mainly divided into image enhancement dehazing, image restoration dehazing, and deep learning-based dehazing. Image enhancement dehazing generally does not consider the principles of fog imaging and achieves dehazing by enhancing the pixels of the image. However, when the fog density distribution is uneven, it is prone to phenomena such as halo and excessive noise, and is mainly used for early dehazing. Image restoration dehazing methods analyze the principles of fog imaging, establish a fog imaging model, and use prior knowledge and assumptions to solve the unknowns and deduce the fog-free image. For example, the image dehazing technique based on dark channel prior proposed by He Kaiming et al. has a certain dehazing effect, but it relies too much on prior knowledge, resulting in poor model generalization ability and a lack of multi-scale structure, leading to uneven dehazing when encountering fog patches in complex scenes. Deep learning-based dehazing algorithms learn from reference images and iteratively update parameters to achieve image dehazing. It has become the mainstream dehazing method in recent years. For example, the CNN-based integrated dehazing network (AOD-Net) has good dehazing effect, but often the network model is complex and has a large number of parameters. To address the problems of poor noise handling capabilities, uneven dehazing of fog patches in complex scenes, complex network models, and large number of parameters in existing dehazing methods, this invention proposes a dual attention dehazing method based on feature enhancement. Summary of the Invention
[0004] Purpose of the invention: The purpose of this invention is to provide a dual attention dehazing network implementation method based on feature enhancement, which can restore details, clarify colors, and uniformly dehaze after processing foggy images.
[0005] Technical solution: The dual attention dehazing network implementation method of the present invention includes the following steps:
[0006] S1, Create the training dataset and the test dataset;
[0007] S2, input the training dataset into the feature-enhanced dual-attention dehazing network model;
[0008] S3, the encoder in the dual attention dehazing network model, uses convolution operations, residual group modules, and dense feature fusion modules to extract feature information from the training dataset;
[0009] S4, the dual attention feature enhancement module in the dual attention dehazing network model uses the Ghost module to enrich features and combines the improved RFB structure to capture features at different scales, and integrates channel attention mechanism and spatial attention mechanism to achieve feature self-optimization;
[0010] S5, the decoder in the dual-attention dehazing network model, uses a dense feature fusion module, an SOS Boosted structure, and convolutional operations to recover features and output the dehazed image.
[0011] Furthermore, the encoder includes a convolution operation, a first feature extraction block, a second feature extraction block, a third feature extraction block, and a fourth feature extraction block connected in sequence. The first feature extraction block, the second feature extraction block, the third feature extraction block, and the fourth feature extraction block each include a residual group module, a dense feature fusion module, and a convolutional layer connected in sequence.
[0012] The decoder includes a fourth feature recovery block, a third feature recovery block, a second feature recovery block, a first feature recovery block, and a convolution operation connected in sequence; the fourth feature recovery block, the third feature recovery block, the second feature recovery block, and the first feature recovery block each include a deconvolution operation, an SOS Boosted module, and a dense feature fusion module connected in sequence.
[0013] The dual attention feature enhancement module includes a Ghost module, an RFB structure, and a fusion channel attention and spatial attention module connected in sequence.
[0014] Furthermore, in step S3, the encoder uses convolution operations, a residual group module, and a dense feature fusion module to extract feature information from the training dataset. The detailed steps are as follows:
[0015] S31, use convolution operations to perform preliminary feature extraction on the training dataset;
[0016] S32, using the residual group module to repair the features of the training dataset;
[0017] S33, the features of the training dataset are downsampled using the convolution operation to reduce the feature size of the training dataset;
[0018] S34, the dense feature fusion module is used to fuse the extracted preliminary features with the output of each previous feature extraction block and input it into the next feature extraction block.
[0019] Furthermore, in step S4, the dual-attention feature enhancement module enriches features and reduces the number of parameters through the Ghost module, and captures features at different scales by combining the RFB structure. Then, it is fused through channel attention mechanism and spatial attention mechanism to ensure dehazing performance while enhancing the network's noise processing capability. The detailed steps are as follows:
[0020] S41, the features output by the encoder are processed by the Ghost module, which uses linear transformation instead of ordinary convolution to enrich the features. The features directly mapped and the features indirectly mapped by linear transformation are fused together, and then the fused features are fed into the RFB structure.
[0021] S42, capture detailed features through the RFB structure, and fuse the output features of the Ghost module at different scales;
[0022] S43, by fusing the channel attention mechanism and the spatial attention mechanism, the features processed by the RFB structure are weighted to achieve adaptive feature optimization and improve the network noise processing capability; then the output features are fed into the decoder.
[0023] Furthermore, in step S5, the decoder uses a dense feature fusion module, an SOS Boosted structure, and convolutional operations to recover features and output the dehazed image; the detailed implementation steps are as follows:
[0024] S51, the features output by the dual attention feature enhancement module are upsampled through deconvolution operation;
[0025] S52, through the SOS Boosted module, the output of the dual attention feature enhancement module, the output of each previous feature recovery block, and the output of the same-level feature extraction block are fused to achieve feature enhancement;
[0026] S53, The dense feature fusion module fuses the features between non-adjacent levels and inputs them into the next feature extraction block;
[0027] S54 recovers the output of the first feature recovery block through convolution operation, and outputs a fog-free image.
[0028] Furthermore, the RFB module is divided into three branches: the first branch is used for overall feature description; the second branch is used for regular scale feature extraction; and the third branch is used for capturing subtle features. The three parallel branches concatenate and stitch together the features they extract, then adjust the number of channels through convolution, and add a residual connection to introduce the initial features into the output of the multi-scale parallel branches.
[0029] Compared with the prior art, the significant advantages of this invention are as follows:
[0030] 1. Network model optimization with significantly reduced number of parameters: This invention uses the Ghost module to replace traditional convolution for feature enrichment, reducing the number of parameters and achieving efficient dehazing;
[0031] 2. For fog in complex scenes, uniform defogging can be achieved: This invention uses RFB structure to extract features at multiple scales and adjusts the receptive field size to enable the network to effectively grasp the overall structure and detailed texture of the image. It is more flexible in extracting fog map features in complex scenes and achieves more uniform defogging.
[0032] 3. Enhanced noise reduction capabilities for foggy images: The dual-channel attention module of this invention combines spatial attention and channel attention mechanisms, ensuring defogging performance while enhancing network anti-interference capabilities, resulting in superior noise reduction performance;
[0033] 4. The restored haze-free image has a more complete structure and clearer details: The dense feature fusion module in this invention closely links the feature information between non-adjacent layers of the encoder and decoder, making up for the spatial information loss in the information transmission of the traditional encoder-decoder structure, preserving the detail information of the original image set to the greatest extent, reducing information loss in feature transmission, and ensuring that the restored image is closer to the real haze-free image. Attached Figure Description
[0034] Figure 1 This is the overall flowchart of the present invention;
[0035] Figure 2 This is a schematic diagram of the feature-enhanced dual-attention dehazing network of the present invention;
[0036] Figure 3 This is a structural diagram of the dual attention feature enhancement module of the present invention;
[0037] Figure 4 This is a structural diagram of the residual group module of the present invention;
[0038] Figure 5 This is a structural diagram of the SOS Boosted module of the present invention;
[0039] Figure 6 This is a structural diagram of the dense feature fusion module of the present invention. Detailed Implementation
[0040] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0041] like Figure 1 As shown, this invention provides a method for implementing a dual-attention dehazing network based on feature enhancement, comprising the following steps:
[0042] Step 1: Create the training dataset and the test dataset;
[0043] Specifically, the training dataset consists of 12,000 pairs of reference images, in which clear images are synthesized into corresponding blurred fog images based on the fog imaging principle to form reference images; the test dataset consists of 30 pairs of reference images and 20 images without reference, in which the reference images have no duplicate images with the training dataset, and the images without reference are real outdoor fog images from Beijing.
[0044] Step 2: Input the training dataset into the feature-enhanced dual-attention dehazing network model;
[0045] like Figure 2 As shown, this is the dual attention dehazing network based on feature enhancement provided by the present invention. The network is based on an encoder-decoder structure, with a dual attention feature enhancement module embedded between the encoder and the decoder. The dual attention feature enhancement module fully captures and fuses features at multiple scales. The architecture will be introduced below in conjunction with specific modules.
[0046] Specifically, the encoder includes a convolution operation, a first feature extraction block, a second feature extraction block, a third feature extraction block, and a fourth feature extraction block, while the decoder includes a fourth feature recovery block, a third feature recovery block, a second feature recovery block, a first feature recovery block, and a convolution operation.
[0047] Specifically, the first feature extraction block, the second feature extraction block, the third feature extraction block, and the fourth feature extraction block each include a residual group module, a dense feature fusion module, and a convolutional layer. The fourth feature recovery block, the third feature recovery block, the second feature recovery block, and the first feature recovery block each include a deconvolution operation, an SOSBoosted module, and a dense feature fusion module.
[0048] Step 3: The encoder in the feature-enhanced dual attention dehazing network extracts feature information from the training dataset.
[0049] The detailed steps of the encoder in the feature-enhanced dual-attention dehazing network, which uses convolutional operations, residual group modules, and dense feature fusion modules to extract feature information from the training dataset, are as follows:
[0050] Step 31: Perform preliminary feature extraction on the training dataset using convolution operations;
[0051] Step 32: Use the residual group module to repair the features of the training dataset;
[0052] Step 33: Use convolution operations to perform downsampling to reduce the feature size of the repaired training dataset;
[0053] Step 34: Use the dense feature fusion module to fuse the extracted preliminary features with the output of each previous feature extraction block, and input the result into the next feature extraction block.
[0054] like Figure 4 As shown, the residual group module that makes up the feature extraction block in the encoder uses bottleneck residual connections to realize the repair and transfer of input features.
[0055] Specifically, when traditional methods use convolutional networks or fully connected networks to pass features, information loss and degradation often occur. To avoid feature distortion while increasing network depth, a residual group module is composed of three bottleneck structures. A residual connection is added to pass the input information directly to the output layer without passing it through the convolutional layer, thus avoiding feature loss and model degradation.
[0056] Step 4: The dual attention feature enhancement module in the feature enhancement-based dual attention dehazing network uses the Ghost module to enrich features and combines the improved RFB structure to fully capture features at different scales. It integrates channel attention and spatial attention mechanisms to achieve efficient self-optimization of features.
[0057] like Figure 3 As shown, the present invention embeds a dual attention feature enhancement module between the encoder-decoder structure to fully capture features and perform multi-scale fusion.
[0058] Specifically, the dual attention feature enhancement module includes the Ghost module, an improved RFB structure, and a fused channel attention and spatial attention module.
[0059] Specifically, the detailed steps of optimizing the network structure using the Ghost module to reduce the number of parameters, applying the improved RFB structure multi-scale fusion features, and combining channel attention and spatial attention mechanisms to enhance anti-interference and noise processing capabilities while ensuring network performance are as follows:
[0060] Step 41: The encoder output features are processed through the Ghost structure, which uses a simple linear transformation instead of traditional convolution to enrich the features. It fuses the directly mapped features and the features indirectly mapped using linear transformation to reduce the number of model parameters.
[0061] Step 42: The features fused in Step 41 are fed into the improved RFB structure. The improved RFB structure is used to capture detailed features. The features processed by the Ghost module are fused at different scales to further enhance the feature expressiveness.
[0062] Step 43: Use channel attention mechanism and spatial attention mechanism to assign weights to the features output by the RFB structure to achieve efficient adaptive optimization of features, and then input the output features into the decoder.
[0063] Specifically, the Ghost module obtains feature maps consistent with ordinary convolutions in two steps: direct mapping using a small number of convolutions (where 256 convolution kernels are normally used, only 128 are needed in this embodiment, reducing computation by half); and enriching features through simple operations, using simple linear transformations to obtain feature maps, with the number of depthwise convolutions used being only half of the normal amount. Specifically, this invention uses 1×1 convolution kernels to extract condensed features from the feature maps input to this module. Next, the linear transformation part uses parallel depthwise separable convolutions to extract features, with convolution kernels of d×d (which can be 3×3, 5×5, 7×7, etc.). Considering hardware limitations, this embodiment uses 3×3 convolution kernels. Finally, the directly mapped feature maps and the linearly transformed feature maps are concatenated. This process effectively reduces model computation and optimizes the model structure.
[0064] Specifically, the improved RFB structure of this invention enhances the network's feature extraction capability by simulating the human eye's visual perception, where objects closer to the center are more important, adjusting the receptive field size, and improving adaptability to small spaces. This structure has three branches. The first branch is used for overall feature description. It first uses a 1×1 convolution to reduce channels, then uses a 5×5 convolution to expand the receptive field. In this embodiment, two 3×3 convolutions replace the 5×5 convolution to reduce the number of parameters. A 3×3 convolution with a dilation rate of 5 is then introduced to achieve large-scale sampling coverage. The second branch is used for regular-scale feature extraction. After increasing the number of channels using a 1×1 convolution, a 3×3 convolution is used to extract features of a regular size. Finally, a 3×3 convolution with a dilation rate of 3 is introduced to further expand the receptive field, following the principle that changes in eccentricity should be consistent with changes in the receptive field. The third branch is used to capture subtle features. A 1×1 convolution focuses on local information, and a 3×3 convolution with a dilation rate of 1 is introduced to give more attention to features closer to the visual center. The three parallel branches concatenate the features they extract, and then adjust the number of channels through 1×1 convolution. In order to avoid information loss and distortion in multi-scale feature fusion, a residual connection is added to introduce the initial features into the output of the multi-scale parallel branches, so as to ensure the integrity of the features and avoid the degradation of the network's expressive ability.
[0065] Specifically, this invention, in order to more efficiently allocate feature resources, accurately recover target region images, suppress distortion caused by complex backgrounds and fog-free areas, and achieve adaptive feature optimization, fully integrates channel attention and spatial attention mechanisms. Attention weights are inferred sequentially from the channel and spatial dimensions, and then multiplied by the input feature map to obtain the optimized feature map. The channel attention mechanism passes the input feature map through width- and height-based global max-pooling and global average-pooling layers respectively, introducing global information. Then, it passes through a shared fully connected layer, adding the output features element-wise, and finally using sigmoid activation to generate channel attention weights. The spatial attention mechanism uses the feature map output from the channel attention mechanism as input for spatial attention, passes through global max-pooling and global average-pooling layers, and concatenates the results along the channel dimension. Then, a 7×7 convolutional layer reduces the dimensionality to one channel, and then a sigmoid function is used to generate spatial attention weights. Finally, the spatial attention weights are multiplied by the input features of the channel attention mechanism to obtain the final generated features.
[0066] Step 5: The decoder of the feature-enhanced dual-attention dehazing network uses a dense feature fusion module, an SOS Boosted structure, and convolutional operations to recover features and output the dehazed image. Detailed steps are as follows:
[0067] Step 51: Use deconvolution to upsample the features output by the dual attention feature enhancement module to increase the feature size.
[0068] Step 52: The SOS Boosted module is used to enhance the output of the dual attention feature enhancement module, the output of each previous feature recovery block, and the output of the same-level feature extraction block to achieve feature enhancement.
[0069] Step 53: Use the dense feature fusion module to fuse features between non-adjacent layers and input them into the next feature extraction block.
[0070] Step 54: Apply convolution operations to recover the output of the first feature recovery block, outputting a haze-free image. For example... Figure 5 As shown, the SOS Boosted module, which makes up the feature recovery block in the decoder, fuses the features input to this module with the output of the dual attention feature enhancement module and the features output of the feature extraction block of the encoder at the same level, so as to fully preserve and utilize the low-level features.
[0071] Specifically, for the feature map j of the input SOS Boosted module in the previous layer... n+1 First, the features are upsampled using a deconvolution layer, and the upsampled features (j) are then... n+1)↑2 and the feature map i obtained from the same level encoder n The sums are then added together; subsequently, feature repair is performed using the residual group module, and the output is subtracted from the feature(j). n+1 )↑2, the final result is the output of the nth layer SOS Boosted module, expressed by the formula below:
[0072]
[0073] in, This indicates that the features have been repaired by the residual group module, and (·)↑2 indicates that the features have been deconvolutioned.
[0074] like Figure 6 As shown, dense feature fusion modules are used in both the feature extraction block of the encoder and the feature recovery block of the decoder to realize information transfer between features of non-adjacent levels and make up for spatial information loss.
[0075] Specifically, in feature processing, there is often a lack of sufficient connections between non-adjacent layer features, leading to information loss. The back projection method constructs an iterative upsampling and downsampling layer feature processing approach, utilizing an iterative error feedback mechanism to calculate upsampling and downsampling errors to guide reconstruction. This invention uses a dense feature fusion module to form a feature restoration block, effectively increasing the connections between non-adjacent layer features and reducing feature loss and distortion. The dense feature fusion module mainly achieves projection through a convolutional layer with a stride of 2 and back projection through a deconvolutional layer. Through iterative projection and back projection operations, the projection error is calculated and parameters are updated to obtain better feature fusion results. The specific process is as follows:
[0076] Let J n J is the output of the SOS Boosted module in the nth layer of the decoder. L J L-1 …J n+1 The output of the dense feature fusion module in the first Ln layers of the decoder is used to enhance one feature J at a time. L-t Enhance current features J n , where t∈{0, 1…Ln-1}, such that feature J n Update. First, calculate the downsampling result after the t-th iteration. and J L-t The difference The formula is shown below:
[0077]
[0078] in, This refers to the projection operation, specifically the output of the SOS Boosted module in the nth layer during the t-th iteration. Downsampling with J L-t Same dimension, That is, the difference between the two.
[0079] Next update The formula is shown below:
[0080]
[0081] in, For the back projection operation, that is, the difference The output of the SOSBoosted module in the nth layer is obtained by upsampling through the deconvolution layer. Same dimensions, then add features This refers to the error feedback mechanism after one iteration; after calculating the downsampling and upsampling errors, the enhanced features are obtained. When t takes the values {0,1,…,Ln-1}, that is, after iterating through all the enhanced features of the first Ln layers, the final enhanced feature J is obtained. n The formula for the operation principle of the dense feature enhancement module in the decoder is shown below:
[0082] J n =D den (J n ,[J L J L-1 ...J n+1 ])
[0083] Among them, D den (·) refers to dense feature fusion performed in the decoder. This invention uses convolutional layers and deconvolutional layers to implement the corresponding backprojection and projection operations.
[0084] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment and should not be construed as limiting the scope of the invention. Those skilled in the art will understand that equivalent variations made to all or part of the processes of the above embodiments still fall within the scope of the invention.
Claims
1. A method for implementing a dual-attention dehazing network based on feature enhancement, characterized in that, Includes the following steps: S1, Create the training dataset and the test dataset; S2, input the training dataset into the feature-enhanced dual-attention dehazing network model; wherein, the dual-attention dehazing network model includes an encoder, a dual-attention feature enhancement module and a decoder; the dual-attention feature enhancement module includes a Ghost module, an RFB structure and a fusion channel attention and spatial attention module connected in sequence; S3, the encoder in the dual attention dehazing network model, uses convolution operations, residual group modules, and dense feature fusion modules to extract feature information from the training dataset; S4, the dual-attention feature enhancement module in the dual-attention dehazing network model, utilizes the Ghost module to enrich features and combines an improved RFB structure to capture features at different scales. It integrates channel attention and spatial attention mechanisms to achieve feature self-optimization, ensuring dehazing performance while enhancing the network's noise processing capabilities. The detailed steps are as follows: S41, the features output by the encoder are processed by the Ghost module, which uses linear transformation instead of ordinary convolution to enrich the features. The features directly mapped and the features indirectly mapped by linear transformation are fused together, and then the fused features are fed into the RFB structure. S42, capture detailed features through the RFB structure, and fuse the output features of the Ghost module at different scales; S5, the decoder in the dual-attention dehazing network model, uses a dense feature fusion module, an SOS Boosted structure, and convolutional operations to recover features and output the dehazed image.
2. The method for implementing a dual-attention dehazing network based on feature enhancement according to claim 1, characterized in that, The encoder includes a convolution operation, a first feature extraction block, a second feature extraction block, a third feature extraction block, and a fourth feature extraction block connected in sequence. The first feature extraction block, the second feature extraction block, the third feature extraction block, and the fourth feature extraction block each include a residual group module, a dense feature fusion module, and a convolutional layer connected in sequence. The decoder includes a fourth feature recovery block, a third feature recovery block, a second feature recovery block, a first feature recovery block, and a convolution operation connected in sequence; the fourth feature recovery block, the third feature recovery block, the second feature recovery block, and the first feature recovery block each include a deconvolution operation, an SOS Boosted module, and a dense feature fusion module connected in sequence.
3. The method for implementing a dual-attention dehazing network based on feature enhancement according to claim 2, characterized in that, In step S3, the encoder uses convolution operations, a residual group module, and a dense feature fusion module to extract feature information from the training dataset. The detailed steps are as follows: S31, use convolution operations to perform preliminary feature extraction on the training dataset; S32, using the residual group module to repair the features of the training dataset; S33, the features of the training dataset are downsampled using the convolution operation to reduce the feature size of the training dataset; S34, the dense feature fusion module is used to fuse the extracted preliminary features with the output of each previous feature extraction block and input it into the next feature extraction block.
4. The method for implementing a dual-attention dehazing network based on feature enhancement according to claim 2, characterized in that, In step S4, the features processed by the RFB structure are weighted by fusing the channel attention mechanism and the spatial attention mechanism to achieve adaptive feature optimization and improve the network's noise processing capability; then the output features are fed into the decoder.
5. The method for implementing a dual-attention dehazing network based on feature enhancement according to claim 2, characterized in that, In step S5, the decoder uses a dense feature fusion module, an SOS Boosted structure, and convolutional operations to recover features and output the dehazed image; the detailed implementation steps are as follows: S51, the features output by the dual attention feature enhancement module are upsampled through deconvolution operation; S52, through the SOS Boosted module, the output of the dual attention feature enhancement module, the output of each previous feature recovery block, and the output of the same-level feature extraction block are fused to achieve feature enhancement; S53, The dense feature fusion module fuses the features between non-adjacent levels and inputs them into the next feature extraction block; S54 recovers the output of the first feature recovery block through convolution operation, and outputs a fog-free image.
6. The method for implementing a dual-attention dehazing network based on feature enhancement according to claim 1, characterized in that, The RFB structure is divided into three branches: the first branch is used for overall feature description; the second branch is used for regular scale feature extraction; and the third branch is used to capture subtle features. The three parallel branches concatenate and stitch together the features they extract, then adjust the number of channels through convolution, and add a residual connection to introduce the initial features into the output of the multi-scale parallel branches.