An image defogging method and system based on a hybrid dual-channel attention mechanism

By employing a hybrid dual-channel attention mechanism-based image dehazing method, which utilizes multi-layer hybrid dilated convolution and attention mechanisms to process image features, the method solves the image degradation problem caused by haze, achieves efficient and simple image dehazing, and improves image recognition accuracy.

CN116167927BActive Publication Date: 2026-02-13CHINA TOWER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211470249.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-23
Publication Date
2026-02-13
Estimated Expiration
2042-11-23

AI Technical Summary

Technical Problem

Dust and smoke particles in smog cause light absorption and scattering during image capture, resulting in color degradation and decreased contrast, which affects the accuracy of image analysis and recognition. Traditional dehazing algorithms have failed to effectively solve the problem of image quality degradation after dehazing.

Method used

An image dehazing method based on a hybrid dual-channel attention mechanism is adopted. Through a feature extraction and fusion module, a K-value generation module, a hybrid dual-channel attention mechanism module, and a clear image generation module, image features are processed using multi-layer hybrid dilated convolution and attention mechanism to generate clear images.

Benefits of technology

It achieves high-quality image dehazing with simple operation, solves the problems of excessive image contrast and loss of detail information after dehazing, and improves the accuracy of image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116167927B_ABST
    Figure CN116167927B_ABST
Patent Text Reader

Abstract

The application relates to an image defogging method and system based on a mixed double-channel attention mechanism and belongs to the technical field of image processing. First, the application makes a picture pass through a feature extraction and fusion module to obtain a multi-level multi-scale fusion feature map, then makes the multi-level multi-scale feature map pass through K value generation modules for mixed hollow convolution, further extracts the features of the fusion feature map, puts the extracted feature map into a mixed double-channel attention module for processing, reobtains a new feature map, calculates the output feature map of the mixed double-channel attention module, and finally obtains a clear picture after defogging. The image defogging method based on the mixed double-channel attention mechanism solves the problems of high contrast after defogging, loss of detail information and the like of a traditional image defogging method, the generated image has the advantages of less noise, rich colors and the like, and can be applied to vehicle automatic driving, aerial image recognition and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image processing, and relates to an image defogging method and system based on a hybrid double-channel attention mechanism. BACKGROUND

[0002] Dust and floating particles such as smog in haze can absorb natural light and cause refraction and scattering of natural light. After the light is refracted and scattered, the light reflected by the target to be observed is mixed, which causes the photographed photo to be far less clear than the photo under normal conditions, and problems such as image color attenuation, contrast reduction and recognizable reduction occur, which have different degrees of influence on image analysis such as segmentation, detection, positioning and classification.

[0003] In the field of image recognition, weather factors often greatly affect the accuracy of recognition. The outdoor image captured in bad weather often has too low contrast and clarity, causing the loss of picture information, and finally causing many computer vision tasks of recognition and perception to be unable to be executed with high quality. The traditional defogging algorithm does not consider the image degradation problem after defogging, and the defogged picture often has problems of too high contrast and low image quality. SUMMARY

[0004] Therefore, the purpose of the present application is to provide an image defogging method and system based on a hybrid double-channel attention mechanism, which does not require complex operation processes and human intervention, and only needs to simply process the image.

[0005] To achieve the above purpose, the present application provides the following technical scheme:

[0006] An image defogging method based on a hybrid double-channel attention mechanism, the method comprising the following steps:

[0007] S1: inputting an input image to a feature extraction and fusion module to obtain a multi-scale fused feature map;

[0008] S2: respectively placing the obtained multi-scale fused feature maps into a K value estimation module for feature extraction, and calculating the K value corresponding to each feature map;

[0009] S3: respectively inputting the feature maps output by the K value generation module of each channel into a hybrid double-channel attention mechanism module to obtain a feature map processed by the attention mechanism;

[0010] S4: performing up-sampling on the feature map processed by the attention mechanism to restore the size to the same size as the original input image;

[0011] S5: further feature fusion is performed on the feature maps with the same size to obtain a new feature map;

[0012] S6: the feature map is input into a clear picture generation module, a picture after fog removal is generated through calculation, a clear picture is generated, and the definition of the generation formula is as follows:

[0013] J(x) = K(x) * I(x) - K(x) + b

[0014] where K(x) is the result of the K value generation module, I(x) is the matrix of the original picture, and b is a constant.

[0015] Optionally, in S1, multi-layer feature extraction and fusion technology is used to extract and fuse the features of the original picture, so that the feature fusion of the image is more sufficient.

[0016] Optionally, in S2, the K value generation module uses a multi-layer mixed hollow convolution to deepen the network depth to extract deep features of the picture.

[0017] Optionally, in S2, different hollow convolution kernels with different hollow rates are used for feature extraction in each convolution layer in the K value generation module.

[0018] Optionally, in S3, a mixed attention mechanism channel is adopted in the mixed double-channel attention module.

[0019] The image defogging system based on the mixed double-channel attention mechanism based on the method includes a feature extraction and fusion module, a K value generation module, a double-channel mixed attention mechanism module, and a clear image generation module:

[0020] The feature extraction and fusion module generates multiple scale thumbnails by downsampling the image, then uses multi-layer convolution on the thumbnail to generate corresponding feature maps, and then up-samples and fuses the feature maps to obtain multiple scale fused feature maps.

[0021] The K value generation module uses multi-layer mixed hollow convolution to further extract features of the input image, and keeps the input and output feature map sizes of each layer of hollow convolution layer unchanged.

[0022] The mixed double-channel attention mechanism module uses a mixed attention mechanism module of two parallel channel attention mechanisms and spatial attention mechanisms to process the feature map.

[0023] The clear image generation module directly generates a clear image feature matrix from the output feature matrix of the attention mechanism module through calculation.

[0024] Optionally, the feature extraction and fusion module uses multi-layer feature extraction and multi-layer feature fusion technology.

[0025] Optionally, the mixed double-channel attention mechanism module uses an attention mechanism module that mixes two parallel channel attention and spatial attention.

[0026] Optionally, the K value generation module uses more atrous convolution layers to extract features from the input picture.

[0027] Optionally, the atrous factor of the atrous convolution layer is set to keep the size of the feature map output by each layer consistent with the size of the input feature map.

[0028] The present application has the advantages that: the present application extracts and fuses features from the original picture, uses a multi-layer mixed atrous convolution method to further extract picture features after feature fusion, and uses a mixed double-channel attention mechanism module to process the output feature map of the K value generation module, solving the problems of picture background information interference in existing dehazing methods, and high contrast and loss of picture detail information after picture dehazing. The operation is simple, only the input of the picture containing fog is required, and the output is the picture after dehazing. The network model has strong fusion performance and can be fused with most image processing networks. It can be embedded in various video recognition application picture preprocessing modules and has good popularization value.

[0029] Other advantages, objects, and features of the present application will be apparent to those skilled in the art from the following specification and accompanying drawings, and will be learned from the study of the following, or will be taught from the practice of the present application. The objects and other advantages of the present application can be achieved and obtained by the following specification. BRIEF DESCRIPTION OF DRAWINGS

[0030] In order to make the objects, technical solutions and advantages of the present application clearer, the preferred detailed description of the present application will be combined with the drawings to describe the present application, wherein:

[0031] Figure 1 is a flowchart of the present application;

[0032] Figure 2 is a flowchart of the feature extraction and fusion module;

[0033] Figure 3 is a processing flowchart of the K value generation module;

[0034] Figure 4 is a flowchart of the mixed double-channel attention mechanism module. DETAILED DESCRIPTION

[0035] The present application is described herein with reference to particular embodiments thereof, which provide for a thorough and enabling disclosure of the application. The skilled person will readily understand other advantages and objects of the present application from the description of the present application given herein. The present application can be carried out in other ways than those specifically set forth herein without departing from the essential characteristics of the application. The present embodiments are therefore to be considered in all respects as illustrative and not restrictive, and all changes coming within the meaning and equivalency range of the appended claims are intended to be embraced therein.

[0036] The accompanying drawings are included to provide a further understanding of the application, and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and together with the description serve to explain the principles of the application. In the drawings:

[0037] The same or similar components in the drawings of the embodiments of the present application correspond to the same or similar components; in the description of the present application, it should be understood that the orientation or position relationship indicated by the terms "upper", "lower", "left", "right", "front", "back" and the like is based on the orientation or position relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore the terms describing the position relationship in the drawings are only for illustrative purposes, and cannot be understood as a limitation of the present application, for ordinary skilled persons in the art can understand the specific meaning of the above terms according to the specific circumstances.

[0038] The image defogging method based on the mixed double-channel attention mechanism of the present application first makes the picture pass through the feature extraction and fusion module to obtain a multi-level and multi-scale fusion feature map, then passes the multi-level and multi-scale feature map through the K value generation module to perform mixed hole convolution on the fusion feature map, further extracts the features of the fusion feature map, and puts the extracted feature map into the mixed double-channel attention module for processing to obtain a new feature map, then calculates the output feature map of the mixed double-channel attention module, and finally obtains a clear picture after defogging.

[0039] The image defogging method based on the mixed double-channel attention mechanism mainly includes a feature extraction and fusion module, a K value generation module, a mixed double-channel attention mechanism module and a clear picture generation module.

[0040] The feature extraction and fusion module generates a plurality of scale thumbnails by downsampling the image, then uses a plurality of convolution layers to generate corresponding feature maps, and then fuses the feature maps by upsampling to obtain a multi-scale fusion feature map.

[0041] The K value generation module uses multi-layer mixed dilated convolution to extract features of the input image, and keeps the input and output sizes of each layer of dilated convolution unchanged.

[0042] The dual-channel mixed attention mechanism module uses a mixed attention mechanism module of two parallel channel attention mechanisms and spatial attention mechanisms to process the feature map.

[0043] The clear image generation module directly generates a feature matrix of a clear image by calculating the output feature matrix of the attention mechanism module.

[0044] Referring to Figures 1-4 , the method specifically comprises the following steps:

[0045] S1, first, the input picture is passed through three down-sampling layers to obtain 1 / 2, 1 / 4, and 1 / 8 of the thumbnail of the original picture, and feature extraction is performed on the thumbnail of each down-sampling layer to obtain feature maps of three thumbnail images of different sizes, and the feature map of the 1 / 8 original picture is directly output as a feature Figure 1 , then the 1 / 8 thumbnail is 2 times up-sampled and fused with the 1 / 4 feature map, and the fused feature Figure 2 , then the 1 / 4 thumbnail feature map of the original picture is 2 times up-sampled and fused with the 1 / 2 thumbnail of the original picture, and the fused feature Figure 3 .

[0046] S2, the three feature maps output by the feature extraction and fusion module are input into three parallel K value generation modules, in each K value generation module, first layer dilated convolution with an expansion rate of 1 is performed on the input picture, and an activation function is used to obtain a feature map x1, then the x1 is subjected to second layer dilated convolution with an expansion rate of 2, and an activation function is used to obtain a feature map x2, the x1 and x2 are fused to obtain a fused feature map concat1, the obtained concat1 is subjected to third layer dilated convolution with an expansion rate of 3, and an activation function is used to obtain a feature map x3, then the x2 and x3 are fused to obtain a fused feature map concat2, the obtained concat2 is subjected to fourth layer dilated convolution with an expansion rate of 1, and an activation function is used to obtain a feature map x4, the x1, x2, x3, and x4 are fused to obtain a fused feature map concat3, then the concat3 is subjected to fifth layer dilated convolution with an expansion rate of 2 and an activation function to obtain a feature map x5, then the x2, x3, x4, and x5 are fused to obtain a fused feature map concat4, then the concat4 is subjected to sixth layer dilated convolution with an expansion rate of 3 and an activation function to obtain a feature map x6, and the obtained x6 is the K value of the picture.

[0047] S3, input the output of the K value generation module of each channel into a mixed double-channel attention mechanism module mixed by the channel attention mechanism module and the spatial attention mechanism module. In branch 1 of the double-channel attention mechanism, after the input feature map passes through the channel attention mechanism, the size of the feature map changes from [B, C, H, W] to [B, C, 1, 1], and the feature information is reduced, so it is necessary to multiply it with the input feature map to restore its dimension. Similarly, after passing through the spatial attention mechanism, the size of the feature map changes from [B, C, H, W] to [B, 1, H, W], and after multiplying with the input feature, the output feature map of branch 1 is obtained. After the attention of the two branches, the obtained features are added to enrich the information of the feature map, and the feature map after the attention mechanism is obtained. In branch 2 of the double-channel attention mechanism, the input feature map first passes through the spatial attention mechanism module, and the size of the feature map changes from [B, C, H, W] to [B, 1, H, W]. In order to keep the dimension of the feature map unchanged, it is necessary to multiply it with the input feature map, and then input it into the channel attention mechanism module. After passing through the channel attention mechanism module, the size of the feature map changes from [B, C, H, W] to [B, C, 1, 1], and finally multiplied with the input feature map, the output feature map of branch 2 is obtained. The feature maps of the two branches are superimposed to obtain the output feature map of the module.

[0048] S4, the feature thumbnail after the mixed double-channel attention mechanism module is upsampled to obtain a feature map with the same size as the original input.

[0049] S5, superimpose the upsampled feature map and the feature map with the original size to obtain the final feature map.

[0050] S6, the final feature map is input into the clear picture generation module, and the clear picture is finally generated after calculation according to the generation formula. The definition of the generation formula is as follows:

[0051] J(x) = K(x) * I(x) - K(x) + b

[0052] Where K(x) is the result of the K value generation module, I(x) is the matrix of the original picture, and b is a constant.

[0053] Finally, it should be pointed out that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should be covered by the claims of the present application.

Claims

1. An image defogging method based on a hybrid dual-channel attention mechanism, characterized in that: The method comprises the following steps: S1: inputting an input image into a feature extraction and fusion module to obtain a plurality of multi-scale fused feature maps; S2: placing the obtained plurality of multi-scale feature maps into a K value estimation module respectively for feature extraction, and calculating the K value corresponding to each feature map to obtain a plurality of K values; S3: inputting the feature map corresponding to each K value into a mixed double-channel attention mechanism module respectively to obtain a plurality of feature maps processed by the attention mechanism; S4: up-sampling the plurality of feature maps processed by the attention mechanism to restore the same size as the original input image to obtain a plurality of feature maps with the same size; S5: superimposing the plurality of feature maps with the same size to obtain a final feature map; S6: inputting the final feature map into a clear picture generation module as K(x) to generate a clear picture by calculation, and the definition of the generation formula is as follows: J(x) = K(x) * I(x) - K(x) + b Wherein, I(x) is the matrix of the original picture, and b is a constant.

2. The image defogging method based on the hybrid dual-channel attention mechanism according to claim 1, characterized in that: In S1, the multi-layer feature extraction and fusion technology is used to extract and fuse the features of the original image, so that the feature fusion of the image is more sufficient.

3. The image defogging method based on the hybrid dual-channel attention mechanism according to claim 2, characterized in that: In S2, the K value generation module uses multi-layer mixed hollow convolution to deepen the network depth to extract the deep features of the picture.

4. The image defogging method based on the hybrid dual-channel attention mechanism according to claim 2, characterized in that: In S2, different hollow convolution kernels with different hole rates are used for feature extraction in each convolution layer in the K value generation module.

5. The image defogging method based on the hybrid dual-channel attention mechanism according to claim 4, characterized in that: In S3, the mixed double-channel attention mechanism module adopts two parallel channel attention and spatial attention mixed attention mechanism channels.

6. The image defogging system based on the hybrid dual-channel attention mechanism based on the method of any one of claims 1-5, characterized in that: The system comprises a feature extraction and fusion module, a K value generation module, a double-channel mixed attention mechanism module, and a clear image generation module: The feature extraction and fusion module generates a plurality of scale thumbnails by down-sampling the image, then uses multi-layer convolution to generate corresponding feature maps, and then up-samples and fuses the feature maps to obtain a plurality of scale fused feature maps; The K value generation module uses multi-layer mixed hollow convolution to further extract the features of the input image, and keeps the input and output feature map sizes of each layer of hollow convolution layer unchanged; The mixed double-channel attention mechanism module uses a mixed attention mechanism module of two parallel channel attention mechanisms and spatial attention mechanisms to process the feature maps; The clear image generation module directly generates a clear image feature matrix by calculating the output feature matrix of the attention mechanism module.

7. The image defogging system based on the hybrid dual-channel attention mechanism according to claim 6, characterized in that: The feature extraction and fusion module uses multi-layer feature extraction and multi-layer feature fusion technology.

8. The image defogging system based on the hybrid dual-channel attention mechanism according to claim 6, characterized in that: The mixed double-channel attention mechanism module uses a mixed attention mechanism module of two parallel channel attention and spatial attention.

9. The image defogging system based on the hybrid dual-channel attention mechanism according to claim 6, characterized in that: The K value generation module uses more hollow convolution layers to extract the features of the input picture.

10. The image defogging system based on the hybrid dual-channel attention mechanism according to claim 6, characterized in that: The hole factor of the hollow convolution layer is set to keep the size of the output feature map consistent with the size of the input feature map.

Citation Information

Patent Citations

  • Monitoring video leaf shielding detection method based on scene depth information perception

    CN111582074A

  • End-to-end defogging method based on residual edge information fusion

    CN112508802A