An image dehazing method and system based on multi-scale information selection attention mechanism

By introducing multi-scale information selection attention mechanism and attention mechanism into the image defog removal algorithm, the multi-scale feature information of the image is extracted and fused, and the problem of existing algorithms ignoring the spatial features and channel features of different scales of the image is achieved, achieving a more efficient image defog removal effect.

CN114663309BActive Publication Date: 2025-05-13SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210289695.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-23
Publication Date
2025-05-13
Estimated Expiration
2042-03-23

AI Technical Summary

Technical Problem

When processing foggy images, existing image defogging algorithms tend to ignore spatial feature information and channel feature information at different scales of the image, resulting in unsatisfactory defogging effect.

Method used

The image defogging method based on the multi-scale information selection attention mechanism is adopted to extract the multi-scale feature information of the image through a parallel multi-scale convolutional neural network and perform feature fusion. Use spatial attention and channel attention mechanisms to design attention groups to enhance information extraction capabilities and improve the pertinence of image fog removal.

Benefits of technology

It achieves a more efficient and targeted image fog removal effect, maximizing the image structure and detail information, and improving the image fog removal performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114663309B_ABST
    Figure CN114663309B_ABST
Patent Text Reader

Abstract

The present invention discloses an image defogging method and system based on a multi-scale information selection attention mechanism, comprising: obtaining image samples with high and low frequency prior information added after preprocessing a foggy image; performing parallel multi-scale multi-layer convolution operations on the image samples using multiple convolution branches, extracting multi-scale features by inter-layer crossover, fusing the multi-scale features to obtain sample fusion features; using an attention group including multiple cascaded multi-scale feature selection attention modules, extracting and splicing multi-scale selection attention feature maps combining spatial attention and channel attention for the sample fusion features, and obtaining fusion attention features after splicing; training a defogging network according to the fusion attention features, and obtaining a fog-free image by using the trained defogging network for the foggy image to be processed. A more efficient and targeted defogging effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image defogging, and in particular to an image defogging method and system based on a multi-scale information selection attention mechanism. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Bad weather that causes visual impairment can cause great interference to the current visual system. Among them, foggy weather has attracted much attention due to its high frequency and wide impact range. Bad foggy weather can cause imaging equipment to produce poor quality images, such as severe distortion, blurring, loss of details, etc. Fog leads to poor imaging quality, which will cause certain obstacles to computer vision tasks such as target detection, tracking and segmentation, and also bring certain challenges to practical applications such as traffic monitoring, intelligent navigation, and scene surveying. Therefore, defogging foggy images and restoring image clarity are of great significance to the normal implementation of a series of subsequent computer vision tasks and the normal production and life of human beings.

[0004] The current image defogging algorithms can be mainly divided into three categories: image processing-based enhancement methods, physical model-based defogging algorithms, and deep learning-based defogging methods. Among them, the image processing-based enhancement methods use existing and mature digital image processing technologies to improve image quality and achieve image defogging. This type of method does not explore the causes of image degradation, but directly enhances the image, which may cause the image to lose some information or even introduce noise to distort the image.

[0005] The defogging algorithm based on physical models predicts atmospheric light values ​​and other parameters by building an atmospheric scattering model and using prior knowledge to achieve defogging. The defogging effect of this type of method is generally stronger than the enhancement method based on image processing, but it is too dependent on physical models and prior knowledge, and the estimation deviation of parameters will directly lead to unsatisfactory defogging effect.

[0006] The deep learning-based dehazing method mainly restores the haze-free image directly by building an end-to-end convolutional neural network. It is currently the most commonly used dehazing method. Although it improves the dehazing performance to a certain extent, it pays insufficient attention to the hazy pixel areas and important feature channel information of the image, and the dehazing effect is still unsatisfactory. Summary of the invention

[0007] In order to solve the above problems, the present invention proposes an image defogging method and system based on a multi-scale information selection attention mechanism, which takes the high and low frequency information of the image as additional priors for defogging, extracts feature information of different scales of the image through a parallel multi-scale convolutional neural network, and performs feature fusion on it. Finally, attention groups are designed based on spatial attention and channel attention mechanisms to achieve a more efficient and targeted defogging effect.

[0008] In order to achieve the above object, the present invention adopts the following technical solution:

[0009] In a first aspect, the present invention provides an image defogging method based on a multi-scale information selection attention mechanism, comprising:

[0010] After preprocessing the foggy image, image samples with added high and low frequency prior information are obtained;

[0011] Multiple convolution branches are used to perform parallel multi-scale and multi-layer convolution operations on image samples, and multi-scale features are extracted by inter-layer crossover. The multi-scale features are fused to obtain sample fusion features.

[0012] An attention group including multiple cascaded multi-scale feature selection attention modules is used to extract and splice the multi-scale selection attention feature map combining spatial attention and channel attention for the sample fusion feature, and the fusion attention feature is obtained after splicing;

[0013] The defogging network is trained according to the fused attention features, and the trained defogging network is used to obtain the fog-free image for the processed foggy image.

[0014] As an optional implementation, preprocessing of the foggy image includes: using a Laplace operator to extract high-frequency components of the foggy image, using a Gaussian filter to extract low-frequency components of the foggy image, and cascading the foggy image with the corresponding high-frequency components and low-frequency components to obtain an image sample.

[0015] As an optional implementation, the parallel multi-scale multi-layer convolution operation includes: among the multiple convolution branches, each convolution branch includes multiple convolution layers, and the multiple convolution branches perform feature extraction on image samples in parallel, and the next layer input of each branch is the output of the previous layer of the branch and the output of the previous layer of other branches, so as to extract multi-scale features.

[0016] As an optional implementation, the extraction process of the multi-scale selective attention feature map includes: using multi-layer convolution branches to perform parallel multi-scale single-layer convolution operations on the sample fusion features to extract information of different scales and splice them, extracting attention features from the spliced ​​features, and obtaining an attention feature map that combines spatial attention and channel attention; after adding the attention feature map and the sample fusion features element by element, the obtained features are repeated with multi-scale single-layer convolution operations and attention feature extraction to obtain a multi-scale selective attention feature map.

[0017] As an optional implementation, the process of attention feature extraction includes: performing global maximum pooling and global average pooling on the splicing features to obtain two channel descriptors, using one-dimensional convolution to aggregate k channel information in the neighborhood of the two channel descriptors, adding the features after the one-dimensional convolution element by element, and obtaining the channel attention feature value after the sigmoid function operation, and obtaining the input feature of the spatial attention after the channel attention feature value and the splicing feature are multiplied element by element;

[0018] The input features of the spatial attention are globally max-pooled and globally mean-pooled along the channel axis to obtain two spatial context descriptors. The two spatial context descriptors are channel-wise concatenated to obtain a valid spatial feature descriptor. The spatial context information is aggregated by dilated convolution for the valid spatial feature descriptor. The spatial attention feature value is obtained based on the spatial context information. The spatial attention feature value is element-wise multiplied with the input features of the spatial attention to obtain the attention feature map.

[0019] As an optional implementation, the splicing process of the multi-scale selective attention feature map includes: splicing the multi-scale selective attention feature maps obtained by each multi-scale feature selective attention module to obtain a fused attention feature; wherein the multi-scale selective attention feature map obtained by the previous multi-scale feature selective attention module is the input of the next multi-scale feature selective attention module.

[0020] As an optional implementation, the process of training the defogging network based on the fused attention features includes adding the fused attention features to the foggy image element by element, and then training the defogging network using L1 loss.

[0021] In a second aspect, the present invention provides an image defogging system based on a multi-scale information selection attention mechanism, comprising:

[0022] The high- and low-frequency information extraction module is configured to obtain image samples with added high- and low-frequency prior information after preprocessing the foggy image;

[0023] The multi-scale feature extraction module is configured to perform parallel multi-scale multi-layer convolution operations on the image samples using multiple convolution branches, extract multi-scale features by inter-layer crossover, and obtain sample fusion features after fusing the multi-scale features;

[0024] The attention group module is configured to use an attention group including a plurality of cascaded multi-scale feature selection attention modules to extract and splice the multi-scale selection attention feature map combining spatial attention and channel attention for the sample fusion feature, and obtain the fusion attention feature after splicing;

[0025] The defogging processing module is configured to train the defogging network according to the fused attention features, and use the trained defogging network to obtain a fog-free image for the foggy image to be processed.

[0026] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in the first aspect is performed.

[0027] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, wherein when the computer instructions are executed by a processor, the method described in the first aspect is performed.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] In order to achieve the goal of clarifying foggy images while maximally maintaining the original structure and detail information of the image, the present invention proposes an image defogging method and system based on a multi-scale information selection attention mechanism, taking the high and low frequency information of the image as additional priors for defogging, extracting feature information of different scales of the image through a parallel multi-scale multi-layer convolutional neural network, and effectively combining the feature information of different scales of the image, and finally designing an MSAB attention group based on a spatial attention mechanism and a channel attention mechanism, introducing an attention mechanism to enhance the information extraction capability, improve the attention to key areas of the image, and thus more targeted defogging, thereby improving the effect of image defogging; solving the problem that when the existing model extracts features from foggy images, it ignores the extraction and aggregation of spatial feature information of different scales of the image, which may lead to the loss of image detail information; and the existing model treats channel features and pixel features in foggy images equally, resulting in insufficient attention to foggy pixel areas and important feature channel information of the image, thereby leading to poor defogging effect.

[0030] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0032] Figure 1 A framework diagram of an image defogging method based on a multi-scale information selection attention mechanism provided in Example 1 of the present invention;

[0033] Figure 2 A structural diagram of a multi-scale feature selection attention module provided in Example 1 of the present invention;

[0034] Figure 3 A structural diagram of a feature attention module provided in Example 1 of the present invention;

[0035] Figure 4 A structural diagram of a channel attention module provided in Example 1 of the present invention;

[0036] Figure 5 This is a structural diagram of the spatial attention module provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0037] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0038] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0039] It should be noted that the terms used herein are only for describing specific embodiments, and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0040] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0041] Example 1

[0042] like Figure 1 As shown, this embodiment provides an image defogging method based on a multi-scale information selection attention mechanism, including:

[0043] S1: After preprocessing the foggy image, an image sample with high and low frequency prior information is obtained;

[0044] S2: Use multiple convolution branches to perform parallel multi-scale and multi-layer convolution operations on image samples, extract multi-scale features through inter-layer crossover, and fuse the multi-scale features to obtain sample fusion features;

[0045] S3: Using an attention group including multiple cascaded multi-scale feature selection attention modules, extract and splice the multi-scale selection attention feature map combining spatial attention and channel attention for the sample fusion feature, and obtain the fusion attention feature after splicing;

[0046] S4: The defogging network is trained based on the fused attention features, and the trained defogging network is used to obtain the fog-free image for the processed foggy image.

[0047] In step S1, the preprocessing of the foggy image includes: extracting the high-frequency component and the low-frequency component of the foggy image, and cascading the foggy image with the corresponding high-frequency component and the low-frequency component to obtain an image sample; specifically including:

[0048] S1-1: Extract high-frequency components of foggy images using Laplacian operator;

[0049] The Laplace operator is used to enhance the edge and texture of the foggy image. The Laplace operator essentially uses the second-order differential to sharpen the image, increase the difference between pixels in the neighborhood, and make the sudden change part of the image more obvious. This embodiment uses the Laplace operator kernel shown in Table 1;

[0050] Table 1 Laplacian operator kernel

[0051]

[0052]

[0053] S1-2: Use Gaussian filtering to extract the low-frequency components of the foggy image;

[0054] In order to extract low-frequency information, this embodiment performs Gaussian filtering on the foggy image to remove high-frequency details. First, a mask is set, and then the grayscale values ​​of the image in the template are weighted averaged, and then the weighted average is assigned to the central pixel of the template until all pixels of the entire foggy image are scanned.

[0055] The formula for the two-dimensional Gaussian function is as follows:

[0056]

[0057] Wherein, x and y represent the coordinate points in the template; σ is the standard deviation; in order to achieve a better blurring effect, this embodiment uses a Gaussian template with a window size of 15 and the standard deviation σ is set to 3.

[0058] S1-3: Cascade the foggy image with its corresponding high-frequency component and low-frequency component to obtain an image sample with added high- and low-frequency prior information;

[0059] Assume that the given foggy image is I, and the low-frequency component obtained after Gaussian filtering is I LF , the high frequency component obtained after Laplace operation is I HF ; Cascade the foggy image I with its corresponding low-frequency component and high-frequency component to obtain image sample I concat as follows:

[0060] I concat =I∞I LF ∞I HF

[0061] Among them, ∞ represents cascade, that is, connection in the channel direction.

[0062] This embodiment uses high and low frequency information as additional prior information, thereby being able to extract richer feature information that can effectively distinguish foggy and non-fog images.

[0063] In step S2, the image sample I with added high and low frequency prior information is concat Through parallel multi-scale multi-layer convolutional neural networks, multi-scale features are extracted and fused to obtain sample fusion features, including:

[0064] S2-1: Image sample I concat A parallel multi-scale multi-layer convolutional neural network is used to extract multi-scale features of image samples by inter-layer crossover. The parallel multi-scale multi-layer convolutional neural network includes multiple convolutional branches, each convolutional branch includes multiple convolutional layers, multiple convolutional branches perform feature extraction in parallel, and the input of the next layer of each branch is the output of the previous layer of the branch and the output of the previous layer of other branches.

[0065] This embodiment uses two convolution branches, each of which includes two convolution layers. At the same time, the convolution kernel sizes of each convolution branch are 3×3 and 5×5 respectively, and the input of the parallel multi-scale multi-layer convolutional neural network is F0. F0 passes through the first convolution layer of the two convolution branches respectively. The output of the first convolution layer is expressed as follows:

[0066] F1 3×3 =f 3×3 (F0;η0 3×3 );

[0067] F1 5×5 =f 5×5 (F0;η0 5×5 );

[0068] Among them, F1 n×n represents the convolution output of the first layer with a scale of n×n, f n×n (.) represents a convolution operation of scale n×n, η0 n×n represents the hyperparameters of the convolution of size n×n.

[0069] In order to further improve the representation ability of the network, this embodiment introduces inter-layer multi-scale information fusion technology to cross-fuse features of different scales. The formula is as follows:

[0070] F2 3×3 =f 3×3 ((F1 3×3 +F1 5×5 );η1 3×3 );

[0071] F2 5×5 =f 5×5 ((F1 5×5 +F1 3×3 );η1 5×5 );

[0072] Among them, F2 n×n represents the convolution output of the second layer with a scale of n×n, η1 n×n represents the convolution hyperparameters of the second layer with a scale of n×n.

[0073] In this embodiment, the activation functions of the above convolutional layers all use the LeaklyReLu activation function with α being 0.5.

[0074] S2-2: Fusing multi-scale features to obtain more informative sample fusion features F n-1 :

[0075] F n _1=F2 3×3 ∞F2 5×5 ;

[0076] Among them, ∞ represents the connection in the channel direction.

[0077] In step S3, the attention group is designed based on the spatial attention mechanism and the channel attention mechanism, and the attention group includes three cascaded multi-scale feature selection attention modules MSAB; Figure 2 As shown, each multi-scale feature selection attention module MSAB includes a parallel multi-scale single-layer convolution module and a feature attention module FAM; Figure 3As shown, the feature attention module FAM includes a channel attention module CAM and a spatial attention module SAM, and the channel attention module CAM and the spatial attention module SAM are combined into the feature attention module FAM in a residual connection manner.

[0078] In this embodiment, the extraction of multi-scale selective attention feature map includes the following steps:

[0079] S3-1: Through parallel multi-scale single-layer convolution modules, the sample fusion feature F n-1 Multi-layer convolution branches are used to perform parallel multi-scale single-layer convolution operations to extract feature information of different scales and perform splicing;

[0080] The parallel multi-scale single-layer convolution module of this embodiment uses two convolution branches, each of which includes a convolution layer. The convolution kernel sizes of the two convolution branches are 1×1 and 3×3 respectively. A 3×3 convolution layer is connected after the two convolution branches. The feature information of different scales is passed through the 3×3 convolution layer to obtain the splicing feature F. The formula is expressed as follows:

[0081] F=f 3×3 (f 3×3 (F n -1)∞f 1×1 (F n-1 ))

[0082] Among them, f n×n (.) represents a convolution operation of size n×n.

[0083] S3-2: Use the feature attention module to extract the attention feature of the splicing feature F, and obtain the attention feature map combining spatial attention and channel attention; specifically including:

[0084] S3-2-1: Use the channel attention module for the splicing feature F, assign different weighted information to different channel features, and obtain the channel attention feature value;

[0085] like Figure 4 As shown in the figure, in the channel attention module, for the spliced ​​feature F of size C×H×W, first, global maximum pooling and global average pooling are used to obtain two 1×1×C channel descriptors from the spatial information to represent the maximum pooling feature and the average pooling feature respectively;

[0086] Then, a one-dimensional convolution with a kernel length of k is used to aggregate the information of k channels in the neighborhood of the channel descriptor;

[0087] Finally, the two features after one-dimensional convolution are added element by element and operated by sigmoid function to obtain the channel attention feature value M c (F); the formula is as follows:

[0088]

[0089] Among them, σ represents the sigmoid function, Represents a one-dimensional convolution operation with a convolution kernel size of k; the value of k is:

[0090]

[0091] Among them, C represents the number of channels of the concatenated feature F, and odd represents the odd number closest to the value.

[0092] S3-2-2: Attention to the channel eigenvalue M c (F) Broadcast expansion is performed in the two dimensions of the space respectively, and multiplied element by element with the concatenated feature F to obtain the input feature F′ of the spatial attention module; the formula is as follows:

[0093]

[0094] in, Represents element-wise multiplication.

[0095] S3-2-3: If Figure 5 As shown in Figure 1, in the spatial attention module, for the feature map F′ with an input size of C×H×W, first, global maximum pooling and global mean pooling are performed along the channel axis to generate two different 1×H×W spatial context descriptors;

[0096] Then, the two spatial context descriptors are channel-concatenated to generate a valid spatial feature descriptor, and the effective spatial feature descriptor is subjected to dilated convolution to efficiently aggregate the spatial context information.

[0097] Finally, the spatial context information is processed through the sigmoid function to generate the spatial attention feature value M s (F); the formula is as follows:

[0098]

[0099] Among them, ∞ represents channel splicing, It indicates that the convolution kernel size is 3×3 and the dilation rate is 2.

[0100] S3-2-4: Attention to the spatial eigenvalue M s (F) Perform broadcast expansion in the two dimensions of the space and multiply it element-by-element with the feature map F′ to obtain the attention feature map F″. The formula is as follows:

[0101]

[0102] S3-3: Fusion of samples with features F n-1Add element-wise to the attention feature map F″:

[0103]

[0104] in, It means element-by-element addition;

[0105] The added feature F n-1 'Repeat the above multi-scale single-layer convolution operation and attention feature extraction operation, and finally obtain a multi-scale selected attention feature map; among them, the multi-scale single-layer convolution module selects convolution kernels of different sizes, and the convolution kernel sizes of the second two convolution branches are 3×3 and 5×5 respectively.

[0106] In this embodiment, the multi-scale selective attention feature maps of multiple cascaded multi-scale feature selection attention modules are spliced; specifically, the splicing in the channel direction is performed by residual connection, and then the spliced ​​feature F MSAB After two convolutional layers, the final fused attention feature F is obtained Attention ; Among them, the convolution kernel sizes of the two convolutional layers are 1×1 and 3×3 respectively; the formula is as follows:

[0107] F MSAB =F MSAB 1 ∞F MSAB 2 ∞F MSAB 3

[0108] Among them, F MSAB n Represents the output of the nth MSAB in the network architecture.

[0109] In step S4, the fusion attention feature F Attention It is added element by element to the original foggy image I, and the defogging network is trained using L1 loss, and finally a clear fog-free image J is output; the formula of the L1 loss function is as follows:

[0110]

[0111] Among them, J represents a haze-free image, I represents a haze-containing image, and MISA represents a dehazing network.

[0112] Example 2

[0113] This embodiment provides an image defogging system based on a multi-scale information selection attention mechanism, including:

[0114] The high- and low-frequency information extraction module is configured to obtain image samples with added high- and low-frequency prior information after preprocessing the foggy image;

[0115] The multi-scale feature extraction module is configured to perform parallel multi-scale multi-layer convolution operations on the image samples using multiple convolution branches, extract multi-scale features by inter-layer crossover, and obtain sample fusion features after fusing the multi-scale features;

[0116] The attention group module is configured to use an attention group including a plurality of cascaded multi-scale feature selection attention modules to extract and splice the multi-scale selection attention feature map combining spatial attention and channel attention for the sample fusion feature, and obtain the fusion attention feature after splicing;

[0117] The defogging processing module is configured to train the defogging network according to the fused attention features, and use the trained defogging network to obtain a fog-free image for the foggy image to be processed.

[0118] It should be noted that the above modules correspond to the steps described in Example 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.

[0119] In further embodiments, there is also provided:

[0120] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in Embodiment 1 is performed. For the sake of brevity, it will not be described in detail here.

[0121] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0122] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0123] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method described in Example 1 is completed.

[0124] The method in Example 1 can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it is not described in detail here.

[0125] Those skilled in the art will appreciate that the units, i.e., algorithm steps, of the various examples described in conjunction with this embodiment can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0126] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. An image dehazing method based on multi-scale information selection attention mechanism, characterized in that: include: After preprocessing the foggy image, image samples with added high and low frequency prior information are obtained; Multiple convolution branches are used to perform parallel multi-scale and multi-layer convolution operations on image samples, and multi-scale features are extracted by inter-layer crossover. The multi-scale features are fused to obtain sample fusion features. The parallel multi-scale multi-layer convolution operation includes: each of the multiple convolution branches includes multiple convolution layers, the multiple convolution branches extract features from the image samples in parallel, and the next layer input of each branch is the output of the previous layer of the branch and the output of the previous layer of other branches, so as to extract multi-scale features; An attention group including multiple cascaded multi-scale feature selection attention modules is used to extract and splice the multi-scale selection attention feature map combining spatial attention and channel attention for the sample fusion feature, and the fusion attention feature is obtained after splicing; The extraction process of the multi-scale selective attention feature map includes: using multi-layer convolution branches to perform parallel multi-scale single-layer convolution operations on sample fusion features to extract information of different scales and splice them, extracting attention features from the spliced ​​features, and obtaining an attention feature map combining spatial attention and channel attention; after adding the attention feature map to the sample fusion features element by element, repeating the multi-scale single-layer convolution operation and attention feature extraction on the obtained features to obtain a multi-scale selective attention feature map; The defogging network is trained according to the fused attention features, and the trained defogging network is used to obtain the fog-free image for the processed foggy image.

2. The image defogging method based on multi-scale information selection attention mechanism as claimed in claim 1, characterized in that: The preprocessing of the foggy image includes: extracting the high-frequency component of the foggy image by using the Laplace operator, extracting the low-frequency component of the foggy image by using the Gaussian filter, and obtaining the image sample by cascading the foggy image with the corresponding high-frequency component and low-frequency component.

3. The image defogging method based on multi-scale information selection attention mechanism as claimed in claim 1, characterized in that: Note that the feature extraction process includes: performing global maximum pooling and global average pooling on the spliced ​​features to obtain two channel descriptors, and using one-dimensional convolution to aggregate the two channel descriptors in the neighborhood. The one-dimensional convolution features are added element by element, and the channel attention feature value is obtained after the sigmoid function operation. The channel attention feature value is multiplied element by element with the concatenation feature to obtain the input feature of the spatial attention. The input features of the spatial attention are globally max-pooled and globally mean-pooled along the channel axis to obtain two spatial context descriptors. The two spatial context descriptors are channel-wise concatenated to obtain a valid spatial feature descriptor. The spatial context information is aggregated by dilated convolution for the valid spatial feature descriptor. The spatial attention feature value is obtained based on the spatial context information. The spatial attention feature value is element-wise multiplied with the input features of the spatial attention to obtain the attention feature map.

4. The image defogging method based on multi-scale information selection attention mechanism as claimed in claim 1, characterized in that: The splicing process of the multi-scale selective attention feature map includes: splicing the multi-scale selective attention feature maps obtained by each multi-scale feature selective attention module to obtain a fused attention feature; wherein, the multi-scale selective attention feature map obtained by the previous multi-scale feature selective attention module is the input of the next multi-scale feature selective attention module.

5. The image defogging method based on multi-scale information selection attention mechanism as claimed in claim 1, characterized in that: The process of training the dehazing network based on the fused attention features includes adding the fused attention features to the hazy image element-wise and then training the dehazing network using the L1 loss.

6. An image defogging system based on multi-scale information selection attention mechanism, characterized in that: include: The high- and low-frequency information extraction module is configured to obtain image samples with added high- and low-frequency prior information after preprocessing the foggy image; The multi-scale feature extraction module is configured to perform parallel multi-scale multi-layer convolution operations on the image samples using multiple convolution branches, extract multi-scale features by inter-layer crossover, and obtain sample fusion features after fusing the multi-scale features; The parallel multi-scale multi-layer convolution operation includes: each of the multiple convolution branches includes multiple convolution layers, the multiple convolution branches extract features from the image samples in parallel, and the next layer input of each branch is the output of the previous layer of the branch and the output of the previous layer of other branches, so as to extract multi-scale features; The attention group module is configured to use an attention group including a plurality of cascaded multi-scale feature selection attention modules to extract and splice the multi-scale selection attention feature map combining spatial attention and channel attention for the sample fusion feature, and obtain the fusion attention feature after splicing; The extraction process of the multi-scale selective attention feature map includes: using multi-layer convolution branches to perform parallel multi-scale single-layer convolution operations on sample fusion features to extract information of different scales and splice them, extracting attention features from the spliced ​​features, and obtaining an attention feature map combining spatial attention and channel attention; after adding the attention feature map to the sample fusion features element by element, repeating the multi-scale single-layer convolution operation and attention feature extraction on the obtained features to obtain a multi-scale selective attention feature map; The defogging processing module is configured to train the defogging network according to the fused attention features, and use the trained defogging network to obtain a fog-free image for the foggy image to be processed.

7. An electronic device, characterized in that: The method comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 5 is completed.

8. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • A Multi-Scale Feature Fusion Network based on GANs for Haze Removal

    AU2020100274A4

  • Image defogging method and system based on global feature fusion attention network

    CN113344806A