Image Blind Deblurring Method with Local Double-Domain Mutual Attention and Global Channel Self-Attention
The hybrid network architecture of WFMAN and GSCAN with FFGF effectively addresses the limitations of existing deblurring methods by integrating spatial and frequency domain features, resulting in improved image clarity and efficiency suitable for mobile devices.
Patent Information
- Application Number
- CN202411616510.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-11-13
AI Technical Summary
The existing image defuzzing method based on convolutional neural networks and visual Transformer has problems such as limited receptive field, high computational complexity, overfitting and excessive smoothing, which limits its applicability in image recovery.
The image blind defuzzing method of local dual-domain mutual attention and global channel self-attention is adopted. By constructing the window's frequency domain-space mutual attention network (WFMAN) and global airspace channel self-attention network (GSCAN), combined with the frequency domain filter-gated feature fusion module (FFGF), it realizes efficient fusion and defuzzing of features.
While maintaining the model's simplicity architecture, the image clarity is significantly improved, the peak signal-to-noise ratio (PSNR) is increased by 0.20dB-2.19dB, and the structural similarity (SSIM) is increased by 0.004-0.02, which is suitable for mobile devices.
Smart Images

Figure CN119579458B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and relates to an image blind deblurring method based on local double-domain mutual attention and global channel self-attention. Background Art
[0002] Existing dynamic scene blind deblurring methods are mainly divided into two categories: deep learning methods based on convolutional neural networks and deep learning methods based on vision transformers.
[0003] In recent years, convolutional neural networks have been widely used in the field of image deblurring due to their powerful representation ability. Xu et al. first introduced deep learning into non-blind deblurring of blurred images and proposed a 5-layer convolutional neural network (CNN) to achieve non-blind deblurring. Tao et al. started from the coarse scale of the blurred image and gradually restored the latent image at a higher resolution, proposing a scale-recurrent network (SRN). Cho et al. used a single U-shaped network to simulate a multi-cascade U-shaped network and introduced asymmetric feature fusion, proposing a multi-input multi-output U-net (MIMO-Unet). Chen et al. decomposed existing methods with good effects and extracted their basic components, and at the same time removed or replaced the non-linear activation function, proposing a simple non-linear activation free network (NAFNet). Mao et al. combined the Fourier transform, ReLU activation function, and inverse Fourier transform into a traditional convolutional residual block, proposing a new residual FFT-ReLU block to achieve the fusion of the spatial domain and the frequency domain.
[0004] The Transformer architecture has been widely applied in the field of image deblurring in recent years due to its large receptive field and achieved good results. Chen et al. were the first to apply the vision Transformer architecture to image restoration, proposing (Image Processing Transformer: IPT), which achieved good results. Inspired by U-Net, Wang et al. proposed a U-shaped image transformer architecture suitable for various image restoration tasks. Zamir et al. applied the self-attention mechanism across feature dimensions rather than spatial dimensions and introduced a gating mechanism in the feed-forward neural network to filter out favorable features. Li et al. proposed a new network architecture for global, regional, and local modeling by fusing three attention mechanisms such as the anchor-based stripe self-attention mechanism. Kong et al. proposed a Fourier Transformer based on the convolution theorem, which transforms the self-attention matrix multiplication into a point multiplication in the matrix frequency domain, and introduces a learnable weight in the frequency domain to screen the frequency domain features of the blurred image.
[0005] However, although the deblurring methods based on convolutional neural networks have achieved some success in image restoration, they still face challenges such as limited receptive field, easy overfitting, vanishing gradients, and limited ability to handle spatially variant blur. While the methods based on vision Transformers improve the deblurring effect by capturing the long-range dependencies of images through the self-attention mechanism, they usually require more training data, have high computational complexity and large number of parameters, and may lead to the problem of over-smoothing, limiting their applicability.
[0006] In summary, in the methods of image deblurring based on deep neural networks, convolutional operations are the basic operations of these methods. However, convolutional operation is a spatially invariant local operation with a very limited receptive field. Even though many methods use larger and deeper models to make up for the limitations of convolution, simply increasing the capacity of the deep model does not always bring better performance. On the other hand, although the Transformer architecture can make up for the problem of limited receptive field of convolutional operations, the calculation of its scaled dot-product attention leads to quadratic space and time complexity, greatly limiting its scope of use. Although there are currently some optimization methods for the Transformer architecture, this may affect the deblurring performance. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide an image blind deblurring method based on local dual-domain mutual attention and global channel self-attention.
[0008] To achieve the above purpose, the present invention provides the following technical solutions:
[0009] An image blind deblurring method based on local double-domain mutual attention and global channel self-attention, comprising the following steps:
[0010] S1: Construct a window-based frequency-spatial mutual attention network (WFMAN), which has two branches in the spatial domain and the frequency domain. The features from the frequency domain are activated by a Gaussian error linear unit and multiplied by the features in the spatial domain as weights to achieve feature fusion of the mutual attention mechanism.
[0011] S2: Construct a global spatial channel self-attention network (GSCAN), which has only one path in the spatial domain. After passing through the global channel attention mechanism, the features from itself are multiplied by the weights from itself to perform feature fusion of the self-attention mechanism.
[0012] S3: Construct a frequency-domain filter gating feature fusion module (FFGF). By introducing a learnable quantization matrix W, adaptively determine the shallow information from the encoder to be retained, and fully fuse the information between the encoder and decoder through a simple gating operation.
[0013] S4: Connect WFMAN and GSCAN in series to form a basic module; use multiple basic modules, FFGF modules and refinement layers to construct a new U-shaped network.
[0014] S5: Use the image data sets before and after deblurring to train and test the new U-shaped network to obtain an image blind deblurring model.
[0015] S6: Use the image blind deblurring model to deblur the acquired images.
[0016] Further, the window-based frequency-spatial mutual attention network WFMAN described in step S1 specifically includes:
[0017] For the input of the i-th encoder First, perform a layer normalization operation on it, and then divide it into upper and lower branches. Each branch includes a feature extraction stage and a feature interaction stage.
[0018] The upper branch is used for operations in the spatial domain, and its feature extraction stage is described as follows: First, perform a 1×1 convolution and a 3×3 depthwise convolution on the features in sequence; then perform a SimpleGate operation, split the input features into two equal parts along the channel dimension, and directly perform an arithmetic dot product on these two parts; then perform a simplified channel attention mechanism on the features, first perform a global average pooling operation on the features, then perform a 1×1 convolution operation, and finally multiply it pointwise to the original features; then aggregate the features through a 1×1 convolution operation, and finally perform a residual connection operation to obtain the output Feature of the feature extraction part of the upper branch top :
[0019]
[0020] where represents the input of the i-th encoder; LN() represents the layer normalization operation; Convlxl() represents the convolution operation with a convolution kernel size of 1×1; DConv3x3() represents the depthwise convolution operation with a convolution kernel size of 3×3; SG() represents the SimpleGate mechanism operation; SCA() represents the simple channel attention mechanism operation
[0021] The lower branch is used for operations in the frequency domain, and its feature extraction stage is described as follows: First, divide the features into windows; then perform Fourier transform, linear layer operation, GELU non-linear activation, and inverse Fourier transform on the window features to obtain effective information about the blur pattern; finally, convert the window features after Fourier transform into features of the original size to obtain the output Feature of the feature extraction part of the WFMAN lower branch bot :
[0022]
[0023] where FFT() represents the two-dimensional fast Fourier transform operation; Linear() represents the linear transformation operation in the complex domain; GELU() represents the Gaussian error linear unit activation function; IFFT() represents the two-dimensional fast inverse Fourier transform operation
[0024] In the feature interaction stage of the upper branch, perform a layer normalization operation and a linear mapping operation on the upper branch feature Feature top ; In the feature interaction stage of the lower branch, perform operations on the lower branch feature Feature botPerform layer normalization operation; then perform a 1×1 convolution and a 3×3 depthwise convolution; then perform GELU non-linear activation, take the features of the lower branch as weights, and dot-multiply them to the features of the upper branch to achieve cross-attention operation, so as to learn the frequency-space dual-domain representation to utilize kernel-level and pixel-level features for deblurring; finally, perform a 1×1 convolution operation and a residual connection operation to obtain the output WFMA of WFMA out :
[0025] WFMA out = Conv1x1{Linear(Norm(Feature top ))
[0026] ⊙ GELU(Dconv3x3(Conv1x1(Norm(Feature bot )))) + Feature top
[0027] where Feature top represents the output of the upper branch feature extraction part of WFMA; Feature bot represents the output of the lower branch feature extraction part of WFMA; Linear represents the linear transformation operation in the real number field.
[0028] Furthermore, the global spatial-channel self-attention network GSCAN described in step S2 specifically includes:
[0029] Feature extraction: Take the output WFMA of WFMA out as the input GSCAN of GSCAN in , first perform layer normalization operation; then sequentially perform a 1×1 convolution and a 3×3 depthwise convolution; then perform a simple gating SimpleGate operation, split the input features along the channel dimension into two equal parts, and directly perform arithmetic dot-multiplication on these two parts; then perform a simplified channel attention mechanism on the features, first perform a global average pooling operation on the features, then perform a 1×1 convolution operation, and finally dot-multiply it to the original features; then aggregate the features through a 1×1 convolution operation, and finally perform a residual connection operation to obtain the output Feature of the GSCAN feature extraction part GSCAN :
[0030] Feature GSCAN = Conv1x1(SCA(SG(DConv3x3(Conv1x1(Norm(GSCAN in )))))) + GSCAN in
[0031] Feature Fusion: Input Feature GSCAN into the two input ends of the feature fusion part respectively. For the input Feature GSCAN , perform layer normalization and linear mapping operations on the upper branch; on the lower branch, perform layer normalization, 1×1 convolution, 3×3 depthwise convolution, and GELU activation operations; then multiply the weights from the lower branch into the upper branch to complete the self-to-self fusion of features and achieve the effect of self-attention; finally, perform 1×1 convolution and residual connection to obtain the output GSCAN of GSCAN 0ut :
[0032] GSCAN out = Conv1x1{Linear(Norm(feature GSCAN ))
[0033] ⊙GELU(Dconv3x3(Conv1x1(Norm(Feature GSCAN ))))}+Feature GSCAN .
[0034] Furthermore, the frequency-domain filtering gated feature fusion module FFGF described in step S3 specifically includes:
[0035] For the output of the i-th layer encoder , first perform 1×1 convolution on it, then expand the feature by PATCH, then perform Fourier transform to transform the feature into the frequency domain, then multiply by a quantization matrix W to adaptively select useful encoder information, and finally perform inverse Fourier transform and PATCH folding operations to change the feature to the size when input;
[0036] Then input the corresponding decoder layer feature with deep information, and perform channel-wise concatenation operation with the filtered encoder layer feature, then use 1×1 convolution to integrate the concatenated information, and finally use SimpleGate to achieve effective fusion of features to obtain the output FFGF of FFGF out :
[0037]
[0038] where represents the output of the i-th layer encoder, represents the output of the (i + 1)-th layer decoder; W represents the quantization matrix; Cat represents the concatenation operation along the channel dimension.
[0039] Furthermore, the definition of the basic module formed by cascading WFMAN and GSCAN described in step S4 is as follows:
[0040]
[0041] wherein represents the input of the i-th layer encoder or decoder, represents the output of the i-th layer encoder or decoder; WFMAN represents the window-based frequency-domain and spatial-domain mutual attention network; GSCAN represents the global spatial-channel self-attention network.
[0042] Furthermore, the novel U-shaped network described in step S4 includes a 3×3 convolution that maps channels to 64, three encoder layers, two downsampling operations, two upsampling operations, two FFGFs, three decoder layers, a refinement layer, and a 3×3 convolution that remaps channels to 3; wherein:
[0043] The first encoder layer EB1, the second encoder layer EB2, the first decoder layer DB1, and the second decoder layer DB2 are all composed of a basic module;
[0044] The third encoder layer EB3 and the third decoder layer DB3 are both composed of 4 identical basic modules stacked together;
[0045] The refinement layer is composed of 2 identical basic modules stacked together;
[0046] The input of the second decoder layer
[0047] The input of the first decoder layer
[0048] The downsampling operation includes a bilinear interpolation with a scale of 0.5 and a 3×3 convolution that doubles the number of channels;
[0049] The upsampling operation includes a bilinear interpolation with a scale of 2 and a 3×3 convolution that halves the number of channels.
[0050] The beneficial effects of the present invention are as follows: Compared with the methods that have achieved remarkable results in the field of blind deblurring of dynamic scene images in recent years, the method proposed by the present invention can reconstruct clearer image quality while maintaining a simple architecture. Both objective evaluation metrics and subjective visual effects have been significantly improved. Among them, the Peak Signal to Noise Ratio (PSNR) has increased by an average of 0.20 dB - 2.19 dB, and the Structural Similarity (SSIM) has increased by an average of 0.004 - 0.02.
[0051] Other advantages, objectives, and features of the present invention will, to some extent, be elaborated in the subsequent description, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following description. Brief Description of the Drawings
[0052] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0053] Figure 1 It is a schematic diagram of the window-based frequency-domain to spatial-domain mutual attention network WFMAN structure;
[0054] Figure 2 It is a schematic diagram of the global spatial-domain channel self-attention network GSCAN structure;
[0055] Figure 3 It is a schematic diagram of the frequency-domain filtering gating feature fusion module FFGF structure;
[0056] Figure 4 It is a schematic diagram of the novel U-shaped network structure;
[0057] Figures 5 - 9 They are all comparison diagrams of the results of the present invention and existing methods. Detailed Embodiments
[0058] The following illustrates the embodiments of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0059] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0060] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0061] The present invention provides an image blind deblurring method based on local dual-domain mutual attention and global channel self-attention. The present invention first proposes a novel U-shaped network, which is composed of multiple basic blocks, and each basic block is composed of two sub-networks with specific functions connected in series. The network achieves strong image deblurring performance while maintaining a small number of parameters and low computational complexity. In order to enable each basic module in the model to extract image features as much as possible to achieve a good restoration effect, a window-based frequency-spatial mutual attention network (WFMAN) and a global spatial channel self-attention network (GSCAN) are respectively proposed. By connecting these two sub-networks in series to form the basic module of the entire model, a better feature extraction effect is achieved. Among them, the window-based frequency-spatial mutual attention network (WFMAN) has two branches in the spatial domain and the frequency domain: the features from the frequency domain will be activated by a Gaussian error linear unit (GELU), and then used as weights to multiply the features in the spatial domain to achieve feature fusion of the mutual attention mechanism. The global spatial channel self-attention network (GSCAN) has only one path in the spatial domain. After passing through the global channel attention mechanism, the network will perform feature fusion of the self-attention mechanism in the form of multiplying the features from itself by the weights from itself. Finally, in order to better fuse the shallow features from the encoder and the deep features from the decoder, a simple and efficient feature fusion mechanism - frequency-domain filter gating feature fusion (FFGF) is designed. It adaptively controls the encoder features after Fourier transform through matrix W, and splices this feature with the output of the decoder in the channel dimension, and finally performs a simplified gating mechanism operation to achieve a good fusion effect.
[0062] The window-based frequency-spatial mutual attention network (WFMAN) proposed by the present invention is as Figure 1As shown. Currently, in the field of image deblurring, existing methods often focus on the extraction of spatial domain features while ignoring the potential relationship between blurred and clear image pairs in the frequency domain. However, frequency domain information is crucial for understanding the degree and pattern of image blurring because it can reveal the frequency distribution of the image and the characteristics of the blur kernel. In addition, spatial domain information is also important for high-quality image restoration. However, existing spatial domain-based feature extraction methods, such as traditional convolutional neural networks (CNNs) or Transformer networks, usually suffer from limited receptive fields or high computational complexity. To solve these problems, the present invention first designs a channel attention mechanism that reduces computational complexity while maintaining a global receptive field. Then, a window-based frequency domain feature extraction method is designed by dividing the input features into different small windows and performing frequency domain feature extraction on the windows. Finally, a feature fusion strategy similar to mutual attention is designed to interact the frequency domain information and spatial domain information in a new way.
[0063] As Figure 1 shown, the window-based frequency-spatial mutual attention network (WFMAN) proposed by the present invention can be divided into upper and lower branches, which respectively represent features from the spatial domain and features from the frequency domain. In Figure 1 the pink background part in is the feature extraction part, while the green background part is the feature interaction part. For the input of the i-th encoder, first, a layer normalization operation (Layer Normalization: LN) needs to be performed on it to eliminate the scale differences between different channels and promote the stability of training. After the layer normalization operation, the network is divided into upper and lower branches, which will be introduced in detail below.
[0064] For the upper branch, which operates in the spatial domain, it is necessary to perform a 1×1 convolution and a 3×3 depthwise convolution on the features in sequence to enhance the feature representation ability. Then, perform the SimpleGate (SG) operation. By splitting the input features into two equal parts along the channel dimension and directly performing arithmetic dot multiplication on these two parts, not only self-fusion of features is achieved, but also a non-linear operation is introduced. Next, a channel attention mechanism needs to be applied to the features. Since the premise is to maintain low complexity and low parameter count, a simple channel attention mechanism (Simple Channel Attention: SCA) is selected. Traditional channel attention compresses spatial information into channels, then applies a multi-layer perceptron to calculate the channel attention and uses it to weight the feature map. Through observation, it can be found that channel attention can be regarded as a special gating mechanism, thus simplifying channel attention and only retaining the two most important functions: aggregating global information and channel information interaction. Choose to first perform global average pooling on the input features, then follow it with a 1×1 convolution operation, and finally multiply it pointwise to the original features. By introducing the simplified channel attention operation, not only global information is captured, but also the computational efficiency is high. After completing the channel attention operation, a 1×1 convolution operation will be performed next, with the purpose of aggregating features, and finally a residual connection operation. So far, the feature extraction part of the upper branch has been introduced, and finally the output Feature of the feature extraction part of the upper branch of WFMAN can be obtained top :
[0065]
[0066] where represents the input of the i-th encoder; LN() represents the layer normalization operation; Convlxl() represents the convolution operation with a convolution kernel size of 1×1; DConv3x3() represents the depthwise convolution operation with a convolution kernel size of 3×3; SG() represents the simple gating mechanism operation; SCA() represents the simple channel attention mechanism operation
[0067] For the lower branch, which operates in the frequency domain, first, it is necessary to perform the operation of dividing the window on the features. Because after dividing the window, not only can the complexity be reduced, but also the local features can be focused, forming an interaction with the global features of the upper branch. Then, perform Fourier transform, linear layer operation, GELU non-linear activation, and inverse Fourier transform on the window features to obtain effective information about the blur pattern, so as to facilitate image deblurring using kernel-level information. Finally, it is necessary to convert the window features after Fourier transform into features of the original size. So far, the feature extraction part of the lower branch has been introduced, and finally the output Feature of the feature extraction part of the lower branch of WFMAN can be obtained bot :
[0068]
[0069] Among them, FFT() represents a two-dimensional fast Fourier transform operation; Linear() represents a linear transformation operation in the complex number domain; GELU() represents a Gaussian error linear unit activation function; IFFT() represents a two-dimensional fast inverse Fourier transform operation.
[0070] Finally, it is necessary to use a feedforward neural network to fuse the features (Feature top ) of the upper branch with the features (Feature bot ). To better fuse the features of the upper and lower branches, the present invention designs a network with two inputs and one output. For the features of the upper branch from the spatial domain, layer normalization operation and linear mapping operation are performed on them; while for the features of the lower branch from the frequency domain, layer normalization operation also needs to be performed on them first, and then a 1×1 convolution and a 3×3 depthwise convolution operation need to be performed successively. Then, GELU non-linear activation needs to be performed on the features, aiming to regard the features of the lower branch as weights, so that they can be dot-multiplied to the features of the upper branch at the end, thus completing an operation similar to mutual attention. Through such an operation, frequency-space dual-domain representation can be learned to utilize kernel-level and pixel-level features for deblurring. Finally, 1×1 convolution operation and residual connection operation also need to be performed. Finally, the output of WFMAN, WFMAN out :
[0071]
[0072] Among them, Feature top represents the output of the upper branch feature extraction part of WFMAN; Feature bot represents the output of the lower branch feature extraction part of WFMAN; Linear represents a linear transformation operation in the real number domain.
[0073] In the window-based frequency-domain to spatial-domain mutual attention network (WFMAN) proposed by the present invention, by separately extracting the features of the spatial domain and the frequency domain, and then using an operation similar to the mutual attention mechanism for feature fusion, the final output is obtained. Through the WFMAN of the present invention, not only kernel-level and pixel-level features are utilized for image deblurring through frequency-space dual-domain representation, but also the main modules in the model are simplified, ensuring the simplicity and efficiency of the model.
[0074] Global features are very important for understanding the overall structure and content of an image, and help to maintain the consistency and naturalness of the image during the deblurring process. If only the window-based frequency-domain to spatial-domain mutual attention network (WFMAN) is used as the basic module of the model, then the model will focus too much on local features and ignore global features. Therefore, to make up for the defect of WFMAN in global information, a global spatial-channel self-attention network (GSCAN) is designed to be cascaded with WFMAN, so that features of the image can be extracted both locally and globally, ensuring that the model fully mines the blurred information.
[0075] As Figure 2 shown, the present invention proposes a global spatial-channel self-attention network (GSCAN), which only contains one branch from the spatial domain. First, the input features will be processed for feature extraction in the brown background area of the figure, and then the features will be respectively input into the two input ends of the feature interaction part in the green background area, and finally the output of GSCAN is obtained. Such a design ensures that the network can extract sufficient global information and forms a complement to the local information that WFMAN focuses on.
[0076] For the output of WFMAN, WFMAN out , which is also the input of GSCAN, GSCAN in , first needs to perform the operation of layer normalization (Norm), and then sequentially perform a 1×1 convolution and a 3×3 depthwise convolution to enhance the expression ability of the features. Then comes the operation of SimpleGate (SG) and the operation of simplified channel attention (SCA), as well as 1×1 convolution and residual connection operations. Finally, the output Feature of the GSCAN feature extraction part can be obtained GSCAN :
[0077] Feature GSCAN
[0078] =Conv1x1(SCA(SG(DConv3x3(Conv1x1(Norm(GSCAN in ))))))+GSCAN in (4)
[0079] After the feature extraction operation, the features need to be fused and enhanced. Different from WFMAN, there is only one path in the spatial domain in GSCAN. Therefore, Feature GSCAN is respectively input into the two input ends of the feature fusion part to make the features fuse with themselves, achieving the effect of self-attention. For the input Feature GSCAN, perform layer normalization and linear mapping operations on the upper branch; while on the lower branch, perform layer normalization, 1×1 convolution, 3×3 depthwise convolution, and GELU activation operations. Then multiply the weights from the lower branch into the upper branch to complete the self-to-self feature fusion. Finally, perform 1×1 convolution and residual connection to obtain the output of GSCAN, GSCAN out :
[0080]
[0081] In the global spatial-channel self-attention network (GSCAN) proposed by the present invention, emphasis is placed on the extraction and fusion of global information. Feature extraction is only performed in the spatial domain, and self-to-self fusion is carried out. GSCAN and the previously introduced WFMAN form a complementarity of local and global information. By concatenating these two sub-networks, a basic module (BasicBlock) of the entire model is constructed, which can well complete the deblurring operation.
[0082] In the deblurring network using a U-shaped structure, the feature fusion between the encoder layer and the corresponding decoder layer is an operation that cannot be ignored. However, some existing methods often only use simple addition operations or concatenation operations, and do not achieve good fusion effects. Since the features of the decoder layer often have more fine-grained information, while the features of the encoder layer are often more primitive and rough, not all encoder information and decoder information contribute to the restoration of the potential clear image. If they are simply added or concatenated, some effective information cannot be utilized, resulting in a deterioration of the deblurring effect. To solve the above problems, the present invention proposes a frequency-domain filtering gating feature fusion strategy. By introducing a learnable quantization matrix W, it can adaptively determine which shallow information from the encoder should be retained, and a SimpleGate (SG) operation is also added to fully fuse the information between the encoder and decoder.
[0083] As Figure 3 shown, for the output of the i-th layer encoder First, perform 1×1 convolution on it, and then expand the feature by PATCH. Here, the PATCH size is set to 8. Next, perform Fourier transform on the feature to transform the feature into the frequency domain, then multiply by a quantization matrix W to adaptively select useful encoder information, and finally perform inverse Fourier transform and PATCH folding operations to change the feature to the size at the input. At this time, the screening of the effective information of the encoder has been completed. Next, the features of the corresponding decoder layer with deep information Perform the input and perform a channel-wise concatenation operation (Concatenate: Cat) with the filtered encoder layer features, then use a 1×1 convolution to integrate the concatenated information, and finally use SimpleGate to achieve effective feature fusion. Finally, the output FFGF of FFGF can be obtained. out :
[0084]
[0085] Among them represents the output of the i-th layer encoder, represents the output of the (i + 1)-th layer decoder; W represents the quantization matrix; Cat represents the concatenation operation along the channel dimension.
[0086] In the frequency-domain filtering gated feature fusion strategy (FFGF) proposed in the present invention, a learnable quantization matrix W is used to adaptively screen useful encoder layer features, and after concatenating them with decoder layer features in the channel dimension, a simple gating operation is performed, fully integrating the effective information between the encoder and decoder. FFGF not only achieves high-performance feature fusion, but also does not add particularly complex operations, only increasing a small part of the computational complexity, ensuring the lightweight of the entire model.
[0087] In previous deep learning-based methods, the strategy from coarse to fine has been proven effective for deblurring. However, in order to pursue excellent performance, these methods often achieve it by stacking the number of network layers, depth, and module quantity, which not only leads to a huge number of parameters but is not necessarily effective. With the development of intelligent devices, more and more mobile devices, such as smartphones, cameras, and drones, are urgently in need of image deblurring algorithms to improve image quality. However, due to the limited computing power of mobile devices, some large models with excellent performance are not suitable for deployment on mobile devices. Therefore, it is particularly crucial to design a deblurring model that is both lightweight and efficient.
[0088] The overall model architecture proposed in the present invention is as Figure 4 shown. A U-shaped network structure of sub-network concatenation type is proposed in the present invention. A single-layer U-net architecture is selected and only two downsampling and upsampling operations are performed to ensure the lightweight of the network. In addition, in a concatenated manner, the proposed WFMAN and GSCAN are concatenated as the basic modules of the model, so that local and global information can be fully utilized, making the network reach an excellent level in performance. Finally, a refinement layer is added to the network, which can further restore the details of the image. Such a design not only meets the requirements of mobile devices for computing efficiency, but also ensures the efficiency and accuracy of the deblurring task.
[0089] As Figure 4 shown, EBi represents the i-th encoder, represents the input of the i-th encoder, represents the output of the i-th encoder. DB i represents the i-th decoder, represents the input of the i-th decoder, represents the output of the i-th decoder, where i ∈ {1, 2, 3}. BasicBlock represents the basic module of the network, which is obtained by connecting WFMAN and GSCAN in series. The definition of BasicBlock is as follows:
[0090]
[0091] where represents the input of the i-th layer encoder or decoder, represents the output of the i-th layer encoder or decoder; WFMAN represents the window-based frequency-domain and spatial-domain mutual attention network; GSCAN represents the global spatial-channel self-attention network.
[0092] The entire U-shaped network includes a 3×3 convolution that maps channels to 64, three encoder layers, two downsampling operations, two upsampling operations, two FFGFs, three decoder layers, a refinement layer, and a 3×3 convolution that remaps channels to 3. EB1, EB2, DB1, and DB2 are all composed of one BasicBlock, while EB3 and DB3 are both composed of four consecutive BasicBlocks stacked together, and the refinement layer is composed of two BasicBlocks stacked together. Among them The downsampling operation includes a bilinear interpolation with a scale of 0.5 and a 3×3 convolution that doubles the number of channels; the upsampling operation includes a bilinear interpolation with a scale of 2 and a 3×3 convolution that halves the number of channels.
[0093] Finally, by combining the WFMAN, GSCAN, FFGF, and the sub-network series-connected U-shaped network architecture proposed in the present invention, the final network model is obtained. This model not only meets the requirements of lightweight and can be deployed on mobile devices, but also achieves excellent deblurring effects and has strong application value.
[0094] Experimental verification:
[0095] Ablation experiment: The present invention uses peak signal-to-noise ratio and structural similarity as quantitative indicators to prove the effectiveness of the proposed WFMAN, GSCAN, and FFGF through the GOPRO dataset. The unit of the number of parameters is million.
[0096] The method proposed by the present invention includes WFMAN, GSCAN, and FFGF. It should be noted that when only WFMAN is used without using GSCAN, or only GSCAN is used without using WFMAN, the number of sub-networks used is doubled to achieve the same effect as when WFMAN and GSCAN are connected in series. When FFGF is not used, the method of direct addition is used for feature fusion.
[0097] Table 1 shows the ablation experiments of each part proposed by the present invention.
[0098] Table 1
[0099] WFMAN GSCAN FFGF PSNR (dB) MACs (G) Params. (M) √ 33.20 59.5 18.12 √ 33.08 59.25 8.94 √ √ 33.22 59.37 13.53 √ √ 33.38 62.18 18.23 √ √ 33.28 61.93 9.05 √ √ √ 33.39 62.06 13.64
[0100] The following three points can be seen from Table 1:
[0101] 1) When only WFMAN and GSCAN proposed by the present invention are used, the obtained PSNRs are 33.20 and 33.08 respectively, and the computational complexity remains around 59. This performance is already competitive among existing SOTA methods. When WFMAN and GSCAN are connected in series and FFGF is added to the network, the PSNR of this method reaches 33.39, while the computational complexity is only 62.06 and the number of parameters is only 13.64. Whether in terms of the performance of the model or its complexity, it has a huge advantage compared with existing methods.
[0102] 2) When only GSCAN is used, the PSNR of the network is not advantageous, but it can maintain a very low number of parameters; when only WFMAN is used, the PSNR of the network performs well, but it brings a large number of parameters, reaching 18.12. However, if the sub-network connection strategy proposed by the present invention is adopted, and WFMAN and GSCAN are connected in series, not only does it not increase the additional computational cost, but also a small increase in PSNR is achieved. Most importantly, this connection method significantly reduces the number of parameters of the network to 13.53, which is about 75% of the original, thus achieving an excellent balance between performance and model scale.
[0103] 3) When the proposed FFGF module is added to the network, the PSNR of the network is significantly increased by 0.17 to 0.2, achieving a large increase. It is worth mentioning that adding FFGF only increases the number of parameters of the network by 0.11 and the computational complexity by 2.69. It can be seen that FFGF is very important for the entire network. It not only greatly improves the performance, but also only requires a low cost.
[0104] To prove the superiority of the method proposed in the present invention in terms of quantitative metrics, the present invention was compared with fifteen existing methods on the GoPro and HIDE datasets. Specifically, all methods were trained on the GoPro training set and then tested on the GoPro test set and the HIDE test set respectively. Table 2 shows the comparison results between the present invention and other existing methods.
[0105] Table 2
[0106]
[0107] Table 2 shows the average PSNR and average SSIM of all methods on the GoPro test set and the HIDE test set, as well as the computational complexity and the number of parameters of these methods. As shown in Table 2, the method proposed in the present invention not only outperforms most of the existing methods on the GoPro dataset, but also has the lowest computational complexity and the number of parameters among all methods. To prove the superiority of the method proposed in the present invention in terms of subjective visual effects, a comparison of the deblurring effects was made with five existing methods. As Figures 5 - 9 shown, in each figure, from left to right are the blurred image, MIMO-UNet+, MPRNet, HINet, NAFNet, Uformer-B, the clear image inferred by the present method, and the real clear image. It can be seen from the figure that the images deblurred by the existing methods have varying degrees of distortion, blurring, and artifacts, while the method of the present invention can not only obtain richer details and sharper edges, but also achieve the highest PSNR. Table 2 and Figures 5 - 9 show that the method proposed in the present invention outperforms the existing methods in both qualitative and quantitative metrics.
[0108] In the above embodiments, the mention of "this embodiment" in the specification means that the specific features, structures, or characteristics described in connection with the embodiment are included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.
[0109] In the above embodiments, although the present invention has been described in connection with specific embodiments of the present invention, many substitutions, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) can be used in the embodiments discussed. The embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims.
[0110] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements any one of the methods in this embodiment.
[0111] This embodiment also provides an electronic terminal, including: a processor and a memory;
[0112] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes any one of the methods in this embodiment.
[0113] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to the computer program. The foregoing computer program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk or optical disc that can store program codes.
[0114] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store a computer program, the communication interface is used for communication, and the processor and the transceiver are used to run the computer program to make the electronic terminal execute each step of the above method.
[0115] In this embodiment, the memory may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0116] The above-mentioned processor may be a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it may also be a digital signal processor (Digital Signal Processing, abbreviated as DSP), an application specific integrated circuit (Application SpecificIntegrated Circuit, abbreviated as ASIC), a field programmable gate array (Field-Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0117] The present invention can be used in many general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0118] The present invention may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including storage devices.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention may be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. An image blind deblurring method based on local dual-domain mutual attention and global channel self-attention, characterized in that: Including the following steps: S1: Construct a window-based frequency-domain and spatial-domain mutual attention network WFMAN, which has two branches in the spatial domain and the frequency domain. The features from the frequency domain are activated by a Gaussian error linear unit and multiplied by the features in the spatial domain as weights to achieve feature fusion of the mutual attention mechanism; S2: Construct a global spatial-domain channel self-attention network GSCAN, which has only one path in the spatial domain. After passing through the global channel attention mechanism, the features from itself are multiplied by the weights from itself to perform feature fusion of the self-attention mechanism; S3: Construct a frequency-domain filtering gated feature fusion module FFGF. By introducing a learnable quantization matrix W, adaptively determine the shallow information from the encoder to be retained, and fully fuse the information between the encoder and the decoder through a simple gating operation; S4: Connect WFMAN and GSCAN in series to form a basic module; use multiple basic modules, FFGF modules and refinement layers to construct a new U-shaped network; S5: Use the image datasets before and after deblurring to train and test the new U-shaped network to obtain an image blind deblurring model; S6: Use the image blind deblurring model to deblur the acquired images.
2. The image blind deblurring method based on local double-domain mutual attention and global channel self-attention according to claim 1, characterized in that: The window-based frequency-domain and spatial-domain mutual attention network WFMAN described in step S1 specifically includes: The input for the i-th encoder First, perform layer normalization on it, and then divide it into upper and lower branches. Each branch includes a feature extraction stage and a feature interaction stage; The upper branch is used for operations in the spatial domain, and its feature extraction stage is described as follows: First, perform a 1×1 convolution and a 3×3 depthwise convolution on the features in sequence; then perform a SimpleGate operation, split the input features into two equal parts along the channel dimension, and directly perform an arithmetic dot product on these two parts; then perform a simplified channel attention mechanism on the features, first perform a global average pooling operation on the features, then perform a 1×1 convolution operation, and finally multiply it pointwise to the original features; then aggregate the features through a 1×1 convolution operation, and finally perform a residual connection operation to obtain the output Feature of the feature extraction part of the upper branch top : Among them represents the input of the i-th encoder; LN() represents the layer normalization operation; Conv1x1() represents the convolution operation with a convolution kernel size of 1×1; DConv3x3() represents the depthwise convolution operation with a convolution kernel size of 3×3; SG() represents the simple gating mechanism operation; SCA() represents the simple channel attention mechanism operation; The lower branch is used for operations in the frequency domain, and its feature extraction stage is described as follows: First, the features are divided into windows; then, the window features are subjected to Fourier transform, linear layer operation, GELU non-linear activation, and inverse Fourier transform to obtain effective information about the fuzzy pattern; finally, the window features after Fourier transform are converted into features of the original size to obtain the output Feature of the feature extraction part of the lower branch of WFMAN bot : where FFT() represents a two-dimensional fast Fourier transform operation; Linear() represents a linear transformation operation in the complex domain; GELU() represents a Gaussian error linear unit activation function; IFFT() represents a two-dimensional inverse fast Fourier transform operation; In the feature interaction stage of the upper branch, the upper branch feature Feature top Perform layer normalization and linear mapping operations; in the feature interaction stage of the lower half branch, the lower half branch feature Feature bot Perform layer normalization; then perform a 1×1 convolution and a 3×3 depth-wise convolution operation; then perform GELU nonlinear activation, use the features of the lower branch as weights, and multiply them by the features of the upper branch to implement mutual attention operation, so as to learn the frequency-space dual-domain representation to use kernel-level and pixel-level features for deblurring; finally perform a 1×1 convolution operation and a residual connection operation to obtain the output of WFMAN out : WFMAN out = Conv1x1{Linear(Norm(Feature top )) ⊙GELU(Dconv3x3(Conv1x1(Norm(Feature bot ))))+Feature top Among them, Feature top represents the output of the upper half-branch feature extraction part of WFMAN; Feature bot represents the output of the lower half-branch feature extraction part of WFMAN; Linear represents a linear transformation operation in the real number field.
3. The image blind deblurring method based on local dual-domain mutual attention and global channel self-attention according to claim 1, characterized in that: The global spatial-domain channel self-attention network GSCAN described in step S2 specifically includes: Feature extraction: Take the output of WFMAN, WFMAN out as the input of GSCAN, GSCAN in , first perform layer normalization; then sequentially perform a 1×1 convolution and a 3×3 depthwise convolution; then perform the SimpleGate operation, split the input features along the channel dimension into two equal parts, and directly perform an arithmetic dot product on these two parts; then perform a simplified channel attention mechanism on the features, first perform global average pooling on the features, then perform a 1×1 convolution operation, and finally dot multiply it to the original features; then aggregate the features through a 1×1 convolution operation, and finally perform a residual connection operation to obtain the output Feature of the GSCAN feature extraction part GSCAN : Feature GSCAN = Conv1x1(SCA(SG(DConv3x3(Conv1x1(Norm(GSCAN in )))))) + GSCAN in Feature fusion: Input Feature GSCAN into the two input ends of the feature fusion part respectively. For the input Feature GSCAN , perform layer normalization and linear mapping operations on the upper branch; on the lower branch, perform layer normalization, 1×1 convolution, 3×3 depthwise convolution, and GELU activation operations; then multiply the weights from the lower branch into the upper branch to complete the self-to-self fusion of features and achieve the effect of self-attention; finally, perform 1×1 convolution and residual connection to obtain the output GSCAN out of GSCAN: GSCAN out = Conv1x1{Linear(Norm(Feature GSCAN )) ⊙GELU(Dconv3x3(Conv1x1(Norm(Feature GSCAN ))))}+Feature GSCAN 。 4. The image blind deblurring method based on local double-domain mutual attention and global channel self-attention according to claim 1, wherein: The frequency-domain filtering gated feature fusion module FFGF described in step S3 specifically includes: For the output of the i-th layer encoder First, perform a 1×1 convolution on it, then expand the features by PATCH, then perform a Fourier transform to transform the features into the frequency domain, then multiply by a quantization matrix W to adaptively select useful encoder information, and finally perform an inverse Fourier transform and PATCH folding operation to change the features back to the size at input; Then, the corresponding decoder layer features with deep information are input, and concatenated with the filtered encoder layer features on the channel dimension. Then, 1×1 convolution is used to integrate the concatenated information, and finally, SimpleGate is used to achieve effective feature fusion, obtaining the output FFGF of FFGF out : Among them represents the output of the i-th layer encoder, represents the output of the (i + 1)-th layer decoder; W represents the quantization matrix; Cat represents the concatenation operation along the channel dimension.
5. The image blind deblurring method based on local double-domain mutual attention and global channel self-attention according to claim 1, wherein: The connection of WFMAN and GSCAN in series to form a basic module described in step S4 is defined as follows: Among them represents the input of the i-th layer encoder or decoder, represents the output of the i-th layer encoder or decoder; WFMAN represents a window-based frequency-domain to spatial-domain cross-attention network; GSCAN represents a global-based spatial-domain channel self-attention network.
6. The blind image deblurring method with local dual-domain mutual attention and global channel self-attention according to claim 1, characterized in that: The new U-shaped network described in step S4 includes a 3×3 convolution that maps the channels to 64, three encoder layers, two downsampling operations, two upsampling operations, two FFGFs, three decoder layers, a refinement layer, and a 3×3 convolution that remaps the channels back to 3; where: The first encoder layer EB1, the second encoder layer EB2, the first decoder layer DB1, and the second decoder layer DB2 are each composed of a basic module; The third encoder layer EB3 and the third decoder layer DB3 are each composed of 4 identical basic modules stacked; The refinement layer is composed of 2 identical basic modules stacked; Input to the second decoder layer Input to the first decoder layer The downsampling operation includes a bilinear interpolation with a scale of 0.5 and a 3×3 convolution that doubles the channels; the upsampling operation includes a bilinear interpolation with a scale of 2 and a 3×3 convolution that halves the channels.
Citation Information
Patent Citations
Remote sensing image super-resolution reconstruction method based on multi-stage space-frequency combination
CN118052712A
Single-image defogging method based on detail restoration
WO2024178979A1