Non-uniform motion image deblurring method based on multi-scale gating module

By designing a multi-scale gating network MSGN in image defuzzing technology, using the multi-scale gating module MGB and structured feedforward network SFFN, combined with the structured fusion strategy module SFS, the problem of difficulty in capturing the global features of the image in the existing technology is solved, and a higher quality image defuzzing effect is achieved.

CN120219231APending Publication Date: 2025-06-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510297047.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing image defuzzing technology is difficult to fully capture the global features of the image, and the local characteristics based on the convolution operation limit the defuzzing effect, and the high computational complexity of the self-attention mechanism also limits its efficiency and performance balance in practical applications.

Method used

A non-uniform motion image defuzzing method based on multi-scale gating module is designed. By constructing a multi-scale gating network MSGN, the multi-scale gating module MGB and structured feedforward network SFFN are used, and the structured fusion strategy module SFS is combined to realize the multi-scale fusion of images and the effective utilization of structured information.

Benefits of technology

This method can not only retain local details, but also combine global structural information to provide a more comprehensive recovery effect, significantly improve the quality of defuzzing results, and better reconstruct the details and global structure of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219231A_ABST
    Figure CN120219231A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image restoration, and particularly relates to a non-uniform motion image deblurring method based on a multi-scale gating module, which comprises the following steps: constructing a multi-scale gating network MSGN, and carrying out image deblurring through the constructed multi-scale gating network MSGN; according to the method, a multi-scale gating module MGB is designed, structural information of each block is integrated into a feature block consistent with the block in position and size through block averaging operation, the structural information is fused into pixel-level features through a point-by-point multiplication mechanism, and the edge and texture recovery capacity is remarkably improved; in addition, the feature of each pixel is directly fused with the structural information of the corresponding block through the structured feed-forward network SFFN, so that the modeling capability of the network on the multi-scale structural information is effectively enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image restoration, and particularly relates to a method for deblurring non-uniform motion images based on a multi-scale gating module. Background Art

[0002] During the image acquisition process, due to factors such as camera shake and object movement, motion blur is likely to occur, resulting in a decline in image quality. Existing image deblurring techniques mostly rely on statistical prior knowledge and manually designed regularization terms, and estimate the blur kernel and iteratively reconstruct the clear image. However, these methods have limited effects when dealing with real-scene images. With the development of deep learning, deblurring methods based on convolutional neural networks (CNNs) have made significant progress. These methods usually use the U-Net architecture to capture local features and estimate residual information, but due to the local characteristics of convolutional operations, it is difficult to comprehensively capture the global features of images. To solve this problem, researchers have developed deblurring methods based on Transformer, which capture global relationships through the self-attention mechanism. However, the high computational complexity of the self-attention mechanism limits the efficiency and performance balance in practical applications. In addition, in recent years, deblurring methods based on multi-layer perceptrons (MLPs) and gating mechanisms have gradually emerged, focusing on improving the performance of network modules while reducing model complexity, but still face the challenge of how to efficiently model global and local features. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a method for deblurring non-uniform motion images based on a multi-scale gating module, including:

[0004] Construct a multi-scale gating network MSGN, and perform image deblurring through the constructed multi-scale gating network MSGN;

[0005] The multi-scale gating network MSGN includes: two 3×3 convolutional layers, three encoders EB1, EB2, EB3 composed of several basic blocks, three decoders DB1, DB2, DB3 composed of several basic blocks, a structured fusion strategy module SFS, and an additional basic block;

[0006] The basic block sequentially integrates a multi-scale gating module MGB and a structured feed-forward network SFFN;

[0007] The structures of the decoders DB1, DB2, DB3 are mirror-symmetric to those of the encoders EB1, EB2, EB3;

[0008] Performing image deblurring through the constructed multi-scale gating network MSGN includes:

[0009] For the input image The multi-scale gated network MSGN extracts shallow features through a 3×3 convolutional layer where H, W, and C represent height, width, and channel dimensions respectively;

[0010] The extracted shallow features are passed to three layers of encoders EB1, EB2, and EB3, and the encoders gradually reduce the spatial resolution and increase the channel depth;

[0011] The output features of the last layer decoder DB3 gradually upsample the low-resolution features through three layers of decoders DB3, DB2, and DB1, while reducing the number of channels in each layer; among them, during the gradual upsampling process of the decoder, the output features of decoder DB3 and decoder DB2 are respectively fused with the output features of their mirror-symmetric decoders through the structured fusion strategy module SFS and then input into the next layer of decoder; bilinear interpolation is used for both downsampling and upsampling;

[0012] Based on the output features of decoder DB1, an additional basic block further refines the high-resolution features;

[0013] Based on the output features of the additional basic block, a 3×3 convolutional layer is used to adjust the number of channels and generate a residual image By adding the residual image R back to the input image, the restored image I R = I + R.

[0014] Advantages of the present invention:

[0015] 1. The multi-scale gated module MGB designed in the present invention extracts multi-scale information by using block averaging operation and effectively fuses pixel-level and structure-level information through a gated mechanism. This method enables the image restoration process to not only retain local details but also combine global structure information, providing a more comprehensive restoration effect; this multi-scale fusion strategy can effectively process blurred features of different scales and improve the quality of the deblurring result.

[0016] 2. The structured feedforward network SFFN designed in the present invention can consider both local information and global structure information during the image restoration process, making the deblurring effect more accurate, thereby better reconstructing the details and global structure of the image. Compared with traditional feedforward networks that only rely on pixel information, SFFN introduces the structure information of the image through block averaging operation, enabling the model to better understand the structural relationship of the image when processing complex blurred images, thereby effectively improving the restoration degree of details and global consistency during the deblurring process.

[0017] 3. When the Structured Fusion Strategy module SFS designed by the present invention combines the features of the encoder and the decoder, it simultaneously considers pixel-level and structure-level information, optimizing the feature fusion process. This strategy can enhance the degree of information fusion in different feature extraction stages, thereby improving the defocusing quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic diagram of the overall structure of the multi-scale gated network MSGN of a method for defocusing non-uniform motion images based on a multi-scale gated module according to the present invention;

[0019] Figure 2 It is a schematic diagram of the defocusing result of the car edge in the GoPro dataset by a method for defocusing non-uniform motion images based on a multi-scale gated module according to the present invention;

[0020] Figure 3 It is a schematic diagram of the defocusing result of the text on the shopping bag in the.HIDE dataset by a method for defocusing non-uniform motion images based on a multi-scale gated module according to the present invention;

[0021] Figure 4 It is a schematic diagram of the performance of a method for defocusing non-uniform motion images based on a multi-scale gated module according to the present invention on the RealBlur-J and RealBlur-R datasets. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] A method for defocusing non-uniform motion images based on a multi-scale gated module, as Figure 1 shown, where (a) the overall network structure, including a basic block integrating MGB and SFFN, and a backbone network of SFS for encoder-decoder feature fusion; (b) the MGB module proposed by the present invention for utilizing multi-scale and structural information; (c) the SFFN module proposed by the present invention; (d) the block average (PA) operation that sequentially performs average pooling and broadcast operations; (e) the overall process of the SFS module proposed by the present invention. includes:

[0024] Construct a multi-scale gated network MSGN, and perform image defocusing through the constructed multi-scale gated network MSGN;

[0025] The multi-scale gated network MSGN includes: two 3×3 convolutional layers, three encoders EB1, EB2, and EB3 composed of several basic blocks, three decoders DB1, DB2, and DB3 composed of several basic blocks, a structured fusion strategy module SFS, and an additional basic block;

[0026] The basic block sequentially integrates a multi-scale gating module MGB and a structured feed-forward network SFFN. A block averaging operation is used in both structures, that is, the feature map is divided into blocks along the pixel dimension, and all pixel values of each block are set to the mean value of the pixels within the block. Among them, MGB includes: a block averaging operation, a layer normalization, two 1×1 convolutional layers, and a 3×3 depthwise separable convolutional layer. The specific processing process is as follows: First, the input feature map passes through layer normalization, a 1×1 convolution, and a 3×3 depthwise separable convolution to obtain a roughly extracted feature map, whose feature dimension is twice that of the input feature map; Then, the feature map is evenly split into two parts, F x and F y along the channel dimension, and then these two parts are element-wise multiplied to obtain the feature map F p . After that, F x is subjected to a block averaging operation, and then element-wise multiplied with F y to obtain the feature map F s . Finally, the feature maps F p and F s are concatenated along the channel dimension, passed through a 1×1 convolution to adjust the number of channels to be the same as the input feature Figure 1 map, and then added to the input feature map to obtain the output feature map of the module. SFFN includes: a block averaging operation, a layer normalization, two 1×1 convolutions, a 3×3 depthwise separable convolution, and a GEGLU activation unit. The specific processing process is as follows: First, the input feature map passes through a layer normalization and then a block averaging operation to obtain the feature map F pa . Then, the feature map F pa is concatenated with the feature map that has not undergone the block averaging operation along the channel dimension and passed through a 1×1 convolution. After that, the output feature map passes through a 3×3 depthwise separable convolution, a GEGLU activation unit, and a 1×1 convolution. Finally, the output feature map is added to the input feature map to obtain the output feature map of the module.

[0027] The structures of the decoders DB1, DB2, and DB3 are mirror-symmetric to those of the encoders EB1, EB2, and EB3;

[0028] The MSGN proposed by the present invention adopts a hierarchical U-Net structure, which is an architecture widely used in the image deblurring task. The overall structure is as shown in Figure 1 (a). For the input image The network first extracts shallow features through a 3×3 convolutional layer where H, W, and C represent height, width, and channel dimensions respectively. Then, these features are passed to a three-layer encoder that gradually reduces the spatial resolution and increases the channel depth to extract hierarchical context-aware representations. Each encoder layer (EB1, EB2, EB3) consists of several base blocks, and each base block sequentially integrates the proposed MGB and SFFN modules. The decoder (DB1, DB2, DB3) has a structure mirror-symmetric to the encoder, gradually upsampling the low-resolution features while reducing the number of channels at each layer. Bilinear interpolation is used for both downsampling and upsampling. To enhance the feature fusion between the encoder and decoder, the present invention designs an SFS module that combines MGB and SFFN to achieve better feature integration by leveraging the structural attributes of the image. In the final stage of the network, additional base blocks are used to further refine the high-resolution features. Finally, a 3×3 convolutional layer is used to adjust the number of channels and generate the residual image Restored image I R I is obtained by adding the residual R back to the input image R = I + R The main innovation of the present invention lies in the design of the MGB and SFFN modules in the base block, and the SFS module for encoder-decoder feature fusion

[0029] Image deblurring is performed through the constructed multi-scale gated network MSGN, including:

[0030] For the input image The multi-scale gated network MSGN extracts shallow features through a 3×3 convolutional layer where H, W, and C represent height, width, and channel dimensions respectively;

[0031] The extracted shallow features are passed to the three-layer encoder EB1, EB2, EB3. The encoder gradually reduces the spatial resolution and increases the channel depth, and uses bilinear interpolation for downsampling to reduce the resolution;

[0032] The output features of the last layer decoder DB3 are gradually upsampled by the three-layer decoder DB3, DB2, DB1 for the low-resolution features while reducing the number of channels at each layer; among them, during the gradual upsampling process of the decoder, the output features of decoder DB3 and decoder DB2 are respectively fused with the output features of their mirror-symmetric decoders through the structured fusion strategy module SFS and then input to the next layer decoder; after each layer of decoder is processed, bilinear interpolation is used for upsampling to increase the resolution and then input to the SFS module for fusion with the encoder features of the same scale;

[0033] Based on the output features of the decoder DB1, additional basic blocks further refine the high-resolution features;

[0034] Based on the output features of the additional basic blocks, a 3×3 convolutional layer is used to adjust the number of channels and generate a residual image By adding the residual image R back to the input image, the restored image I is obtained R = I + R.

[0035] The design goals of the MGB include:

[0036] 1. Utilize multi-scale information while preserving texture details and make the feature gradients consistent with the clear image;

[0037] 2. Use structural information to guide pixel recovery through an efficient fusion mechanism;

[0038] 3. Integrate structural-level and pixel-level information in a single module.

[0039] The multi-scale gated module MGB, as Figure 1 (b) shows, includes: a block averaging operation, a layer normalization, two 1×1 convolutional layers, and a 3×3 depthwise separable convolutional layer. The specific processing process is as follows: First, the input feature map passes through layer normalization, a 1×1 convolution, and a 3×3 depthwise separable convolution in sequence to obtain a roughly extracted feature map, whose feature dimension is twice that of the input feature map; Then, this feature map is evenly split into two parts, F x and F y along the channel dimension, and then these two parts are element-wise multiplied to obtain the feature map F p . After that, F y is subjected to a block averaging operation, and then element-wise multiplied with F x to obtain the feature map F s . Finally, the feature maps F p and F s are concatenated along the channel dimension, passed through a 1×1 convolution to adjust the number of channels to be the same as the input feature Figure 1 and then added to the input feature map to obtain the output feature map of the module.

[0040] To achieve Goal 1, a block averaging (PA) operation is introduced to utilize multi-scale information and preserve texture details. As Figure 1 (d) shows, given the input feature map FFF, the PA operation combines non-overlapping average pooling and broadcast operations to restore the resolution of the feature map. The specific formula is:

[0041]

[0042] where represents non-overlapping average pooling with a block size of k×k, Broadcastk×k It means that each pooling value is copied k×k times to restore the original resolution. By averaging the features in the local block, this operation preserves the texture pattern while capturing multi-scale information and reducing the gradient domain gap between the features and the clear image.

[0043] To achieve Goal 2, Pixel Gating (PG) is designed to use the structural information obtained by the PA operation to guide pixel restoration. PG is represented by the following formula:

[0044] PG(F) = F ⊙ PA(F)

[0045] where F is the input feature map and ⊙ represents element-wise multiplication. This design enables the network to selectively amplify or suppress specific feature information according to the structural importance.

[0046] To achieve Goal 3, pixel-level gating and structure-level gating operations are combined. The pixel-level gating uses a Simple Gating (SG) operation, which divides the input feature map into two parts along the channel dimension and modulates them through element-wise multiplication. By applying SG and PG to the same feature map simultaneously and weighting their outputs, MGB can capture the complex interactions between local and structural features.

[0047] In summary, the formula for MGB is:

[0048] F x ,F y = Split(Conv 3×3 (Conv 1×1 (Norm(F in ))))

[0049] F p = F x ⊙ F y ,F s = F x ⊙ PA(F y )

[0050] F mgb = F in + Conv 1×1 (Concat(F p ,F s ))

[0051] where F in represents the input feature map, Norm(·) represents layer normalization, Conv 1×1 represents a convolution operation with a 1×1 convolution kernel, DWConv 3×3 represents a depthwise separable convolution with a 3×3 convolution kernel, Split(·,·) represents splitting the feature map into two equal parts along the channel dimension, F x and F yrespectively represent the feature maps after being segmented in different directions, ⊙ represents the element-wise dot product operation, PA(·) represents the block average operation, and F p and F s respectively represent the pixel fusion feature and the structure fusion feature, Concat(·,·) represents concatenating two feature maps in the channel dimension, and F mgb represents the output feature map of the MGB module.

[0052] By cascading two types of gating, the MGB module proposed by the present invention can achieve sufficient interaction between the pixel domain and the structure domain features, thereby enhancing the overall feature representation ability.

[0053] The structured feed-forward network SFFN includes: a block average operation, a layer normalization, two 1×1 convolutions, a 3×3 depthwise separable convolution, and a GEGLU activation unit. Its specific processing process is as follows: First, the input feature map goes through a layer normalization and then a block average operation to obtain the feature map F pa . Then, the feature map F pa is concatenated with the feature map that has not undergone the block average operation along the channel dimension and passes through a 1×1 convolution. After that, the output feature map passes through a 3×3 depthwise separable convolution, a GEGLU activation unit, and a 1×1 convolution. Finally, the output feature map is added to the input feature map to obtain the output feature map of the module.

[0054] Traditional feed-forward networks (FFNs) rely on pixel-level information and cannot make full use of cross-scale or structure information. Inspired by the block average operation, structural information and multi-scale features are introduced into the FFN to more effectively guide pixel recovery. As Figure 1 (c) shows, the structural features obtained through PA are concatenated with the original feature map in the channel dimension, and a convolution is applied to adjust the number of channels, and then a dot product attention operation is performed. The specific formula is:

[0055] F pa = PA(Norm(F mgb ))

[0056] F p ' a = DWConv 3×3 (Conv 1×1 (Concat(F mgb ,F s )))

[0057]

[0058] where, F mgb represents the output feature map of the MGB module, Norm(·) represents the layer normalization, PA(·) represents the block average operation, and Fpa denotes a feature map containing structural information, Concat(·,·) denotes concatenating two feature maps along the channel dimension, and Conv 1×1 denotes a convolution with a kernel size of 1×1, and DWConv 3×3 denotes a depthwise separable convolution with a kernel size of 3×3, and F pa ′ denotes the feature map obtained after fusion, denotes the GEGLU activation function, and F out denotes the output feature map of the SFFN module.

[0059] Most networks usually fuse encoder and decoder features through concatenation or skip connections. However, decoder features usually contain more refined recovery results, which are different from encoder features. To better combine encoder and decoder features while considering pixel-level and structural-level information, this application proposes a novel structured fusion strategy (SFS), as shown in Figure 1 (e).

[0060] The data fusion process of the structured fusion strategy module SFS includes:

[0061] Given the encoder feature and the decoder feature Concatenate the two along the channel dimension and fuse the features through a 1×1 convolution;

[0062] Pass the fused feature map into the MGB and SFFN modules to filter and refine the features at the pixel level and the structural level;

[0063] The refined features are based on increasing the channel dimension through convolution and divided into two parts along the channel dimension and and and Add them together to obtain the fused feature F fuse .

[0064] To verify the effectiveness of the proposed MSGN, it was evaluated on widely used image deblurring datasets, including the GoPro dataset, the HIDE dataset, and the RealBlur dataset. The experiments were divided into two settings to evaluate the performance of MSGN on synthetically blurred datasets and real-world blurred datasets:

[0065] Setting I: Synthetic dataset;

[0066] The model was trained on the GoPro dataset, which contains 2103 pairs of synthetically blurred and clear image pairs, and tested on the GoPro test set (1111 pairs of images) and the HIDE dataset.

[0067] Setting II: Real-world Datasets

[0068] The model is trained on the RealBlur-R and RealBlur-J datasets, which contain 3,758 pairs of image pairs for training respectively, and is evaluated on their respective test sets (each containing 980 pairs of images).

[0069] The same loss function as MIMO-UNet is adopted, and the model is trained using the Adam optimizer with hyperparameters set as β1 = 0.9 and β2 = 0.9. The entire training process consists of 300,000 iterations, with an initial learning rate of 1×10 -3 , which is gradually reduced to 1×10 -7 through the cosine annealing strategy. The training image patch size is 256×256, and the batch size is 16. The data augmentation strategy follows the scheme of Restormer.

[0070] In all modules of MSGN, the patch size of the patch averaging operation is set to 8×8. Two model variants are trained: MSGN-S (small) and MSGN-B (base). The final models include:

[0071] 1. MSGN-S and MSGN-B trained based on I;

[0072] 2. MSGN-B trained based on II.

[0073] The model is implemented based on the PyTorch platform and trained using four Nvidia RTX A100 GPUs.

[0074] In this embodiment, the method of the present invention is compared with the latest methods:

[0075] 3.1 Quantitative comparison:

[0076] Table 1. Comparison results on the GoPro and HIDE datasets under Setting I, where the best values are marked in bold underline and the second-best values are marked in underline:

[0077] Table 1

[0078]

[0079]

[0080] Table 2. Comparison results on the RealBlur-R and RealBlur-J datasets under Setting II, where the best values are marked in bold underline and the second-best values are marked in underline:

[0081] Table 2

[0082]

[0083] For Setup I, the MSGN-S and MSGN-B models were trained on 2103 pairs of images from the synthetic GoPro dataset and evaluated on the GoPro test set (1111 pairs of images) and the HIDE dataset. As shown in Table 1, the results were compared with recent state-of-the-art methods, using the average peak signal-to-noise ratio (PSNR), average structural similarity index (SSIM), and the number of parameters as evaluation metrics. The MSGN models outperformed all CNN-, Transformer-, and MLP-based methods in terms of both PSNR and SSIM on the GoPro and HIDE datasets. Specifically, the MSGN-S model achieved competitive performance with the lowest number of parameters, while the MSGN-B model outperformed all state-of-the-art methods, achieving the best PSNR and SSIM scores. Notably, the average PSNR of the MSGN-B model was 0.13 dB and 0.01 dB higher than that of UFPNet on the GoPro and HIDE datasets, respectively, but its number of parameters was only 32% of that of UFPNet.

[0084] For Setup II, the MSGN-B model was trained on the real-world blur datasets RealBlur-R and RealBlur-J, each containing 3758 pairs of image pairs, and tested on their respective test sets (each containing 980 pairs of image pairs). As shown in Table 2, the method of the present invention outperformed compared with other state-of-the-art methods. Compared with LoFormer-B with a similar number of parameters (the second-best performing method on the RealBlur-R and RealBlur-J datasets), the MSGN-B had an average PSNR improvement of 0.08 dB on RealBlur-R and 0.14 dB on RealBlur-J. These results indicate that the method of the present invention has excellent effects in dealing with real-world blur problems and outperforms existing methods in both synthetic data and real-world data evaluations.

[0085] Qualitative comparison:

[0086] Figures 2 to 4 Visual comparison results of the proposed MSGN of the present invention with recent state-of-the-art image deblurring methods on the GoPro, HIDE, RealBlur-J, and RealBlur-R datasets are shown. The method of the present invention can continuously recover clearer and sharper details in both synthetic and real-world datasets.

[0087] Figure 2 From the GoPro dataset. At Figure 2In it, the edges of the car appear sharper and clearer, demonstrating the ability of the present invention to restore details of fast-moving objects.

[0088] Figure 3 from the HIDE dataset, which focuses on images containing human motion blur. In Figure 3 for the text on the shopping bag, the method of the present invention restores it clearly and readable, while other methods fail to generate clear results.

[0089] For the real-scene dataset, Figure 4 shows the performance of the present invention on the RealBlur-J and RealBlur-R datasets. Figure 4 The second row of Figure 4 is from the RealBlur-J dataset, focusing on fluorescent text. The method of the present invention successfully restores clear and distinguishable text details, while other methods fail to generate usable results. The third row of

[0090] is also from the RealBlur-J dataset, targeting advertisement text. The method of the present invention restores it sharply and accurately, being clearer and more readable than other methods. The last two rows are from the RealBlur-R dataset, corresponding to the third and fourth images in the top row. Even under low-light conditions, the method of the present invention can effectively restore the clarity and accuracy of fluorescent text, demonstrating its robustness in complex lighting scenarios. Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A non-uniform motion image deblurring method based on a multi-scale gating module, characterized in that: include: Construct a multi-scale gating network MSGN, and perform image deblurring through the constructed multi-scale gating network MSGN; The multi-scale gating network MSGN includes: two 3×3 convolutional layers, encoders EB1, EB2, EB3 composed of three layers of basic blocks, decoders DB1, DB2, DB3 composed of three layers of basic blocks, a structured fusion strategy module SFS and an additional basic block; The basic block sequentially integrates a multi-scale gating module MGB and a structured feed-forward network SFFN; The structures of the decoders DB1, DB2, DB3 are mirror-symmetrical to those of the encoders EB1, EB2, EB3; Image deblurring is performed by constructing a multi-scale gating network MSGN, including: For the input image Multi-scale gating network MSGN extracts shallow features through a 3×3 convolutional layer Among them, H, W, and C represent the height, width, and channel dimensions respectively; The shallow features extracted Passed to the three-layer encoder EB1, EB2, and EB3, the encoder gradually reduces the spatial resolution and increases the channel depth, and uses bilinear interpolation downsampling to reduce the resolution; The output features of the last layer of decoder DB3 are gradually upsampled to low-resolution features through three layers of decoders DB3, DB2, and DB1, while reducing the number of channels at each layer. In the process of progressive upsampling of the decoders, the output features of decoders DB3 and DB2 are fused with the output features of the mirror-symmetric decoders through the structured fusion strategy module SFS and then input into the next layer of decoders. After each layer of decoders is processed, bilinear interpolation upsampling is used to increase the resolution, and then input into the SFS module to be fused with the encoder features of the same scale. Based on the output features of decoder DB1, additional basic blocks further refine the high-resolution features; Based on the output features of the additional basic blocks, a 3×3 convolutional layer is used to adjust the number of channels and generate a residual image. By adding the residual image R back to the input image, the restored image I is obtained R =I+R.

2. The method for deblurring a non-uniform motion image based on a multi-scale gating module according to claim 1, characterized in that: The multi-scale gating module MGB comprises: F x ,F y =Split(DWConv 3×3 (Conv 1×1 (Norm(F in )))) F p =F x ⊙F y ,F s =F x ⊙PA(F y ) F mgb =F in +Conv 1×1 (Concat(F p ,F s )) Among them, F in represents the input feature map, Norm(·) represents layer normalization, Conv 1×1 Indicates a convolution operation with a convolution kernel of 1×1, DWConv 3×3 represents a depth-wise separable convolution with a convolution kernel of 3×3. Split(·,·) means that the feature map is evenly divided into two parts in the channel dimension. x and F y They represent the feature maps after being split in different directions, ⊙ represents the element-by-element point multiplication operation, PA(·) represents the block average operation, and F p and F s They represent pixel fusion features and structural fusion features respectively. Concat(·,·) means concatenating two feature maps in the channel dimension. mgb Represents the output feature map of the MGB module.

3. The method for deblurring a non-uniform motion image based on a multi-scale gating module according to claim 1, characterized in that: The structured feedforward network SFFN comprises: F pa =PA(Norm(F mgb )) F pa ′=DWConv 3×3 (Conv 1×1 (Concat(F mgb ,F pa ))) Among them, F mgb represents the output feature map of the MGB module, Norm(·) represents layer normalization, PA(·) represents block average operation, and F pa represents a feature map containing structural information, Concat(·,·) represents concatenating two feature maps in the channel dimension, Conv 1×1 Represents a convolution with a kernel size of 1×1, DWConv 3×3 represents a depthwise separable convolution with a kernel size of 3×3, F pa ′ represents the feature map obtained after fusion, G represents the GEGLU activation function, F out Represents the output feature map of the SFFN module.

4. The method for deblurring a non-uniform motion image based on a multi-scale gating module according to claim 1, characterized in that: The data fusion process of the structured fusion strategy module SFS includes: Given the encoder characteristics and decoder characteristics Concatenate the two along the channel dimension and fuse the features through 1×1 convolution; The fused feature map is passed to the MGB and SFFN modules to filter and refine the features at the pixel level and structure level; The refined features are based on increasing the channel dimension through convolution and are divided into two parts along the channel dimension and and will and Add together to get the fusion feature F fuse ; where H, W, and C represent the height, width, and channel dimensions, respectively.