Image deblurring method and system based on multi-scale space-time guide model
Through the image defuzzing method of multi-scale space-time guided model, the limitations of the prior art in dealing with complex non-uniform motion blur are solved, and better image clarity and quality recovery are achieved.
Patent Information
- Application Number
- CN202510299385.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
Existing image defuzzing models have limitations in dealing with complex non-uniform motion blur, especially in extracting multi-scale features and capturing dynamic changes, resulting in unsatisfactory defuzzing effects.
The image defuzzing method based on multi-scale space-time guided model is adopted, and the blurred image is decomposed into feature maps of different scales through the multi-scale feature decomposition module, the time convolution network is used to capture the time features, and the reconstruction clarity is enhanced through the edge attention-guided reconstruction module.
The blur problem in different scale ranges is effectively solved, the clarity and quality of the image is significantly improved, and the limitations of existing models when dealing with complex non-uniform motion blur.
Smart Images

Figure CN120219232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and particularly relates to an image deblurring method and system based on a multi-scale spatio-temporal guidance model. Background Art
[0002] Image deblurring technology is an important task in the field of image processing, aiming to recover clear images from blurred images. With the development of digital image processing technology, image deblurring has wide applications in multiple fields such as monitoring, autonomous driving, and medical imaging.
[0003] In early research, image deblurring mainly relied on optimization techniques (such as prior-based blur kernel estimation). Although these methods have theoretical support, they require complex parameter adjustment, resulting in high computational complexity and limited applicability.
[0004] With the development of deep learning, convolutional neural networks (CNNs) have shown significant advantages in the field of image deblurring. For example, DeepDeblur introduced a multi-scale CNN architecture to handle deblurring of dynamic scenes, improving the deblurring effect by processing images at different scales; SRN-Deblur proposed a scale-recursive network to enhance deblurring performance by sharing weights between different scales, thereby improving the clarity of images; DeblurGAN utilized a conditional generative adversarial network (GAN) to generate realistic deblurred images and improved the realism of deblurred images through adversarial training.
[0005] However, existing models have limitations in dealing with complex non-uniform motion blur, especially in extracting multi-scale features, resulting in unsatisfactory deblurring effects. In addition, existing models may lack sufficient adaptability when generalized to real-world scenarios.
[0006] In summary, in the field of image deblurring, the main disadvantages of existing technologies are as follows:
[0007] (1) First, existing models handle different-scale features poorly. When dealing with complex non-uniform motion blur, it is difficult to simultaneously take into account the full recovery of global and detailed information. When enhancing certain scale features, it is easy to overly weaken the key information of other scales, resulting in the loss of details in the deblurred image.
[0008] (2) Second, most existing models fail to effectively simulate the dynamic changes during the blurring process and cannot deeply explore the potential dynamic information in blurred images, resulting in affected deblurring effects. When processing blurred images caused by fast-moving objects, ghosting or trailing phenomena are likely to occur.
[0009] (3) Thirdly, the existing models do not make full use of edge information, and it is difficult for the network to accurately focus on key edge features, resulting in poor edge sharpness of the reconstructed image, easy appearance of jagged edges, detail loss or generation of artifacts, seriously affecting the image quality;
[0010] (4) Fourthly, the existing models lack an efficient mechanism for integrating multi-scale features, cannot give full play to the synergistic effect of different scale features, and cannot reasonably optimize and configure the features, thus restricting the improvement of the deblurring effect. Summary of the Invention
[0011] The technical problems to be solved by the present invention are as follows: 1. How to construct a module capable of effectively extracting multi-scale features of an image to comprehensively solve the blurring problems within different scale ranges; 2. How to design a module capable of capturing temporal features: regarding a blurred image as a weighted superposition of multiple "clear images", decomposing the corresponding clear image sequence through deep learning, and simulating the motion trajectory of the image in time; 3. How to use edge information to guide the network to focus on key edge features, thereby improving the sharpness and quality of the reconstructed image; 4. How to integrate the above innovative components to provide a novel and effective single-image deblurring method, overcome the limitations of previous models, and provide excellent performance in restoring image sharpness.
[0012] In view of the above problems of the existing technology, an image deblurring method and system based on a multi-scale spatio-temporal guidance model are provided, which solve the blurring problems within various scale ranges and improve the image quality by using edge details.
[0013] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0014] An image deblurring method based on a multi-scale spatio-temporal guidance model includes the following steps:
[0015] Establish a multi-scale spatio-temporal guidance model and perform model training, and then input a blurred image into the trained multi-scale spatio-temporal guidance model to obtain deblurred images of different scales. The image deblurring steps of the multi-scale spatio-temporal guidance model include:
[0016] Obtain a blurred image, and decompose the blurred image into feature maps of different scales through a multi-scale feature decomposition module;
[0017] For each scale of feature map, use a temporal convolutional network to capture temporal features and enhance the reconstruction sharpness through an edge attention-guided reconstruction module to obtain the reconstructed feature map of each scale;
[0018] The reconstructed feature map of the first scale is directly input into the corresponding encoder block for encoding. For the reconstructed feature maps of other scales, the corresponding feature attention module is used to enhance the feature fusion of the output of the encoder block of the previous scale and the output of the edge attention-guided reconstruction module of this scale by integrating the attention mechanism, and then input into the corresponding encoder block for encoding;
[0019] The output of the encoder block of the last scale is directly input into the corresponding decoder block. For other scales, the corresponding multi-scale feature fusion module is used to merge the features of different scales and then input into the corresponding decoder block for decoding. Then, the output feature image of the decoder block corresponding to each scale is aligned with the blurred image at the corresponding scale, and then the prediction output corresponding to each scale is obtained through residual connection combination to obtain the deblurred images at different scales.
[0020] Furthermore, when the blurred image is decomposed into feature maps of different scales through the multi-scale feature decomposition module, the expression is as follows:
[0021]
[0022] in, represents the feature map of the kth scale, B k is the blurred image of the kth scale obtained from the blurred image B1 by bilinear interpolation. The scale factor is 0.5 to the k-1th power. [·;·] represents the concatenation operation along the channel dimension. Conv represents the convolution operation. Conv d represents the downsampling convolution operation, F k-1 Represents the feature map of the k-1th scale The feature map of the last convolution layer is input into the convolution operation.
[0023] Furthermore, when using a temporal convolutional network to capture temporal features, it includes:
[0024] The input feature map is passed through the convolution layer to extract features. The expression is as follows:
[0025]
[0026] in, Represents the feature map of the kth scale, and Conv represents the convolution operation;
[0027] The extracted features are decomposed into multiple frames through the convolution layer, and the expression is as follows:
[0028]
[0029] Where T is the number of frames, represents the Tth frame of the kth scale;
[0030] Perform one-dimensional temporal convolution on the decomposed frames to extract temporal features, and the expression is as follows:
[0031]
[0032] where w t represents the weight of the temporal convolution kernel corresponding to time t, m is half of the temporal convolution kernel, κ is the size of the temporal convolution kernel. When m is not an integer, it is rounded down. i, j, and c represent the element at the i-th row and j-th column of the feature map in the c-th channel;
[0033] Combine the integrated temporal features to generate the final output The expression is as follows:
[0034]
[0035] where ReLU is the ReLU activation function.
[0036] Furthermore, when enhancing the reconstruction clarity through the edge attention-guided reconstruction module, it includes:
[0037] For the blurred image B at the k-th scale k Calculate the horizontal gradient using the Sobel operator and the vertical gradient According to the horizontal gradient and the vertical gradient Calculate the edge map The expression is as follows:
[0038]
[0039] For the edge map Perform batch normalization, and the expression is as follows:
[0040]
[0041] where Conv1 represents a 1×1 convolution operation, and BN represents batch normalization;
[0042] Multiply the normalized edge map and the temporal features to adjust the weights of different regions, and the expression is as follows:
[0043]
[0044] where ⊙ represents the Hadamard product.
[0045] Further, when using the corresponding feature attention module to enhance feature fusion by integrating the attention mechanism for the output of the encoder block at the previous scale and the output of the edge attention-guided reconstruction module at the current scale, it includes:
[0046] Downsample the output of the encoder block at the k-th scale by a downsampling convolutional layer to reduce the spatial dimension by half while keeping the number of channels, and the expression is as follows:
[0047]
[0048] Multiply the downsampled output of the encoder block at the k-th scale and the output of the edge attention-guided reconstruction module at the (k + 1)-th scale element-wise by Hadamard product to capture their interaction information, obtaining a fused feature map The expression is as follows:
[0049]
[0050] Process the fused feature map by depthwise separable convolution to obtain a fused feature map The expression is as follows:
[0051]
[0052] where DConv represents depthwise separable convolution, Conv1 represents a 1×1 convolution operation for adjusting the channel dimension, and BN represents batch normalization to stabilize the feature values;
[0053] Perform channel attention calculation and spatial attention calculation on the fused feature map The expressions are as follows:
[0054]
[0055] where is the channel attention map, is the spatial attention map, AP and MP are average and max pooling operations respectively, F1 and F2 are fully connected layers for reducing and expanding the channel dimension, σ represents the sigmoid function, and Conv S is the depthwise separable convolution for generating the spatial attention weight to reduce the computational complexity;
[0056] According to the fused feature map channel attention map spatial attention map and the downsampled output of the encoder block at the k-th scale Calculate the final output of the edge attention-guided reconstruction module at the (k + 1)-th scale, and the expression is as follows:
[0057]
[0058] Among them,
[0059] Furthermore, when the multi-scale feature fusion module combines features of different scales, the expression is as follows:
[0060]
[0061] Among them, represents the output of the multi-scale feature fusion module at the η-th scale, k represents the number of different scales, [·;·] represents the concatenation operation along the channel dimension, Conv represents the convolution operation, ↑ represents upsampling, and ↓ represents downsampling.
[0062] Furthermore, when combining the output feature images of the decoder blocks corresponding to each scale with the blurred images at the corresponding scales through residual connections, it specifically includes:
[0063] Concatenate the output feature image of the decoder block corresponding to the current scale with the alignment result corresponding to the previous scale after transposed convolution, and then perform a convolution operation to obtain the alignment result corresponding to the current scale. If the current scale is the last scale, perform a convolution operation on the output feature image of the decoder block corresponding to the current scale to obtain the alignment result;
[0064] Perform a residual connection between the alignment result of the current scale and the blurred image of the current scale to obtain the predicted output of the current scale.
[0065] Furthermore, when the number of different scales is 3, the predicted output expression for each scale is as follows:
[0066]
[0067] Among them, is the predicted output at the k-th scale, B k is the blurred image at the k-th scale, Conv k3 represents a convolution operation with an input channel of C k , an output channel of 3, a kernel size of 3×3, a padding of 1, and a stride of 1. Τ represents the transposed convolution operation, and are the output feature images of the decoder blocks at the 1st scale, 2nd scale, and 3rd scale, respectively.
[0068] Furthermore, during the process of training the model, the loss function expression is as follows:
[0069] Ltotal = αL cont + βL MSFR + γL temporal ,
[0070] where L cont is the content loss, and the expression is as follows:
[0071]
[0072] where K is the number of scales, n k is the total number of elements in the k-th scale, and I k represent the deblurred image and the ground truth image in the k-th scale respectively;
[0073] L MSFR is the multi-scale reconstruction loss, and the expression is as follows:
[0074]
[0075] where denotes the fast Fourier transform;
[0076] L temporal is the temporal smoothness loss, defined as:
[0077]
[0078] where is the total number of elements in scale k at time t, represents the feature at scale k and frame t. The goal of the temporal smoothness loss is to reduce the mutation between adjacent frames, usually measured by the L1 loss.
[0079] The present invention also proposes an image deblurring system based on a multi-scale spatio-temporal guidance model, including:
[0080] A multi-scale feature decomposition module for decomposing a blurred image into feature maps of different scales;
[0081] A temporal convolutional network for capturing the temporal features in the feature maps and simulating the motion trajectory of the image in time;
[0082] An edge attention-guided reconstruction module for guiding the network to focus on key edge features by using edge information to improve the clarity and quality of the reconstructed image;
[0083] A feature attention module for encoding the output of the encoder block of the previous scale and the output of the edge attention-guided reconstruction module of this scale by integrating the attention mechanism to enhance the input corresponding to the encoder block after feature fusion;
[0084] The multi-scale feature fusion module is used to merge the output features of the encoder blocks at different scales and then input them into the corresponding decoder blocks for decoding.
[0085] The multi-scale output module is used to align the output feature images of the decoder blocks corresponding to each scale with the blurred images at the corresponding scales and then combine them through residual connections to obtain the predicted output corresponding to each scale.
[0086] Compared with the prior art, the advantages of the present invention are as follows:
[0087] (1) Through multi-scale feature decomposition, it can process blurs in different scale ranges more comprehensively and can better restore image details compared with the prior art.
[0088] (2) A temporal convolutional network is proposed to decompose the corresponding clear image sequence through deep learning and simulate the motion trajectory of the image in time to capture the dynamic changes during the blurring process.
[0089] (3) The edge attention-guided reconstruction mechanism enables the network to focus on key edge features, significantly improving the clarity and quality of the image and outperforming the prior art in edge restoration.
[0090] (4) The feature attention module and the multi-scale feature fusion module are designed to make full use of the information at each scale and optimize the fusion of features at different scales. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 It is a flowchart of the method according to an embodiment of the present invention.
[0092] Figure 2 It is a schematic structural diagram of the multi-scale spatio-temporal guidance model according to an embodiment of the present invention.
[0093] Figure 3 It is a schematic diagram of the multi-scale feature decomposition module according to an embodiment of the present invention.
[0094] Figure 4 It is a schematic diagram of the edge attention-guided reconstruction module according to an embodiment of the present invention.
[0095] Figure 5 It is a schematic diagram of the encoder block and the decoder block according to an embodiment of the present invention.
[0096] Figure 6 It is a schematic diagram of the feature attention module according to an embodiment of the present invention.
[0097] Figure 7 It is a schematic diagram of the multi-scale feature fusion module according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0098] The present invention will be further described below in conjunction with specific preferred embodiments, but the protection scope of the present invention is not limited thereby.
[0099] Embodiment 1
[0100] An image deblurring method based on a multi-scale spatio-temporal guidance model is disclosed in an embodiment of the present invention, and the purpose is:
[0101] (1) To propose a model that can effectively extract multi-scale features of an image, perform feature extraction and downsampling through multiple convolutional layers, generate feature maps of different scales, comprehensively capture image details from fine textures to overall structures, and be able to process blurs in different scale ranges more comprehensively and accurately, effectively restore image details, and significantly improve the clarity and integrity of the deblurred image;
[0102] (2) Through deep learning, regarding the blurred image as a weighted superposition of multiple "clear images", decomposing the corresponding clear image sequence, simulating the motion trajectory of the image in time, and fully capturing the dynamic changes during the blurring process, thereby effectively improving the deblurring effect and making the processed image more in line with the actual scene;
[0103] (3) Using edge information to guide the network to focus on key edge features, obtaining edge information using the Sobel operator, generating an attention map to guide the network to focus on key edge features, effectively restoring edge details, significantly improving the clarity and quality of the reconstructed image, avoiding problems such as jagged edges and artifacts at the edges, and making the image edges more natural and real;
[0104] (4) By designing a feature attention module and a multi-scale feature fusion module to fuse the multi-scale feature decomposition, temporal convolutional network, and edge attention-guided reconstruction information, more effective feature integration and optimization are achieved. These components work together to improve the accuracy and robustness of the deblurring process.
[0105] As Figure 1 shown, the method of this embodiment includes the following steps:
[0106] S101) Establish a multi-scale spatio-temporal guidance model and perform model training;
[0107] S102) Input the blurred image into the trained multi-scale spatio-temporal guidance model to obtain deblurred images of different scales.
[0108] In step S101 of this embodiment, a multi-scale spatio-temporal guidance model (MSTSGM) is proposed, which includes a multi-scale feature decomposition module (MSFD), a temporal convolutional network (TCN), an edge attention-guided reconstruction module (EAGR), a feature attention module (FAM), and a multi-scale feature fusion module (MSFF). Among them, the temporal convolutional network (TCN) and the edge attention-guided reconstruction module (EAGR) correspond to each scale one by one. For scales other than the first scale, there is a corresponding feature attention module (FAM), and for scales other than the last scale, there is a corresponding multi-scale feature fusion module (MSFF). This model decomposes the input image into multi-scale features through the multi-scale feature decomposition module, captures temporal features through the temporal convolutional network, and enhances the reconstruction clarity through the edge attention-guided reconstruction module and other modules. This integrated method solves the blurring problem in various scale ranges and improves the image quality by using edge details. As Figure 2 shown, the image deblurring process of the multi-scale spatio-temporal guidance model is as follows:
[0109] Obtain a blurred image, and decompose the blurred image into multi-scale feature maps through the multi-scale feature decomposition module (MSFD);
[0110] For the feature maps of each scale, use the temporal convolutional network (TCN) to capture temporal features and enhance the reconstruction clarity through the edge attention-guided reconstruction module (EAGR) to obtain the reconstructed feature maps of each scale;
[0111] Directly input the reconstructed feature map output by the edge attention-guided reconstruction module of the first scale into the corresponding encoder block for encoding. For the reconstructed feature maps output by the edge attention-guided reconstruction modules of other scales, use the corresponding feature attention module (FAM) to enhance the feature fusion of the output of the encoder block of the previous scale and the output of the edge attention-guided reconstruction module of this scale through the integrated attention mechanism, so as to optimize the combined features, and then input the output of the feature attention module into the corresponding encoder block for encoding;
[0112] Directly input the output of the encoder block of the last scale into the corresponding decoder block. For other scales, use the corresponding multi-scale feature fusion (MSFF) module to merge the features of different scales and then input them into the corresponding decoder block for decoding. Then, after aligning the output feature images of the decoder blocks corresponding to each scale with the blurred images at the corresponding scales, combine them through residual connections to generate the predicted output corresponding to each scale, so as to obtain the deblurred images at different scales.
[0113] The following is an explanation of the relevant content.
[0114] As Figure 3As shown, the multi-scale feature decomposition (MSFD) module extracts features and downsamples through multiple convolutional layers, decomposing the input blurred image into multiple feature representations at different scales, capturing details from fine textures to overall structures. This decomposition facilitates the effective solution of blur problems within various scale ranges during subsequent processing. Given an input blurred image B1 with C = 3 channels and a feature map of size H×W, the MSFD module uses multiple convolutional layers for feature extraction and downsampling. Conv represents the convolutional operation, where the kernel size is 3×3, the stride is 1, and the padding is 1 to maintain the spatial dimension of the feature map. The downsampling convolutional operation Conv d uses a stride of 2 and appropriate padding to halve the spatial dimension. When the blurred image is decomposed into feature maps at different scales through the multi-scale feature decomposition module, the process of the MSFD module is represented as:
[0115]
[0116] where B2 and B3 are obtained from B1 through bilinear interpolation with scale factors of 0.5 and 0.25 respectively. [·;·] represents the concatenation operation along the channel dimension. F k represents the intermediate feature map at scale k (the feature map of the last convolutional layer before obtaining ), so F k-1 represents the feature map of the (k - 1)-th scale input to the last convolutional layer in the convolutional operation of the feature map. represents the feature map of the k-th scale. Specifically, and are the final output feature maps at different scales, where has C1 = 32 channels and is a feature map of the original size (size denoted as H1×W1, H1 = H, W1 = W), has C2 = 64 channels and is a feature map of 1 / 2 resolution (size denoted as H2×W2, ), has C3 = 128 channels and is a feature map of 1 / 4 resolution (size denoted as H3×W3, ).
[0117] At each scale k, the input image or its downsampled version is processed through a series of convolutional layers. For k = 1, B1 passes through the convolutional layer Conv and is concatenated with B1 to form an intermediate feature map, which is then processed to produce For k = 2 and k = 3, B k is concatenated with the output of the downsampling convolutional layer Conv d (F k-1 ) and processed to produce
[0118] The MSFD module effectively decomposes the input blurred image into feature representations at different scales, facilitating subsequent image processing tasks. As the scale decreases, the channel dimension increases from 32 to 64 and then to 128, capturing more detailed features at each level. This module effectively decomposes the image, facilitating subsequent processing of blurs in different scale ranges.
[0119] The temporal convolutional network (TCN) of this embodiment processes the output feature map MSFD from the multi-scale feature decomposition (MSFD) module out , capturing the temporal features in the feature map and simulating the motion trajectory of the image over time. The TCN consists of three main components: a decomposition network, a temporal convolutional network, and an image reconstruction network. These components work together to optimize the feature information related to different scales and positions. When using the temporal convolutional network to capture temporal features, the working process of these components is as follows:
[0120] The decomposition network decomposes the input feature map into multiple frames, assuming that the input is a superposition of several relatively clear images. This decomposition is achieved through a convolutional layer that expands the input into multiple frames, with each frame representing a potential clear image in the sequence. The input feature map extracts features through the convolutional layer, and the expression is as follows:
[0121]
[0122] This convolutional operation keeps the number of channels as C k . Furthermore, the extracted feature Z k is decomposed into multiple frames through another convolutional layer, and the expression is as follows:
[0123]
[0124] This convolutional operation converts the feature into C k ×T channels, where T is the number of frames, represents the T-th frame at the k-th scale. The decomposed frame D k is reshaped into (T, C k , H k , W k ), enabling the present invention to model the input as a superposition of T relatively clear and varying images. [·;·] represents the concatenation of feature maps along the frame dimension, and in this embodiment, T = 5 is set.
[0125] The temporal convolutional network performs one-dimensional temporal convolution on the decomposed frames to extract temporal features. This process treats the feature values at different positions in the feature map as evolving along the temporal dimension. The operation uses grouped 1D convolution to capture the dynamic changes of features in the temporal dimension at each specific scale.
[0126] For each scale, the temporal features Obtained by applying temporal convolution to the decomposed frame D k as follows:
[0127]
[0128] where w t represents the weight of the temporal convolution kernel corresponding to time t, m is half of the temporal convolution kernel, and can be expressed as κ is the size of the temporal convolution kernel. When m is not an integer, it is rounded down. i, j, and c represent the element at the i-th row and j-th column of the feature map in the c-th channel. In this embodiment, the convolution kernel κ of the one-dimensional temporal convolution is set to 3, so m = 1.
[0129] The image reconstruction network combines the integrated temporal features to generate the final output This process involves a series of convolutional layers with ReLU activation functions to map the temporal features back to the desired output space:
[0130]
[0131] This combined operation refines the features and generates the final output. TCN is crucial for improving the quality of the reconstructed image by extracting and integrating temporal features, and helps to better restore details and maintain the temporal consistency of the image content by utilizing the temporal information in features at different scales.
[0132] The Edge Attention Guided Reconstruction (EAGR) module in this embodiment aims to enhance image reconstruction by leveraging edge information, guiding the network to focus on key edge features, and improving the clarity and quality of the reconstructed image. This module is crucial for improving the clarity and quality of the reconstructed image by effectively restoring edge details. The EAGR module consists of several key components that work together to achieve edge-guided reconstruction. When enhancing the reconstruction clarity through the edge attention guided reconstruction module, as Figure 4 shown, it includes the following steps:
[0133] Processing the blurred image B at the k-th scale k using the Sobel operator to calculate the horizontal and vertical gradients, respectively obtaining and and Both are single-channel feature maps. The Sobel kernel is initialized as follows:
[0134]
[0135] where S x and S y are the Sobel kernels for calculating the horizontal and vertical gradients. The edge map is generated by calculating the gradient magnitude and is also a single-channel feature map:
[0136]
[0137] Among them, represents an edge map derived from the gradient magnitude. To adjust the channel dimension and generate an attention map, through two consecutive 1×1 convolutions followed by batch normalization:
[0138]
[0139] where Conv1 represents the 1×1 convolution operation and BN represents batch normalization.
[0140] Then, by multiplying the normalized edge map and the temporal features at the corresponding positions to adjust the weights of different regions:
[0141]
[0142] where ⊙ represents the Hadamard product, representing the multiplication of elements at the corresponding positions. This process enhances the reconstruction by focusing on edge details and improves the overall image quality.
[0143] In the encoder block (EB k ) and decoder block (DB k ) of this embodiment, k represents different scales, and they share the same structure, aiming to extract deeper spatial features. The basic module of the encoder block and decoder block is the residual block, as Figure 5 (a) shows, which consists of two convolutional layers with a kernel size of 3×3 and has a residual connection to ensure that the size and number of channels of the feature map remain unchanged before and after convolution, facilitating the serial connection of multiple layers. As Figure 5 (b) shows, the encoder block or decoder block is formed by serially connecting n residual blocks.
[0144] The feature attention module (FAM) of this embodiment aims to enhance feature fusion by integrating the attention mechanism, thereby optimizing the combined features. As Figure 6 shown, the output of the encoder block at the previous scale and the output of the edge attention-guided reconstruction module at this scale are used by the feature attention module corresponding to the k + 1-th scale. When enhancing feature fusion by integrating the attention mechanism, the steps are as follows:
[0145] The input feature is first passed through a downsampling convolutional layer to reduce the spatial dimension by half while maintaining the number of channels. The expression is as follows:
[0146]
[0147] downsampled and Multiply the corresponding elements of the two through the Hadamard product to capture their interaction information, obtaining a fused feature map The expression is as follows:
[0148]
[0149] This operation combines features to form an initial fused feature map. The fused feature map is processed by a Dual Attention Depthwise Separable Convolution (DADSC) module, which integrates channel and spatial attention mechanisms as well as depthwise separable convolution. First, the fused feature map is processed by depthwise separable convolution, obtaining a fused feature map The expression is as follows:
[0150]
[0151] where DConv represents depthwise separable convolution, Conv1 represents a 1×1 convolution operation for adjusting the channel dimension. BN represents batch normalization to stabilize the feature values.
[0152] Then, channel attention calculation and spatial attention calculation are performed on the fused feature map The attention mechanism in DADSC combines channel and spatial attention to precisely focus on key features. The channel attention calculation is:
[0153]
[0154] where is the channel attention map, AP and MP are average and max pooling operations respectively. F1 and F2 are fully connected layers for reducing and expanding the channel dimension. σ represents the sigmoid function. The spatial attention calculation is:
[0155]
[0156] where is the spatial attention map, Conv S is the depthwise separable convolution for generating spatial attention weights to reduce computational complexity. Then, a spatial attention map is generated through the sigmoid activation function. This map is used to selectively focus on important spatial regions of the feature map. The combined attention mechanism applies these weights to the feature map:
[0157]
[0158] The final output of the FAM module at the k+1-th scale can be expressed as:
[0159]
[0160] In this embodiment, k = 1 or 2. This formula ensures that the FAM module adaptively processes features of different scales and enhances the overall feature representation. As the input of EB2, is the input of EB3, and the input of EB1 is
[0161] As Figure 7 shown, the multi-scale feature fusion (MSFF) module of this embodiment aims to merge features of different scales, which is crucial for subsequent image-related tasks. When the MSFF module merges features of different scales, the MSFF module is defined by the following equation:
[0162]
[0163] where the features are fused in ascending order of scale, and this alignment ensures that features from different scales can be effectively combined to capture high-level semantic information and detailed texture information. [·;·] represents the concatenation operation along the channel dimension, Conv represents the convolution operation, ↑ represents upsampling, and ↓ represents downsampling until the feature maps are aligned. represents the output of the multi-scale feature fusion module at the η-th scale, and k represents the number of different scales.
[0164] In this embodiment, k = 3. Therefore, the equations of the multi-scale feature fusion modules corresponding to the first scale and the second scale are as follows:
[0165]
[0166] The MSFF module enhances the model's ability to process complex image data by integrating multi-scale features, thereby improving the performance of tasks such as image deblurring and reconstruction. This method allows the network to utilize complementary information from various scales, thus achieving a more robust and accurate representation of image features.
[0167] In this embodiment, when the output feature images of the decoder blocks corresponding to each scale are combined through residual connections after aligning with the blurred images at the corresponding scales, it specifically includes:
[0168] Concatenate the output feature image of the decoder block corresponding to the current scale with the aligned result corresponding to the previous scale after transposed convolution, and then perform a convolution operation to obtain the aligned result corresponding to the current scale. If the current scale is the last scale, perform a convolution operation on the output feature image of the decoder block corresponding to the current scale to obtain the aligned result;
[0169] The alignment result at the current scale is concatenated with the blurred image residual at the current scale to obtain the predicted output at the current scale.
[0170] As Figure 2 shown, the outputs from the multi-scale feature fusion (MSFF) module, and serve as the inputs to decoder blocks DB1 and DB2 respectively, while decoder block DB3 receives its input from EB3. At each scale, the output feature map undergoes specific processing to align with the corresponding B k and is combined through residual connections to produce the final output
[0171] At scale k = 3, after a convolutional operation with a kernel of 3×3 and matching the channel dimension of B3, and then through a residual connection with B3 to obtain This step utilizes the residual connection to combine the processed predicted output with the information of the original blurred image at this scale, retaining some features of the original image and contributing to improving the accuracy of the deblurred image.
[0172] At scale k = 2, is concatenated with the (denoted as ) after transposed convolution, and the result passes through a convolutional layer and is connected with B2 through a residual connection to obtain The transposed convolution operation adjusts the scale of to match that of for convenient concatenation and fusion, which can integrate the feature information of different scales and further optimize the deblurring effect.
[0173] At scale k = 1, concatenates the feature (denoted as ) after transposed convolution, then passes through a convolutional layer, and finally generates This process fully integrates the feature information of the three scales. Through transposed convolution and convolutional operations, the predicted outputs at different scales are adjusted and fused, and finally a deblurred image with rich details and accurate information is obtained.
[0174] The above process is formulated as:
[0175]
[0176] where Conv k3 represents a convolutional operation with an input channel C k , an output channel of 3, a kernel size of 3×3, a padding of 1, and a stride of 1. is the predicted output at scale k, and Τ represents the transposed convolution operation, which doubles the size of the feature map.
[0177] In step S101 of this embodiment, during the training process of the proposed multi-scale spatio-temporal guidance model (MSTSGM), a composite loss function is used to optimize the performance. This loss function is a weighted sum of three independent loss components: content loss, multi-scale reconstruction loss, and temporal smoothness loss. Each component aims to address different aspects of image deblurring to enhance the overall quality of the reconstructed image, where:
[0178] The content loss measures the difference between the model output and the true sharp image. In this embodiment, a multi-scale content loss function is used, and the content loss (L cont ) is defined as:
[0179]
[0180] where K is the number of scales, and n k is the total number of elements at scale k. and I k represent the deblurred image and the true image at scale k, respectively.
[0181] The multi-scale reconstruction loss minimizes the distance between the input and output in the feature space, which is particularly important for recovering the high-frequency components lost during the deblurring process. The multi-scale reconstruction loss (L MSFR ) measures the L1 distance between the multi-scale true and deblurred images in the frequency domain:
[0182]
[0183] where, denotes the fast Fourier transform, which converts the image signal to the frequency domain.
[0184] The temporal smoothness loss is used to constrain the continuity between the reconstructed frames, ensuring that the generated temporal sequence features remain smooth between adjacent frames. This temporal smoothness loss (L temporal ) is defined as:
[0185]
[0186] where, is the total number of elements at scale k at time t, represents the feature at scale k and frame number t. The goal of the temporal smoothness loss is to reduce the abrupt changes between adjacent frames, usually measured by the L1 loss.
[0187] The total loss of this embodiment is the weighted sum of the above three loss components, defined as:
[0188] L total = αLcont +βL MSFR +βL temporal ,
[0189] The present invention sets α = 1, β = 0.1, and γ = 0.1 as weight parameters for balancing the contributions of different loss components.
[0190] During the model training process of this embodiment, the number of residual blocks n in the encoder block and the decoder block is set to 8, the initial learning rate is set to 0.0001, and it decays to 0.5 times the original every 500 batches. The number of iterations is 3000, the batch size is 4, and the optimizer is Adam. The finally trained multi-scale spatio-temporal guidance model can achieve a good image deblurring effect.
[0191] Embodiment 2
[0192] This embodiment proposes an image deblurring system based on a multi-scale spatio-temporal guidance model, including a microprocessor and a computer-readable storage medium connected to each other. The microprocessor is programmed or configured to implement the steps of image deblurring of the multi-scale spatio-temporal guidance model in Embodiment 1. Specifically, the microprocessor in this embodiment includes the following modules:
[0193] A multi-scale feature decomposition module for decomposing a blurred image into feature maps of different scales;
[0194] A temporal convolutional network for capturing temporal features in the feature map and simulating the motion trajectory of the image in time;
[0195] An edge attention-guided reconstruction module for using edge information to guide the network to focus on key edge features and improve the clarity and quality of the reconstructed image;
[0196] A feature attention module for encoding the output of the encoder block of the previous scale and the output of the edge attention-guided reconstruction module of this scale by integrating the attention mechanism to enhance the input corresponding to the encoder block after feature fusion;
[0197] A multi-scale feature fusion module for merging the output features of the encoder blocks of different scales and then inputting them into the corresponding decoder block for decoding.
[0198] A multi-scale output module for aligning the output feature images of the decoder blocks corresponding to each scale with the blurred image at the corresponding scale and then combining them through a residual connection to obtain the predicted output corresponding to each scale.
[0199] In summary, the present invention proposes an image deblurring method (MSTSGM) based on a multi-scale spatio-temporal guidance model, and proposes a corresponding image deblurring system to perform a novel and effective single-image deblurring method, overcoming the limitations of previous models and providing excellent performance in restoring image clarity.
[0200] The present invention effectively solves the problem of image deblurring with complex non-uniform motion blur by integrating innovative components such as multi-scale feature decomposition, temporal convolutional network, and edge attention-guided reconstruction. MSTSGM can effectively decompose the input blurred image into feature representations of different scales, comprehensively solving the blur problems in different scale ranges. The temporal convolutional network captures the temporal features in the feature map, simulating the temporal movement trajectory of the image to capture the dynamic changes during the blurring process. The edge attention-guided reconstruction module uses edge information to guide the network to focus on key edge features, significantly improving the clarity and quality of the reconstructed image. The multi-scale feature fusion module enhances the model's ability to process complex image data by merging features of different scales. The present invention provides excellent performance in restoring image clarity, overcoming the limitations of the prior art in dealing with complex non-uniform motion blur.
[0201] The above are only the preferred embodiments of the present invention and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention shall fall within the scope of the protection of the technical solution of the present invention.
Claims
1. An image deblurring method based on a multi-scale spatiotemporal guided model, characterized in that: include: A multi-scale spatiotemporal guidance model is established and trained, and then the blurred image is input into the trained multi-scale spatiotemporal guidance model to obtain deblurred images of different scales. The image deblurring step of the multi-scale spatiotemporal guidance model includes: Obtain a blurred image, and decompose the blurred image into feature maps of different scales through a multi-scale feature decomposition module; For the feature maps of each scale, a temporal convolutional network is used to capture the temporal features, and the edge attention is used to guide the reconstruction module to enhance the reconstruction clarity, thus obtaining the reconstructed feature maps of each scale. The reconstructed feature map of the first scale is directly input into the corresponding encoder block for encoding. For the reconstructed feature maps of other scales, the corresponding feature attention module is used to enhance the feature fusion of the output of the encoder block of the previous scale and the output of the edge attention-guided reconstruction module of this scale by integrating the attention mechanism, and then input into the corresponding encoder block for encoding; The output of the encoder block of the last scale is directly input into the corresponding decoder block. For other scales, the corresponding multi-scale feature fusion module is used to merge the features of different scales and then input into the corresponding decoder block for decoding. Then, the output feature image of the decoder block corresponding to each scale is aligned with the blurred image at the corresponding scale, and then the prediction output corresponding to each scale is obtained through residual connection combination to obtain the deblurred images at different scales.
2. The image deblurring method based on the multi-scale spatiotemporal guidance model according to claim 1, characterized in that: When the blurred image is decomposed into feature maps of different scales through the multi-scale feature decomposition module, the expression is as follows: in, represents the feature map of the kth scale, B k is the blurred image of the kth scale obtained from the blurred image B1 by bilinear interpolation. The scale factor is 0.5 to the k-1th power. [·;·] represents the concatenation operation along the channel dimension. Conv represents the convolution operation. Conv d represents the downsampling convolution operation, F k-1 Represents the feature map of the k-1th scale The feature map of the last convolution layer is input into the convolution operation.
3. The image deblurring method based on the multi-scale spatiotemporal guided model according to claim 1, characterized in that: When using a temporal convolutional network to capture temporal features, it includes: The input feature map is passed through the convolution layer to extract features. The expression is as follows: in, Represents the feature map of the kth scale, and Conv represents the convolution operation; The extracted features are decomposed into multiple frames through the convolution layer, and the expression is as follows: Where T is the number of frames, represents the Tth frame of the kth scale; One-dimensional temporal convolution is performed on the decomposed frames to extract temporal features, as expressed as follows: Among them, w t represents the weight of the time convolution kernel corresponding to time t, m is half of the time convolution kernel, κ is the size of the temporal convolution kernel. When m is not an integer, it is rounded down. i, j, c represents the element in the i-th row and j-th column of the feature map under the c-th channel. Combine the integrated temporal features to generate the final output The expression is as follows: Among them, ReLU is the activation function.
4. The image deblurring method based on a multi-scale spatiotemporal guided model according to claim 1, characterized in that: When the reconstruction clarity is enhanced through the edge attention guided reconstruction module, it includes: For the blurred image B of the kth scale k Use the Sobel operator to calculate the horizontal gradient and vertical gradient According to the horizontal gradient and vertical gradient Compute edge graph The expression is as follows: Edge Graph Perform batch normalization, the expression is as follows: Among them, Conv1 represents a 1×1 convolution operation, and BN represents batch normalization; For the normalized edge map And time characteristics Multiply to adjust the weights of different areas. The expression is as follows: Here, ⊙ represents the Hadamard product.
5. The image deblurring method based on a multi-scale spatiotemporal guided model according to claim 1, characterized in that: When using the corresponding feature attention module to enhance feature fusion by integrating the attention mechanism on the output of the encoder block of the previous scale and the output of the edge attention-guided reconstruction module of this scale, it includes: The output of the encoder block at the kth scale By downsampling the convolutional layer to reduce the spatial dimension by half while maintaining the number of channels, the expression is as follows: The output of the encoder block at the downsampled kth scale and the output of the k+1th scale edge attention-guided reconstruction module The Hadamard product is used to multiply the elements at corresponding positions to capture their mutual information and obtain the fused feature map. The expression is as follows: Fusion feature map Through the depth-separable convolution process, the fused feature map is obtained The expression is as follows: Among them, DConv represents depthwise separable convolution, Conv1 represents 1×1 convolution operation, which is used to adjust the channel dimension, and BN represents batch normalization to stabilize the feature value; Fusion feature map Perform channel attention calculation and spatial attention calculation, the expressions are as follows: in, is the channel attention map, is the spatial attention map, AP and MP are the average and maximum pooling operations respectively, F1 and F2 are fully connected layers used to reduce and expand the channel dimension, σ represents the sigmoid function, Conv S It is a depth-wise separable convolution that generates spatial attention weights to reduce computational complexity; According to the fusion feature map Channel Attention Map Spatial Attention Map and the output of the encoder block at the downsampled kth scale The final output of the edge attention-guided reconstruction module at the k+1th scale is calculated and expressed as follows: in, 6. The image deblurring method based on a multi-scale spatiotemporal guided model according to claim 1, characterized in that: When the multi-scale feature fusion module merges features of different scales, the expression is as follows: in, represents the output of the multi-scale feature fusion module of the ηth scale, k represents the number of different scales, [·;·] represents the concatenation operation along the channel dimension, Conv represents the convolution operation, ↑ represents upsampling, and ↓ represents downsampling.
7. The image deblurring method based on a multi-scale spatiotemporal guided model according to claim 1, characterized in that: When the output feature image of the decoder block corresponding to each scale is aligned with the blurred image at the corresponding scale and combined through a residual connection, it specifically includes: The output feature image of the decoder block corresponding to the current scale is concatenated with the alignment result corresponding to the previous scale after transposed convolution, and then the convolution operation is performed to obtain the alignment result corresponding to the current scale. If the current scale is the last scale, the output feature image of the decoder block corresponding to the current scale is convolved to obtain the alignment result; The alignment result of the current scale is connected with the blurred image residual of the current scale to obtain the prediction output of the current scale.
8. The image deblurring method based on the multi-scale spatiotemporal guidance model according to claim 7, characterized in that: When the number of different scales is 3, the prediction output expression of each scale is as follows: in, is the predicted output of the kth scale, B k is the blurred image of the kth scale, Conv k3 Indicates that the input channel C k , a convolution operation with 3 output channels, a kernel size of 3×3, padding of 1, and a stride of 1, where Τ represents a transposed convolution operation, and The output feature images of the decoder blocks at the 1st scale, 2nd scale, and 3rd scale, respectively.
9. The image deblurring method based on a multi-scale spatiotemporal guided model according to claim 1, characterized in that: During the training model, the loss function expression is as follows: L total =αL cont +βL MSFR +γL temporal , Among them, L cont is the content loss, and the expression is as follows: Where K is the number of scales, n k is the total number of elements in the kth scale, and I k Represent the deblurred image and the real image at the kth scale respectively; L MSFR is the multi-scale reconstruction loss, expressed as follows: in, represents fast Fourier transform; L temporal is the time smoothing loss, defined as: in, is the total number of elements of scale k at time t, Representing features with scale k and number of frames t, the goal of the temporal smoothness loss is to reduce the mutations between adjacent frames, which is usually measured by the L1 loss.
10. An image deblurring system based on a multi-scale spatiotemporal guidance model, characterized in that: include: Multi-scale feature decomposition module, used to decompose the blurred image into feature maps of different scales; Temporal convolutional network, which is used to capture the temporal features in the feature map and simulate the motion trajectory of the image in time; The edge attention guided reconstruction module is used to use edge information to guide the network to focus on key edge features and improve the clarity and quality of the reconstructed image; The feature attention module is used to enhance the feature fusion by integrating the attention mechanism to encode the encoder block corresponding to the input after the output of the encoder block of the previous scale and the output of the edge attention-guided reconstruction module of this scale. The multi-scale feature fusion module is used to merge the output features of encoder blocks of different scales and input them into the corresponding decoder blocks for decoding. The multi-scale output module is used to align the output feature image of the decoder block corresponding to each scale with the blurred image at the corresponding scale, and then combine them through residual connections to obtain the prediction output corresponding to each scale.
Citation Information
Cited By
Blurred image sharpening method, system, equipment and medium
CN120807354A