Video motion blur removing method combining single-frame and multi-frame features
Through the U-Net network and the adaptive fusion module combined with multi-frame and single-frame features, the existing video defuzzing method has solved the problem of poor computational complexity and nonlinear motion processing effect, and achieved efficient recovery of clear video frames.
Patent Information
- Application Number
- CN202510202879.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-11
AI Technical Summary
The existing video defuzzing methods have high computational complexity and are difficult to deal with complex or large-scale motions, especially in nonlinear motion scenarios, and have poor edge fuzzing effects.
U-Net network is used for feature extraction and fusion, combining multi-frame and single-frame features, and dynamically adjusting the weights through the adaptive fusion module, using intermediate information and lightweight packet space-time shift algorithm to capture inter-frame relationships, reducing computational complexity.
It improves the accuracy and efficiency of video defuzzing, can effectively restore complex motion and edge texture details, reduce calculation costs, and adapt to information fusion in different scenarios.
Smart Images

Figure CN120298253A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video restoration, and particularly to a video motion deblurring method combining single-frame and multi-frame features. Background Art
[0002] In recent years, video deblurring methods based on deep neural networks have received extensive attention. Early video deblurring research only regarded it as an extension of image deblurring and failed to fully utilize the redundant information between video frames. Compared with single-image deblurring, multi-image deblurring can utilize additional information between images to compensate for the information lost during imaging to a certain extent. These multi-blurred images often originate from video frames or stereo image sets.
[0003] The video deblurring process relies on efficient spatio-temporal modeling techniques, aiming to capture complementary information between adjacent frames to enhance the restoration effect. Early strategies for video or multi-frame deblurring were based on a core observation: the degree of blur in each frame of a video is not consistent. The strategy proposed by Law et al. directly identifies and selects clear pixels in multiple frames as components of the final image. In the study of Wang et al., the researchers first used homography and path-based alignment to align multiple frames, and then aggregated sharp pixels to generate a potential sharp frame. Some methods also combined motion modeling techniques to estimate the blur kernel that varies over time and space and perform deconvolution accordingly, as described by Lee et al. However, these schemes often come with high computational costs and rely on artificial prior knowledge to ensure the optimization convergence of the energy function, which greatly hinders their practical application scope. There are also several methods that use convolutional networks for multi-frame fusion without explicit alignment, but their performance is often suboptimal.
[0004] To enhance the deblurring effect, researchers have explored various complementary multi-frame alignment strategies aimed at optimizing the temporal modeling process. Kim et al. proposed a spatio-temporal flow to establish correspondences across multiple frames for video restoration, which transfers spatio-temporal information from multiple frames to the current frame. In the study of Zhou et al., deformable convolution and dynamic convolution were also applied to implicitly align adjacent frame features, which achieved better deblurring performance. The proposal of the dynamic filter architecture in 2018 enabled researchers to explore methods that can solve the problem of adjacent video frame alignment, such as STFAN and EDVR. STFAN proposed a filter adaptive convolution layer, regarding the alignment and deblurring processes as two filter adaptive convolutions in the feature domain. STFAN applies element-wise convolution kernels to the video frames to be restored, adaptively performing pixel-level feature transformation on the video frames according to the input. STFAN integrates adjacent frame alignment and deblurring into the same framework without explicit motion estimation, which solves the problem of high computational cost in traditional optical flow estimation. EDVR proposed a multi-scale adjacent video frame alignment module with the help of a deformable convolutional network. This module uses a pyramid structure to perform alignment of each adjacent frame with the reference frame in a coarse-to-fine manner at the feature layer to handle large and complex motions. The alignment process utilizes the geometric deformation modeling ability of deformable convolution, enabling the model to handle different degrees of inter-frame jitter.
[0005] The existing methods also have the following deficiencies:
[0006] 1. Many video deblurring methods adopt 3D convolutional networks, optical flow networks, or network structures combined with spatio-temporal attention mechanisms. Although they can effectively capture spatio-temporal information, their computational costs are high, introducing more computational complexity.
[0007] 2. Traditional methods rely on optical flow or deformable convolution for temporal alignment, but these methods have limitations in dealing with complex or large-scale motions, especially the processing effect for non-linear motions is not good.
[0008] 3. The existing alignment-based deblurring methods require complex training and optimization, and introduce additional parameters and computational overhead, which makes their training and application more difficult.
[0009] 4. In some scenarios, some of the existing methods have poor processing effects on edge blurring, especially when dealing with large-scale motions or non-linear motions, problems may easily occur. Summary of the Invention
[0010] Aiming at the deficiencies of the existing technology, the present invention proposes a video de - motion - blur method that combines single - frame and multi - frame features. In the feature extraction stage, three consecutive blurred frames, namely the current frame and the two adjacent frames before and after it, are input into a stacked U - Net network for feature extraction. In the feature fusion stage, the features of the current frame and the adjacent frames before and after it obtained through the feature extraction stage are input into the multi - frame branch to obtain multi - frame features, and then the current - frame features are input into the single - frame branch to obtain single - frame features. Finally, the multi - frame features and the single - frame features are input into the adaptive fusion single - frame and multi - frame feature module, and the weights generated by the intermediate information extraction module are used to guide the adaptive fusion single - frame and multi - frame feature module to generate fusion features. In the feature reconstruction stage, the same network structure as in the feature extraction stage is used to process the fusion features to obtain the reconstructed clear frame, specifically including:
[0011] Step 1: Input three consecutive degraded blurred frames into 5 stacked U - Net networks for feature extraction respectively to obtain the current - frame features and the features of two adjacent frames respectively;
[0012] Step 2: Through the multi - frame feature extraction module, utilize the correlation between adjacent frames to capture the information in the temporal and spatial dimensions to obtain multi - frame feature B iout , including:
[0013] Step 21: Aggregate the current - frame features and the features of two adjacent frames extracted through the feature extraction stage to obtain the primary repaired frame A i , as the input of the multi - frame feature extraction module. First, input the primary repaired frame A i into the first - layer encoder to obtain the first feature A i1 ;
[0014] Step 22: Downsample the first feature A i1 and pass it through the second - layer encoder to obtain the second feature A i2 ;
[0015] Step 23: Downsample the second feature A i2 and pass it through the third - layer encoder to obtain the third feature A i3 ;
[0016] Step 24: Perform a decoding operation on the third feature A i3 . The first - layer decoder includes a spatial shift module. Divide the third feature A i3 equally along the number of channels to obtain two feature groups, including the first - channel feature and the second - channel feature Divide the second - channel feature equally along the number of channels and then perform spatial shifts in different directions and aggregate it to obtain the shifted feature Finally, combine the first - channel feature and the shifted feature The fused feature B is obtained through the fusion layer i3 ;
[0017] Step 25: Upsample the fused feature B i3 and obtain the first decoded feature B through the second decoder i2 ;
[0018] Step 26: Upsample the first decoded feature B i2 and obtain the multi-frame feature B through the third decoder iout ;
[0019] Step 3: Use the intermediate information extraction module of the multi-frame feature extraction module to generate the corresponding intermediate information E i , specifically, the second feature A obtained by the second encoder i2 and the first decoded feature B obtained by the second decoder i2 respectively pass through a 3×3 convolutional layer and an activation function layer, and then aggregate them to obtain the intermediate information E i , and the intermediate information E i is used by the feature fusion module to guide the fusion to recover the intermediate repair frame I i ;
[0020] Step 4: Use the single-frame feature extraction module to extract single-frame features, which consists of two identical extraction modules, both including multi-scale convolution and multi-shape convolution, including:
[0021] Step 41: Use the current frame feature after the feature extraction stage as the input of the first extraction module, and first input it into the multi-scale convolution to obtain the multi-scale feature X′ isc ;
[0022] Step 42: Input the multi-scale feature X′ isc into the multi-shape convolution to obtain the multi-shape feature X′ ish ;
[0023] Step 43: Input the multi-shape feature X′ ish into the channel attention module to extract important information, and then introduce channel-wise multiplication, channel-wise addition, skip connection and convolution group to further extract information to obtain the first output feature X′ of the first extraction module ione ;
[0024] Step 44: Use the first output feature X′ ione as the input of the second extraction module, and perform the same processing on it to obtain the second output feature X′ itwo , and fuse the outputs of the two extraction modules as the final single-frame feature X′ iout ;
[0025] Step 5: Input multi-frame feature B iout and single-frame feature X' iout into the adaptive fusion single-frame and multi-frame feature module. Based on the different importance of multi-frame features and single-frame features in video content, weights are assigned through intermediate information E i . The adaptive fusion single-frame and multi-frame feature module includes a preliminary fusion module and a deep fusion module, including:
[0026] Step 51: The preliminary fusion module has a dual-branch input structure. Take the multi-frame feature B iout from the multi-frame feature extraction module and the single-frame feature X' iout from the single-frame feature extraction module as the two inputs, and add the two inputs channel by channel to obtain the fused feature C;
[0027] Step 52: Pass the fused feature C through a global average pooling layer to obtain the pooled feature C g ;
[0028] Step 53: Let the pooled feature C g pass through a convolutional layer, an activation function layer, and be normalized along the second dimension of the tensor to obtain the smoothed feature C s ;
[0029] Step 54: Multiply the smoothed feature C s with the multi-frame feature B iout and the single-frame feature X' iout channel by channel respectively to obtain the first preliminary fused feature C b and the second preliminary fused feature C x ;
[0030] Step 55: The deep fusion module adopts a three-branch input structure. The three inputs are the intermediate information E i obtained by the intermediate information extraction module, the first preliminary fused feature C b and the second preliminary fused feature C x obtained by the preliminary fusion module. Specifically, input the intermediate information E i into two branches respectively. One branch consists of a convolutional group to obtain the first deep fused feature E ig ; the other branch passes through a common convolutional layer to obtain the second deep fused feature E is ;
[0031] Step 56: Input the first deep fused feature E ig and the second deep fused feature E is into a convolutional layer to obtain the deep feature mask D. Each element in the deep feature mask D represents the weight at the corresponding spatial position;
[0032] Step 57: Upsample the depth feature mask D to adjust the size of the mask, obtaining the resized fused feature D up ;
[0033] Step 58: Multiply the fused feature D up channel - by - channel with the first preliminary fused feature C b and the second preliminary fused feature C x respectively, obtaining the first masked feature C′ b guided by the mask and the second masked feature C′ x ;
[0034] Step 59: Add the first masked feature C′ b and the second masked feature C′ x channel - by - channel, obtaining the intermediate repaired frame I i ;
[0035] Step 6: Aggregate the initially input blurred frame, the primary repaired frame A i from the feature extraction stage and the intermediate repaired frame I i from the feature fusion stage, obtaining the aggregated feature Y e ;
[0036] Step 7: Input the aggregated feature Y e into 5 stacked U - Net networks for feature reconstruction, and output the reconstructed high - definition frame Y i .
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1. In both the feature extraction and feature reconstruction stages of the present invention, 5 stacked U - Net networks are adopted. On each U - Net network, average pooling and bilinear upsampling are used to adjust the feature resolution, and residual blocks are used to extract features. Compared with the previous feature extraction and feature reconstruction methods, this structure improves the learning ability of the model, enabling the network to more accurately distinguish subtle blur and motion information. At the same time, the output of each U - Net network can not only be used for de - blurring but also provide input for subsequent U - Net modules, thus realizing the multiple utilization of features and better restoring clear images.
[0039] 2. The present invention fuses multi - frame feature and single - frame feature information for video de - blurring. Compared with the existing method that uses the aggregation between adjacent frames to restore each clear frame, it can not only restore the blur caused by complex or large - amplitude motion but also well restore the detailed features of edges or textures.
[0040] 3. In the multi-frame branch of the present invention, a lightweight grouped spatio-temporal shift algorithm is used, which is a lightweight and direct technique that can implicitly capture the correspondence between multiple frames to achieve aggregation. Compared with mainstream deep learning methods such as complex network structures like optical flow estimation and deformable convolution, it reduces the computational cost and computational complexity.
[0041] 4. An intermediate information extraction module is also designed in the multi-frame branch of the present invention. In previous methods, the intermediate information generated during the feature extraction process was often ignored, but this information is very important in video deblurring. When the video content changes drastically, the intermediate information captured by the intermediate information extraction module can make up for the deficiencies of the information extracted by the multi-frame feature extraction module. At the same time, it has a relatively high weight and can be used to evaluate whether the multi-frame branch can accurately capture features.
[0042] 5. The single-frame branch of the present invention uses a multi-scale convolution and a multi-shape convolution. Compared with the previous method that only uses a single convolution module, it can more accurately restore the spatial information of different scales and the edge and texture information in different directions. These information make up for the deficiencies of the features extracted by the multi-frame branch in terms of edges and textures.
[0043] 6. The present invention proposes an adaptive fusion single-frame multi-frame feature module. Compared with the existing multi-modal fusion or multi-scale fusion modules, this module can dynamically fuse the inter-frame and intra-frame features, allocate appropriate weights, and enable the model to adaptively select effective information in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is the structural diagram of the deblurring network of the present invention;
[0045] Figure 2 is the structural diagram of the multi-frame feature extraction module of the present invention;
[0046] Figure 3 is the structural diagram of the spatial shift module of the present invention;
[0047] Figure 4 is the structural diagram of the single-frame feature extraction module of the present invention;
[0048] Figure 5 is the structural diagram of the adaptive fusion single-frame multi-frame feature module of the present invention;
[0049] Figure 6 is the partial test result on the GOPRO dataset;
[0050] Figure 7 is the partial test result on the DVD dataset;
[0051] Figure 8are partial test results of the BSD dataset; Detailed implementation mode
[0052] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in combination with the specific implementation modes and with reference to the accompanying drawings. It should be understood that these descriptions are exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0053] The following is a detailed description in combination with the accompanying drawings.
[0054] Aiming at the problems existing in the prior art, the present invention proposes a new video de-motion blur method by using the feature information of single-frame and multi-frame, which can not only promote the development of the video de-motion blur field, but also promote the development of deep learning technology. Specifically, the present invention proposes a video de-motion blur method using single-frame and multi-frame features. In the feature extraction stage of this method, three blurred frames, namely the current frame and the two adjacent frames before and after, are input into 5 stacked U-Net networks for feature extraction. In the feature fusion stage, first, the features of the current frame and the adjacent frames before and after obtained through the feature extraction stage are input into the multi-frame branch to obtain multi-frame features. Then, the current frame feature is input into the single-frame branch to obtain single-frame features. Finally, the multi-frame features and the single-frame features are input into the adaptive fusion single-frame multi-frame feature module for adaptively fusing single-frame and multi-frame information, and the weight generated by the intermediate information extraction module is used to guide the adaptive fusion single-frame multi-frame feature module to generate fusion features. In the feature reconstruction stage, the same network structure as that in the feature extraction stage is used to process the fusion features to obtain the reconstructed clear frame, and the network structure is as Figure 1 shown.
[0055] Step 1: Input three consecutive degraded blurred frames into the stacked U-Net network respectively, and use 5 stacked U-Net networks for feature extraction to obtain the current frame feature X i ' and the features of two adjacent frames. Among them, the middle frame X i is used as the current frame, where T represents the number of the blurred frame.
[0056] The U-Net network includes network modules with convolutional layers, activation functions, attention mechanisms and residual connections, and has functions such as feature extraction, non-linear transformation, attention weighting and residual connection.
[0057] This structure improves the learning ability of the model, enabling the network to more precisely distinguish subtle blurring and motion information. At the same time, the output of each U-Net can not only be used for deblurring, but also provide input for subsequent U-Net modules, thus realizing the multiple utilization of features. This allows the model to gradually refine features and better restore clear images. On each U-Net, average pooling and bilinear upsampling are used to adjust the feature resolution, and residual blocks are used to extract features.
[0058] Step 2: Utilize the correlation between adjacent frames through the multi-frame feature extraction module to capture information in the temporal and spatial dimensions, obtaining multi-frame feature B iout , and the multi-frame feature extraction module is as Figure 2 shown, specifically including:
[0059] Step 21: Aggregate the features of the current frame and two adjacent frames extracted during the feature extraction stage to obtain the primary repair frame A i , which serves as the input to the multi-frame feature extraction module; where A i ∈R h×w×c , where h and w represent height and width, and c represents the number of channels. First, input the primary repair frame A i into the first-layer encoder, which consists of a convolutional layer, a pooling layer, and a Relu activation function. After passing through the first-layer encoder, obtain the first feature A i1 ;
[0060] Step 22: Downsample the first feature A i1 and pass it through the second-layer encoder to obtain the second feature A i2 ;
[0061] Step 23: Downsample the second feature A i2 and pass it through the third-layer encoder to obtain the third feature A i3 ;
[0062] Step 24: Perform a decoding operation on the third feature A i3 . The first-layer decoder consists of a spatial shift module, as Figure 3 shown.
[0063] Specifically, evenly divide the third feature A i3 along the number of channels to obtain two feature groups, including the first channel feature and the second channel feature Evenly divide the second channel feature along the number of channels and then perform spatial shifts in different directions and aggregate it to obtain the shifted feature Specifically as Figure 3 shown, perform spatial shifts in 16 directions. Finally, combine the first channel feature and the shifted feature The fused feature B is obtained through the fusion layer i3 . The fusion layer contains two lightweight convolutional blocks. Each block uses pointwise convolution, depthwise convolution, and a gating layer to avoid heavy computations. Through this step, a larger effective receptive field can be obtained, effectively aggregating inter-frame information and reducing the computational overhead.
[0064] Step 25: Upsample the fused feature B i3 and obtain the first decoded feature B through the second decoder i2 ;
[0065] Step 26: Upsample the first decoded feature B i2 and obtain the multi-frame feature B through the third decoder iout ;
[0066] Step 3: Use the intermediate information extraction module of the multi-frame feature extraction module to generate the corresponding intermediate information E i , specifically: The second feature A obtained from the second encoder i2 and the first decoded feature B obtained from the second decoder i2 are respectively passed through a 3×3 convolutional layer and a Leaky ReLU activation function layer, and then aggregated to obtain the intermediate information E i . The intermediate information E i is used by the feature fusion module to guide the fusion to recover the intermediate repair frame I i .
[0067] Step 4: Extract the single-frame feature using the single-frame feature extraction module, which consists of two identical extraction modules, including multi-scale convolution and multi-shape convolution. The schematic diagram of the structure of the single-frame feature extraction module is as shown in Figure 4 . The single-frame branch is used to extract the feature information that has not been fully explored and utilized by the multi-frame branch. It includes the following steps:
[0068] Step 41: Use the current frame feature X i ' after the feature extraction stage as the input to the first extraction module. First, input it into the multi-scale convolution to obtain the multi-scale feature X′ isc . The multi-scale convolution contains multiple convolutional modules with different kernel sizes. First, use 1×1 convolution and 3×3 convolution to reduce the number of channels of the input feature map to reduce the computational cost, and then use 5×5 convolution and 7×7 convolution to extract multi-scale information to obtain the multi-scale feature X′ isc .
[0069] Step 42: Input the multi-scale feature X′ isc into the multi-shape convolution to obtain the multi-shape feature X′ ish。In multi-shaped convolution, in order to better extract edge and texture features in different directions, the present invention uses a common 3×3 convolution layer and horizontal 1×9 and vertical 9×1 convolution layers. The 1×9 convolution kernel is longer in width, which enables it to capture features in a longer range in the horizontal direction of the input data. The 9×1 convolution kernel is longer in height, which enables it to capture features in a longer range in the vertical direction of the input data. The size of the kernel is determined through experiments in the experiment to obtain the multi-shaped feature X′ ish 。
[0070] Step 43: Input the multi-shaped feature X′ ish into the channel attention module to extract more valuable information, and then introduce channel-wise multiplication, channel-wise addition, skip connection and convolution group to further extract information to obtain the first output feature X′ of the first extraction module ione , and the mathematical expression is as follows:
[0071] X′ ionv = Conv(CA(C sh (C sc (X′ i )))⊙X′ ish )⊕X′ i (1)
[0072] Among them, C sh represents the multi-shaped convolution operation, and C sc represents the multi-scale convolution operation.
[0073] Step 44: Use the first output feature X′ ione as the input of the second extraction module, and perform the same processing on it to obtain the second output feature X′ itwo , and fuse the outputs of the two extraction modules as the final single-frame feature X′ iout .
[0074] After two repeated extraction modules, the single-frame feature extraction module of the present invention has stronger single-frame feature extraction and repair capabilities.
[0075] Step 5: Input the multi-frame feature B iout and the single-frame feature X′ iout into the adaptive fusion single-frame multi-frame feature module. Based on the different importance of the multi-frame feature and the single-frame feature in the video content, weights are assigned through intermediate information. The adaptive fusion single-frame multi-frame feature module includes a preliminary fusion module and a deep fusion module.
[0076] Regarding the phenomenon that the importance of the multi-frame feature and the single-frame feature in the video content is different, and different features have different weights, the present invention proposes an adaptive fusion single-frame multi-frame feature fusion module, including a preliminary fusion module and a deep fusion module. The network structure is asFigure 5 as shown
[0077] Step 51: The preliminary fusion module has a dual-branch input structure. The multi-frame features B of the multi-frame feature extraction module iout and the single-frame feature X' of the single-frame feature extraction module iout are used as two inputs, and the two inputs are added channel by channel to obtain the fusion feature C;
[0078] Step 52: The fusion feature C is passed through a global average pooling layer to obtain the pooled feature C g , which can reduce the number of channels to improve the calculation efficiency;
[0079] Step 53: Let the pooled feature C g pass through a convolutional layer, an activation function layer, and perform Softmax normalization along the second dimension of the tensor to obtain the smoothed feature C s ;
[0080] Step 54: The smoothed feature C s is multiplied channel by channel with the multi-frame feature B iout and the single-frame feature X' iout respectively to obtain the first preliminary fusion feature C b and the second preliminary fusion feature C x ;
[0081] Step 55: The deep fusion module adopts a three-branch input structure. The three inputs are the intermediate information E obtained by the intermediate information extraction module i and the first preliminary fusion feature C and the second preliminary fusion feature C b obtained by the preliminary fusion module x . Specifically, the intermediate information E i is input into two branches respectively. One branch consists of a convolutional group to obtain the first deep fusion feature E ig ; the other branch passes through a common convolutional layer to obtain the second deep fusion feature E is . This can not only extract information at different levels but also complement the two sets of information.
[0082] Step 56: The first deep fusion feature E ig and the second deep fusion feature E is are input into a convolutional layer to obtain the deep feature mask D. Each element in the deep feature mask D represents the weight at the corresponding spatial position, and this weight is used to guide the subsequent adaptive feature fusion;
[0083] Step 57: The deep feature mask D is upsampled to adjust the size of the mask to obtain the resized fusion feature D up, which can ensure that the mask is aligned with the feature map in the spatial dimension, while maintaining a smooth transition between pixel values and reducing some artifacts caused by interpolation;
[0084] Step 58: Multiply the fused feature D up channel by channel with the first preliminary fused feature C b and the second preliminary fused feature C x respectively to obtain the first masked feature C' b and the second masked feature C' x ;
[0085] Step 59: Add the first masked feature C' b and the second masked feature C' x channel by channel to obtain the intermediate repaired frame I i ;
[0086] Step 6: Aggregate the initially input blurred frame, the primary repaired frame in the feature extraction stage, and the intermediate repaired frame I i in the feature fusion stage to obtain the aggregated feature Y e ;
[0087] Step 7: Input the aggregated feature Y e into 5 stacked U-Net networks for feature reconstruction, and output the reconstructed high-definition frame Y i .
[0088] The experimental environment of the present invention is specifically: the graphics card is an NVIDIA GeForce RTX 3090 GPU.
[0089] The present invention uses the commonly used GOPRO, DVD, and BSD datasets for video deblurring experiments and compares with existing methods. Among them, the GOPRO and DVD datasets are artificially synthesized blurred datasets, while the BSD is a real-world blurred dataset captured by a dual-camera system.
[0090] The GOPRO dataset contains 33 pairs of clear and blurred videos with a resolution of 1280×720. Among them, 22 pairs of videos with a total of 2103 pairs of frames are used for training, and 11 pairs of videos with a total of 1111 pairs of frames are used for testing.
[0091] The DVD dataset contains 71 pairs of clear and blurred videos with a resolution of 1280×720. Among them, 61 pairs of videos with a total of 5708 pairs of frames are used for training, and 10 pairs of videos with a total of 1000 pairs of frames are used for testing.
[0092] The BSD dataset is a real-world blurred dataset captured using a dual-camera system, containing 80 pairs of clear and blurred videos, a total of 9000 frames, with a resolution of 640×480. Among them, 60 pairs of videos are used for training, and 20 pairs of videos are used for testing.
[0093] The present invention uses two metrics, namely Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM), to evaluate the deblurring results. PSNR is a traditional metric for measuring the error between the deblurring result and the clear reference image, which is calculated based on the mean squared error (MSE) of pixel values. The higher the value of PSNR, the closer the deblurring result is to the reference image. SSIM measures image quality from three aspects: brightness, contrast, and structural similarity, and has a higher correlation with human visual perception. Its value ranges from 0 to 1, and the higher the value, the better the image quality.
[0094] For training, the present invention sets the Batch size to 1, the video length to 6, and the batch size to 256×256. For testing, the present invention sets the video length to 16 and the image size to the original size. The present invention uses the Adam optimizer, and according to the cosine annealing strategy, the learning rate decreases from 1×10 -4 to 1×10 -7 . The loss function used in training is the L1 loss function.
[0095] For testing, the present invention sets the video length to 16 and the image size to the original size, and uses the test sets of 3 datasets for testing.
[0096] The comparison methods include three currently mainstream methods. Method 1: Using temporal sharpness for cascaded deep video deblurring network (TSP); Method 2: Optic flow-guided Transformer network (FGST); Method 3: Deep discriminative spatio-temporal network for efficient video deblurring (DSTNET). Tables 1, 2, and 3 respectively show the comparison results of the PSNR and SSIM metrics of the method of the present invention and other methods.
[0097] Table 1 Comparison of metrics on the GOPRO dataset
[0098]
[0099] Table 2 Comparison of metrics on the DVD dataset
[0100]
[0101] Table 3 Comparison of metrics on the BSD dataset
[0102]
[0103] As can be intuitively seen from Table 1 to Table 3, the objective evaluation indexes of the method of the present invention on the three data sets are slightly higher than those of the existing methods, indicating that the method of the present invention has technological progress.
[0104] Figure 6 , Figure 7 and Figure 8 respectively show partial visual effects of the method of the present invention on the three data sets. As shown in the figure, the deblurred frames processed by the present invention are very similar to the original clear frames, with better restoration of details and edge contours, and distinguishable clear results are restored in the small blurred icon part and the digital part.
[0105] It should be noted that the above specific embodiments are exemplary. Those skilled in the art can come up with various solutions inspired by the disclosed content of the present invention, and these solutions also belong to the disclosure scope of the present invention and fall within the protection scope of the present invention. Those skilled in the art should understand that the description and drawings of the present invention are illustrative and do not constitute a limitation on the claims. The protection scope of the present invention is defined by the claims and their equivalents.
Claims
1. A video de - motion - blur method combining single - frame and multi - frame features, characterized in that, In the feature extraction stage of the method, three consecutive blurred frames, namely the current frame and the two adjacent frames before and after it, are input into the stacked U-Net network for feature extraction. In the feature fusion stage, the features of the current frame and the adjacent frames before and after it obtained through the feature extraction stage are input into the multi-frame branch to obtain multi-frame features, and then the current frame features are input into the single-frame branch to obtain single-frame features; finally, the multi-frame features and single-frame features are input into the adaptive fusion single-frame multi-frame feature module, and the weight generated by the intermediate information extraction module is used to guide the adaptive fusion single-frame multi-frame feature module to generate fused features; In the feature reconstruction stage, the same network structure as that in the feature extraction stage is used to process the fused features to obtain the reconstructed clear frame, specifically including: Step 1: The three consecutive degraded blurred frames are respectively input into 5 stacked U-Net networks for feature extraction to obtain the features of the current frame and the two adjacent frames respectively; Step 2: The multi-frame feature extraction module captures the information in the temporal and spatial dimensions by using the correlation between adjacent frames, and obtains the multi-frame feature B iout , including: Step 21: Aggregate the current frame features extracted in the feature extraction stage and the features of two adjacent frames to obtain the primary repair frame A i , as the input of the multi-frame feature extraction module, first input the primary repair frame A i into the first layer of the encoder to obtain the first feature A i1 ; Step 22: Downsample the first feature A i1 and obtain the second feature A through the second-layer encoder i2 ; Step 23: Downsample the second feature A i2 and obtain the third feature A through the third-layer encoder i3 ; Step 24: For the third feature A i3 perform a decoding operation. The first-layer decoder includes a spatial shift module that divides the third feature A i3 equally along the number of channels into two feature groups, including the first-channel feature and the second-channel feature Divide the second-channel feature equally along the number of channels and then aggregate it after performing spatial shifts in different directions to obtain a shifted feature Finally, the first-channel feature and the shifted feature pass through a fusion layer to obtain a fused feature B i3 ; Step 25: Take the fused feature B i3 Perform upsampling and pass through the second decoder to obtain the first decoded feature B i2 ; Step 26: Upsample the first decoded feature B i2 and obtain multi-frame feature B through the third decoder layer iout ; Step 3: Use the intermediate information extraction module of the multi-frame feature extraction module to generate the corresponding intermediate information E i , specifically, the second feature A obtained from the second-layer encoder i2 and the first decoded feature B obtained from the second-layer decoder i2 are respectively passed through a 3×3 convolutional layer and an activation function layer, and then aggregated to obtain the intermediate information E i . The intermediate information E i is used by the feature fusion module to guide the fusion to recover the intermediate repair frame I i ; Step 4: Use the single-frame feature extraction module to extract single-frame features, which consists of two identical extraction modules, both of which include multi-scale convolution and multi-shape convolution, including: Step 41: Use the current frame features obtained in the feature extraction stage as the input of the first extraction module. First, input them into the multi-scale convolution to obtain multi-scale features X′ isc ; Step 42: Input the multi-scale feature X' isc into the multi-shape convolution to obtain the multi-shape feature X' ish ; Step 43: Input the multi-shape feature X′ ish into the channel attention module to extract important information, and then introduce channel-wise multiplication, channel-wise addition, skip connection and convolutional group to further extract information to obtain the first output feature X′ of the first extraction module ione ; Step 44: Take the first output feature X′ ione as the input of the second extraction module, and perform the same processing on it to obtain the second output feature X′ itwo , and fuse the outputs of the two extraction modules as the final single-frame feature X′ iout ; Step 5: Multiframe feature B iout and single-frame feature X' iout are input into the adaptive fusion single-frame and multi-frame feature module. Based on the different importance of multi-frame features and single-frame features in video content, weights are assigned through intermediate information E i The adaptive fusion single-frame and multi-frame feature module includes a preliminary fusion module and a deep fusion module, and comprises: Step 51: The preliminary fusion module has a dual-branch input structure. The multi-frame feature B of the multi-frame feature extraction module iout and the single-frame feature X' of the single-frame feature extraction module iout are used as two inputs, and the two inputs are added channel by channel to obtain the fusion feature C; Step 52: Obtain the pooled feature C by passing the fused feature C through the global average pooling layer g ; Step 53: Let the pooled feature C g pass through a convolutional layer, an activation function layer, and be normalized along the second dimension of the tensor to obtain the smoothed feature C s ; Step 54: Multiply the smoothed feature C s separately with the multi-frame feature B iout and the single-frame feature X' iout channel by channel to obtain the first preliminary fusion feature C b and the second preliminary fusion feature C x ; Step 55: The deep fusion module adopts a three-branch input structure, and the three inputs are the intermediate information E obtained by the intermediate information extraction module i and the first preliminary fusion feature C obtained by the preliminary fusion module b and the second preliminary fusion feature C x . Specifically, the intermediate information E i is respectively input into two branches. One branch consists of a convolution group to obtain the first deep fusion feature E ig ; the other branch passes through a common convolution layer to obtain the second deep fusion feature E is ; Step 56: Input the first depth fusion feature E ig and the second depth fusion feature E is into a convolutional layer to obtain a depth feature mask D, where each element in the depth feature mask D represents the weight at the corresponding spatial position; Step 57: Upsample the depth feature mask D to adjust the size of the mask, obtaining the resized fused feature D up ; Step 58: Take the fused feature D up and perform element-wise multiplication with the first preliminary fused feature C b and the second preliminary fused feature C x respectively to obtain the first masked feature C′ b and the second masked feature C′ x ; Step 59: Add the first mask feature C′ b and the second mask feature C′ x channel by channel to obtain an intermediate repaired frame I i ; Step 6: Aggregate the initially input blurred frame, the primary repaired frame A in the feature extraction stage i and the intermediate repaired frame I in the feature fusion stage i to obtain the aggregated feature Y e ; Step 7: Input the aggregated feature Y e into 5 stacked U-Net networks for feature reconstruction, and output the reconstructed high-definition frame Y i .