Video deblurring method combining intra-frame and inter-frame features

By using wavelet transform in the video defuzzing method to separate and process the high-frequency and low-frequency parts of the video frame, combined with the feature fusion of adjacent frames and the current frame, the existing methods have poor results and high computational cost when recovering severe fuzzing frames, achieving more efficient video defuzzing effect and lower computational complexity.

CN119941568APending Publication Date: 2025-05-06CHENGDU UNIV OF INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510006012.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing video defuzzing method is not effective when recovering fuzzy severe frames, and the calculation cost based on recurrent neural network and Transformer methods are high, with large parameters and long inference time, and failing to make full use of the current frame information.

Method used

Two-dimensional discrete wavelet transformation is used to separate the high-frequency part and low-frequency part of the video frame and process it separately. For the low-frequency part, the features of adjacent frames and the current frame are extracted for fusing; for the high-frequency part, the process is performed using convolution operations. Finally, the defuzzing result is obtained through the wavelet inverse transformation fusion.

Benefits of technology

Improves the video debum effect, reduces the computational complexity and network complexity, improves the running speed and reduces memory consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941568A_ABST
    Figure CN119941568A_ABST
Patent Text Reader

Abstract

The invention relates to a method for deblurring a video by using the characteristics of adjacent frames and the characteristics of a current frame, which comprises the following steps of: separating a high-frequency part and a low-frequency part of a video frame by using two-dimensional discrete wavelet transform, extracting the characteristics in the adjacent frames by using a dynamic convolution operation for the low-frequency part so as to fully capture the context information of the video; extracting the characteristics of the current frame according to the decomposability of the motion blur trajectory; carrying out feature fusion by using an intra-frame and inter-frame feature fusion strategy and reconstructing a clear frame; the high-frequency part is directly processed by using convolution operation; and finally, fusing the high-frequency part and the reconstructed low-frequency part by using inverse transformation to obtain a clear video. According to the invention, useful information is found from the blurred video frame to remove the blurring effect, and a new view angle and a new tool can be provided for research in the fields of image and video processing, photography and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing, and in particular to a video deblurring method combining intra-frame and inter-frame features. Background Art

[0002] As one of the core technologies in the current field of artificial intelligence, neural networks have been widely used in various video restoration tasks. Convolutional Neural Networks (CNN) have promoted the development of video deblurring with their powerful image feature extraction capabilities and RNN's ability to capture contextual information of long sequences. Current video deblurring methods are mainly divided into RNN-based methods and Transformer-based methods.

[0003] In 2017, Kim et al. proposed a deep convolutional residual network (STRCNN) based on recurrent neural networks, which propagates information in the time domain using a dynamic hybrid approach. In 2019, Seungjun Nah et al. proposed a method of recurrent neural networks using intra-frame loops, which uses intra-frame (internal) and inter-frame (external) loop schemes to update the hidden state multiple times within a single time step. In 2019, Zhou et al. proposed a spatiotemporal filter adaptive network (STFAN) for video deblurring, which uses element-wise filter adaptive convolution operations to extract blur features. In 2021, Zhong et al. proposed a recurrent neural network-based method (ESTRNN), which uses a residual dense block-based recurrent unit and a global spatiotemporal attention module to generate hierarchical spatial features and fuse high-level features of adjacent frames separately. In 2022, Zhu et al. proposed a deep recurrent neural network with multi-scale bidirectional propagation (RNN-MBP) to exploit inter-frame information of adjacent frames for video deblurring. Lin et al. proposed an optical flow guided Transformer network (FGST) for video deblurring in 2022, using self-attention and recurrent embedding mechanisms to remove blurring effects under the guidance of optical flow. Liang et al. proposed a Transformer-based method (VRT) in 2022, using multi-scale temporal mutual self-attention to predict the features of remote frames, and proposed a parallel warping module to fuse the features of adjacent frames.

[0004] The existing methods have the following shortcomings:

[0005] 1. Methods based on recurrent convolutional neural networks (RNNs) often require additional blurry frame alignment or frame reconstruction modules to restore clear frames, and are not very effective in restoring severely blurred frames. In addition, these RNN-based methods often have a large number of parameters and a long inference time.

[0006] 2. Most methods based on the Transfermor model require a lot of computational costs and perform poorly under low computing power conditions.

[0007] 3. These methods mainly focus on extracting useful information from adjacent frames to remove the blur effect of the current frame, but there is still a lot of information in the current frame that is not utilized. Summary of the invention

[0008] In view of the shortcomings of the prior art, the present invention proposes a video deblurring method combining intra-frame and inter-frame features, characterized in that the deblurring method uses a two-dimensional discrete wavelet transform to separate the high-frequency part and the low-frequency part of a video frame, and after processing the low-frequency part and the high-frequency part respectively, uses an inverse transform to fuse them to obtain a final deblurring result; firstly, the high-frequency part and the low-frequency part of the video frame are separated by using a wavelet transform, and for the low-frequency part, the features of adjacent frames and the features of the current frame are extracted respectively, and then they are fused to reconstruct a clear frame; for the high-frequency part, a convolution operation is used for processing, and finally the processed high-frequency part and the reconstructed low-frequency part are fused by using an inverse wavelet transform to obtain a deblurring result, specifically comprising:

[0009] Step 1: For the input blurred frame sequence N represents the length of the sequence. First, the 3D convolution operation is used to expand the number of channels of the input blurred frame, expanding each image from three RGB channels to 64 channels;

[0010] Step 2: Use two-dimensional discrete wavelet transform to transform the blurred frame sequence after the channel number expansion to separate the high-frequency part and the low-frequency part; specifically, use the first filter G 1 , the second filter G 2 , the third filter G 3 and the fourth filter G 4 The input blurred image is filtered respectively to obtain the features of four parts, namely X LL , X LH , X HL and X HH , where X LL Represents the low-frequency characteristics, X LH , X HL and X HH Represents the high-frequency characteristics;

[0011] Step 3: Extract the adjacent frame features from the low-frequency part features through the adjacent frame feature extraction module. The processing process includes:

[0012] Step 31: Low-frequency feature sequence Perform layer normalization and 3D convolution to obtain low-frequency feature sequences

[0013] Step 32: Sequence the low-frequency features The features are divided into the first feature sequence along the channel dimension The second characteristic sequence and the third characteristic sequence Three parts;

[0014] Step 33: Use continuous pooling operations and 2D convolution operations on the first feature sequence Perform downsampling operation and output the first downsampling feature map K (1) ;

[0015] Step 34: The first downsampled feature map K (1) As the convolution kernel for the first feature sequence Perform convolution operation on each frame in and output the first convolution feature sequence

[0016] Step 35: Use continuous pooling operations and 2D convolution operations to the third feature sequence Perform downsampling operation and output the second downsampling feature map K (2) ;

[0017] Step 36: Subtract the second downsampled feature map K (2) As the convolution kernel for the third feature sequence Perform convolution operation on each frame in to obtain the second convolution feature sequence

[0018] Step 37: Use element-wise multiplication to transform the first convolution feature sequence The second characteristic sequence and the second convolutional feature sequence Multiply, that is, multiply each element corresponding to each frame to obtain an enhanced feature sequence

[0019] Step 38: Use 3D convolution to enhance the feature sequence Channel dimension reduction is performed, the number of channels becomes 64, and then the LeakyReLU activation function is used for activation, and then the enhanced feature sequence is enhanced using element addition. and low-frequency feature sequences Addition, that is, adding each corresponding element in each frame to obtain the feature sequence of adjacent frames

[0020] Step 4: Use the feature extraction module of the current frame to extract the features of the current frame. The feature extraction module contains three identical submodules, namely the channel gating module, which represents the input feature sequence as the initial feature The channel gating module is used to extract the features of the current frame. The feature extraction module of the current frame is as follows: Figure 5 As shown, the processing process includes:

[0021] Step 41: Initial feature X of the input i Perform layer normalization, 3×3 convolution and 1×1 convolution to obtain convolution features The convolution feature The number of channels increases from 64 to 128;

[0022] Step 42: Convolutional features Divided into two parts along the channel dimension, the fourth feature and the fifth characteristic

[0023] Step 43: The fourth feature and the fifth characteristic 2D convolution is performed separately, and the size of the convolution kernel is 5×5 and 3×3 respectively. The convolution kernels of different sizes are used to extract features at different spatial scales. Then the Sigmoid activation function is used for activation to obtain the fourth spatial feature. and the fifth spatial feature

[0024] Step 44: Add the fourth spatial feature and the fifth spatial feature Use element-wise multiplication, then perform 2D convolution with a convolution kernel size of 3×3 to obtain the enhanced features of the current frame.

[0025] Step 45: Convolutional features Perform 3×3 convolution and then use adaptive average pooling for downsampling to obtain the downsampled feature W i ;

[0026] Step 46: Downsample the feature W i and the enhanced features of the current frame Multiply by channel multiplication to get the first fusion feature

[0027] Step 47: First fusion feature Perform a 3×3 convolution and then use the LeakyReLU activation function to activate and get the first current frame feature Y i ;

[0028] Step 48: Sequence the low-frequency features Segmentation is performed in the horizontal and vertical directions respectively to obtain the low-frequency horizontal features after segmentation. and low frequency vertical features

[0029] Step 48: Segment the low-frequency level features and low frequency vertical features Perform feature extraction according to steps 41 to 47 to obtain the current frame level features and the vertical features of the current frame

[0030] Step 49: Set the current frame horizontal features and the vertical features of the current frame Add element by element, then use two residual modules to fuse, and then add the first current frame feature Y i Add, and then use two residual modules to fuse to get the final current frame features

[0031] Step 5: Forward fusion of feature sequences, feature sequences of adjacent frames Every two frames of images and the current frame sequence A frame of image in is fused through the intra-frame and inter-frame feature fusion modules. The processing process specifically includes:

[0032] Step 51: Combine adjacent frame features Current frame features and adjacent frame features Splice from the channel dimension, then perform 3×3 convolution and use the LeakyReLu function for activation to obtain the second fusion feature At this time, the second fusion feature The number of channels has been changed to 128;

[0033] Step 52: Combine the second fusion features along the channel dimension Divided into two parts, the first channel characteristics and the second channel characteristics First channel characteristics and the second channel characteristics The number of channels is 64, and then 3×3 convolution and LeakyReLU activation functions are used for activation, and the obtained features are combined with the first channel features. Second channel characteristics Perform element multiplication and then add to get channel fusion features

[0034] Step 53: Channel fusion features Use dynamic convolution operation to perform downsampling, first use a series of operations to get the third downsampling feature The third downsampling feature The width and height are both 5, using the third downsampling feature As convolution kernel and channel fusion features Perform convolution operation;

[0035] Step 54: Input the result of the convolution in step 53 into 15 consecutive residual modules to obtain the third fusion feature

[0036] Step 6: Since only two frames can be fused at a time, forward fusion is performed in a cyclic manner to obtain the features after forward fusion.

[0037] Step 7: Forward fusion of the feature sequence in step 6 Perform backward fusion to obtain backward fusion features

[0038] Step 8: For the high-frequency feature sequence and Perform 3×3 convolution, LeakyReLu and 3×3 convolution operations in sequence to obtain the first high-frequency feature Second high frequency feature and the third high frequency feature

[0039] Step 9: Use inverse wavelet transform to transform the three high-frequency features obtained in step 8 and the backward fusion features Transform back and then change the number of channels to 3 to get the final deblurred frame.

[0040] Therefore, this study proposes to develop a new video deblurring method by combining the information of adjacent frames and the current frame. This method can not only provide a powerful means for the field of video deblurring, but also promote the development of deep learning technology. Compared with the existing methods, the beneficial effects of the present invention are:

[0041] 1. When extracting features of adjacent frames, the present invention constructs dynamic convolution kernel parameters for the input sequence to perform feature extraction. Compared with previous feature extraction methods, such as attention mechanism and deformable convolution, it improves the feature extraction strength while reducing the complexity of the network.

[0042] 2. The present invention extracts features of the current frame while extracting features of adjacent frames. Compared with various methods based on adjacent frames, it increases the representation strength and richness of the features and can restore a clearer picture in the fusion stage.

[0043] 3. The present invention uses wavelet transform to separate high-frequency features and low-frequency features of each frame, and uses different methods to process the two features respectively, thereby improving the deblurring effect and reducing the computational complexity.

[0044] 4. Previous methods are based on recurrent neural networks (RNNs) and Transformers. Recurrent neural networks run in a global loop, occupying large amounts of memory and taking a long time to reason. The calculation of the self-attention mechanism in Transformers also requires huge computing resources. The present invention is constructed using a convolutional neural network (CNN), which only uses loops (i.e., local loops) in the feature fusion stage. Compared with the previous two methods, it increases the running speed while also reducing memory consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is the overall structure diagram of the deblurring model of the present invention;

[0046] Figure 2 It is a structural diagram of a feature extraction module of adjacent frames of the present invention;

[0047] Figure 3 is a structural diagram of a channel gating module of the present invention;

[0048] Figure 4 is a structural diagram of the residual module of the present invention;

[0049] Figure 5 is a structural diagram of a feature extraction module of a current frame of the present invention;

[0050] Figure 6 It is a structural diagram of the intra-frame and inter-frame feature fusion module of the present invention;

[0051] Figure 7 is a structural diagram of a forward fusion unit of the present invention;

[0052] Figure 8 is a structural diagram of a backward fusion unit of the present invention;

[0053] Fig. 9 These are some test results on the GOPRO dataset;

[0054] Fig.10 Here are some test results on the DVD dataset;

[0055] Fig.11 These are some test results on the BSD dataset. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.

[0057] The following is a detailed description with reference to the accompanying drawings.

[0058] In view of the problems existing in the prior art, the present invention proposes a new video deblurring method by combining the information of adjacent frames and the current frame. The method can not only provide a powerful means for the field of video deblurring, but also promote the development of deep learning technology. Specifically, the present invention proposes a video deblurring method that combines intra-frame and inter-frame features. The deblurring method uses wavelet transform to separate the high-frequency part and the low-frequency part of the video frame, and after processing the low-frequency part and the high-frequency part separately, they are fused using inverse transform to obtain the final deblurring result. First, the high-frequency part and the low-frequency part of the video frame are separated by wavelet transform. For the low-frequency part, the features of the adjacent frames and the features of the current frame are extracted respectively, and then they are fused to reconstruct a clear frame; for the high-frequency part, a convolution operation is used for processing. Finally, the processed high-frequency part and the reconstructed low-frequency part are fused using inverse wavelet transform to obtain the deblurred result, such as Figure 1 shown.

[0059] Step 1: For the input blurred frame sequence N represents the length of the sequence. First, the 3D convolution operation is used to expand the number of channels of the input blurred frame. Each image is expanded from three RGB channels to 64 channels. Each channel contains different information. The purpose is to obtain more feature representations of the blurred frame.

[0060] Step 2: Use two-dimensional discrete wavelet transform to transform the blurred frame sequence after the channel number expansion to separate the high-frequency part and the low-frequency part.

[0061] The wavelet function used in the present invention is the Haar Wavelet Function, which is used to filter each frame of the image according to the following formula:

[0062]

[0063] Among them, X(i,j) represents the pixel value of the i-th row and j-th column after filtering, G(m,n) represents the value of the m-th row and n-th column of the filter, and I represents the input image. There are four filters in total, and each filter is a two-dimensional matrix, as follows:

[0064]

[0065] Using the first filter G 1 , the second filter G 2 , the third filter G 3 and the fourth filter G 4 The input blurred image is filtered respectively to obtain the features of four parts, namely X LL , X LH , X HLand X HH Among them, X LL Represents the low-frequency characteristics, X LH , X HL and X HH Indicates the high-frequency characteristics. LL , X LH , X HL and X HH Still has 64 channels, but the width and height are halved;

[0066] Step 3: Extract the adjacent frame features through the feature extraction module of the adjacent frames for the low-frequency part features, such as Figure 2 As shown, the feature extraction module of adjacent frames mainly extracts spatiotemporal information, including:

[0067] Step 31: Low-frequency feature sequence Perform layer normalization and 3D convolution to obtain low-frequency feature sequences

[0068] The purpose of layer normalization is to adjust the distribution of data and reduce the range of value differences, thereby improving the generalization of the model. The purpose of 3D convolution is to increase the number of feature channels and increase the feature representation. The number of channels increases from 64 to 768.

[0069] Step 32: Sequence the low-frequency features The features are divided into the first feature sequence along the channel dimension The second characteristic sequence and the third characteristic sequence There are three parts, and the number of channels in each part is 256;

[0070] Step 33: Use continuous pooling operations and 2D convolution operations on the first feature sequence Perform downsampling operation and output the first downsampling feature map K (1) , the width and height both become 5;

[0071] The specific operation order is: average pooling -> 3×3 convolution -> 3×3 convolution -> maximum pooling -> 3×3 convolution -> 3×3 convolution -> adaptive average pooling -> 1×1 convolution. The purpose of this step is to bring the feature information together.

[0072] Step 34: The first downsampled feature map K (1) As the convolution kernel for the first feature sequence Perform convolution operation on each frame in and output the first convolution feature sequence Mathematical expression:

[0073]

[0074] in, Represents the result after convolution of each frame.

[0075] Step 35: Use continuous pooling operations and 2D convolution operations to the third feature sequence Perform downsampling operation and output the second downsampling feature map K (2) , the width and height variables are both changed to 3;

[0076] The order of operations is: average pooling -> 3×3 convolution -> 3×3 convolution -> maximum pooling -> 3×3 convolution -> 3×3 convolution -> adaptive average pooling -> 1×1 convolution.

[0077] Step 36: Subtract the second downsampled feature map K (2) As the convolution kernel for the third feature sequence Perform convolution operation on each frame in to obtain the second convolution feature sequence as follows:

[0078]

[0079] in, Represents the result after convolution of each frame.

[0080] Step 37: Use element-wise multiplication to transform the first convolution feature sequence The second characteristic sequence and the second convolutional feature sequence Multiply, that is, multiply each element corresponding to each frame to obtain an enhanced feature sequence

[0081] The first down-sampled feature map K of two convolution kernels (1) and the second downsampled feature map K (2) The sizes of are 5×5 and 3×3 respectively. Convolution kernels of different sizes have different receptive fields and can extract features at different spatial scales.

[0082] Step 38: Use 3D convolution to enhance the feature sequence Channel dimension reduction is performed, the number of channels becomes 64, and then the LeakyReLU activation function is used for activation, and then the enhanced feature sequence is enhanced using element addition. and low-frequency feature sequences Addition, that is, adding each corresponding element in each frame to obtain the feature sequence of adjacent frames Skip connections are used to prevent feature degradation.

[0083] Step 4: Use the feature extraction module of the current frame to extract the features of the current frame. The feature extraction module contains three identical submodules, namely, the channel gating module, with the following structure: Figure 3 As shown. The input feature sequence is represented as the initial feature The channel gating module is used to extract the features of the current frame. The feature extraction module of the current frame is as follows: Figure 5 As shown, the specific steps include:

[0084] Step 41: Initial feature X of the input i Perform layer normalization, 3×3 convolution and 1×1 convolution to obtain convolution features The convolution feature The number of channels increases from 64 to 128;

[0085] Step 42: Convolutional features Divided into two parts along the channel dimension, the fourth feature and the fifth characteristic The fourth characteristic and the fifth characteristic The number of channels is 64;

[0086] Step 43: The fourth feature and the fifth characteristic 2D convolution is performed separately, and the size of the convolution kernel is 5×5 and 3×3 respectively. The convolution kernels of different sizes are used to extract features at different spatial scales. Then the Sigmoid activation function is used for activation to obtain the fourth spatial feature. and the fifth spatial feature

[0087] Step 44: Add the fourth spatial feature and the fifth spatial feature Use element-wise multiplication, then perform 2D convolution with a convolution kernel size of 3×3 to obtain the enhanced features of the current frame.

[0088] Step 45: Convolutional features Perform 3×3 convolution and then use adaptive average pooling for downsampling to obtain the downsampled feature W i ; Downsample feature W i is a vector with 64 channels and 1 width and height;

[0089] Step 46: Downsample the feature W i and the enhanced features of the current frame Multiply by channel multiplication to get the first fusion feature The expression for channel multiplication is as follows:

[0090]

[0091] Among them, k, i, and j represent the number of channels, width, and height, respectively.

[0092] Step 47: First fusion feature Perform a 3×3 convolution and then use the LeakyReLU activation function to activate and get the first current frame feature Y i .

[0093] Step 48: Characterize the low frequency part Segmentation is performed in the horizontal and vertical directions respectively to obtain the low-frequency horizontal features after segmentation. and low frequency vertical features

[0094] Because the fuzzy trajectory in any direction can be regarded as a combination of horizontal and vertical components, splitting from the horizontal and vertical directions allows the model to focus more on feature extraction in a single direction;

[0095] Step 48: Segment the low-frequency level features and low frequency vertical features Perform feature extraction according to steps 41 to 47 to obtain the current frame level features and the vertical features of the current frame

[0096] Step 49: Set the current frame horizontal features and the vertical features of the current frame Add element by element, and then use two residual modules for fusion. The residual module structure is as follows Figure 4 As shown, the first and the next current frame features Y i Add, and then use two residual modules to fuse to get the final current frame features

[0097] Step 5: Forward fusion of feature sequences, feature sequences of adjacent frames Every two frames of images and the current frame sequence A frame of image in is fused through intra-frame and inter-frame feature fusion modules, including:

[0098] Step 51: Combine adjacent frame features Current frame features and adjacent frame features Splice from the channel dimension, then perform 3×3 convolution and use the LeakyReLu function for activation to obtain the second fusion feature At this time, the second fusion feature The number of channels is changed to 128; the intra-frame and inter-frame feature fusion modules are as follows Figure 6 shown.

[0099] Step 52: Combine the second fusion features along the channel dimension Divided into two parts, the first channel characteristics and the second channel characteristics First channel characteristics and the second channel characteristics The number of channels is 64, and then 3×3 convolution and LeakyReLU activation functions are used for activation, and the obtained features are combined with the first channel features. Second channel characteristics Perform element multiplication and then add to get channel fusion features

[0100] Step 53: Channel fusion features Use dynamic convolution operation to perform downsampling, first use a series of operations to get the third downsampling feature The third downsampling feature The width and height are both 5, using the third downsampling feature As convolution kernel and channel fusion features Perform a convolution operation, where p represents the convolution kernel used for fusion.

[0101] The order of operations is: average pooling->residual module->residual module->residual module->max pooling->residual module->residual module->residual module->adaptive average pooling->1×1 convolution.

[0102] Step 54: Input the result of the convolution in step 53 into 15 consecutive residual modules to obtain the third fusion feature

[0103] Step 6: Since only two frames can be fused at a time, forward fusion is performed in a cyclic manner to obtain the features after forward fusion. as follows:

[0104]

[0105] Among them, F represents the intra-frame and inter-frame feature fusion module, represents the i+1th frame after forward fusion, and represents the adjacent frame features of the i+1th frame and the ith frame, Represents the current frame features of the i+1th frame, and the fusion method is as follows Figure 7 shown.

[0106] Step 7: Forward fusion feature sequence Perform backward fusion to obtain backward fusion features The mathematical expression is as follows:

[0107]

[0108] in, represents the i-1th frame after backward fusion, and represents the i-1th frame and the i-th frame after forward fusion, Represents the current frame features of the i-1th frame, and the fusion method is as follows Figure 8 shown.

[0109] Step 8: For the high-frequency feature sequence and Perform 3×3 convolution, LeakyReLu and 3×3 convolution operations in sequence to obtain the first high-frequency feature Second high frequency feature and the third high frequency feature

[0110] Step 9: Use inverse wavelet transform to transform the three high-frequency features obtained in step 8 and the backward fusion features Transform back and then change the number of channels to 3 to get the final deblurred frame. The mathematical expression is as follows:

[0111]

[0112] Among them, D i represents the final deblurred frame, IHT represents inverse wavelet transform, B i Indicates a blurred frame.

[0113] The experimental environment of the present invention is specifically as follows: the graphics card is an NVIDIA RTX 3090 GPU, and the CPU is an Intel(R) Core(TM) i7-13700K.

[0114] The present invention uses the commonly used GOPRO, DVD and BSD datasets for video deblurring to conduct experiments and compare with existing methods. The GOPRO and DVD datasets are artificially synthesized blur datasets, while the BSD is a real blur dataset taken.

[0115] The GOPRO dataset contains 33 pairs of blurry videos and clear videos, a total of 3214 pairs of images, and the resolution of each image is 1280 × 720. 22 videos are used for training and 11 videos are used for testing.

[0116] The DVD dataset contains 71 videos, consisting of 6708 pairs of images, each with a resolution of 1280 × 720. 61 videos are used for training and 10 videos are used for testing.

[0117] The BSD dataset is a real-world blur dataset collected using a dual-camera system, containing 80 videos, 9,000 blurry and clear image pairs, and the resolution of each image is 640x480. 60 videos are used for training and 20 videos are used for testing.

[0118] The present invention uses two indicators, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM), to evaluate the deblurring results. PSNR evaluates the quality of deblurring by measuring the error between the deblurred image and the clear image. The larger the value, the better. SSIM measures the similarity of images from three parts: brightness, contrast and structure between two images. The closer the value is to 1, the higher the similarity of the two images.

[0119] In the training phase, the present invention uses supervised learning to train the model. Each blurry-clear image pair is cropped into a size of 240×240 and input into the model. The batch size is set to 1, and the video length of each batch is 18. The AdamW optimizer is used for optimization, and the initial learning rate is set to 1×10 -4 , the cosine annealing strategy is used to adjust the learning rate, and finally the learning rate is reduced to 1×10 -7 , and trained for 600,000 iterations.

[0120] The loss functions used in training are Charbonnier Loss and L1 Loss. For Charbonnier Loss, the mathematical expression is as follows:

[0121]

[0122] Where p is the pixel value output by the model, t is the pixel value of the corresponding clear image, and ∈ is a small constant, which is set to 1×10 -12 . For L1 Loss, the calculation formula is as follows:

[0123]

[0124] When calculating the final loss, the weights of the two losses are 1 and 0.1, as follows:

[0125] L=L Charbonnier ×1+L L1 ×0.1 (11)

[0126] During the testing phase, the input video length was set to 36, and the test sets of the three datasets were used for testing.

[0127] The comparison methods include the following methods. Method 1: Spatiotemporal filter adaptation network (STFAN); Method 2: Multi-scale bidirectional propagation deep recurrent neural network (RNN-MBP); Method 3: Optical flow guided transformer network (FGST). Table 1 is the comparison of PSNR and SSIM on the GOPRO dataset; Table 2 is the comparison of PSNR and SSIM on the DVD dataset; Table 3 is the comparison of PSNR and SSIM on the BSC dataset.

[0128] Table 1 Comparison of indicators on the GOPRO dataset

[0129]

[0130] Table 2 Comparison of indicators on the DVD dataset

[0131]

[0132] Table 3 Comparison of indicators on the BSD dataset

[0133]

[0134] It can be seen intuitively from Tables 1 to 3 that the objective evaluation indicators of the method of the present invention on the three data sets are slightly higher than those of the existing methods, indicating that the method of the present invention is a technical advancement.

[0135] Fig. 9 , Fig.10 and Fig.11 Partial visual effects of the method of the present invention on three data sets are shown respectively. As shown in the figure, the deblurred frame processed by the present invention is very close to the original clear frame, with better detail restoration, and clear numbers that can be distinguished are restored in the smaller blurred digital parts.

[0136] It should be noted that the above specific embodiments are exemplary, and those skilled in the art can come up with various solutions inspired by the disclosure of the present invention, and these solutions also belong to the disclosure scope of the present invention and fall within the protection scope of the present invention. Those skilled in the art should understand that the present invention description and its drawings are illustrative and do not constitute a limitation of the claims. The protection scope of the present invention is defined by the claims and their equivalents.

Claims

1. A video deblurring method combining intra-frame and inter-frame features, characterized in that: The deblurring method uses a two-dimensional discrete wavelet transform to separate the high-frequency part and the low-frequency part of the video frame, and after processing the low-frequency part and the high-frequency part respectively, uses an inverse transform to fuse them to obtain the final deblurring result; firstly, the high-frequency part and the low-frequency part of the video frame are separated by using a wavelet transform, and for the low-frequency part, the features of the adjacent frames and the features of the current frame are extracted respectively, and then the features are fused to reconstruct a clear frame; For the high-frequency part, the convolution operation is used for processing, and finally the processed high-frequency part and the reconstructed low-frequency part are fused using the inverse wavelet transform to obtain the deblurred result, which includes: Step 1: For the input blurred frame sequence N represents the length of the sequence. First, the 3D convolution operation is used to expand the number of channels of the input blurred frame, expanding each image from three RGB channels to 64 channels; Step 2: Use two-dimensional discrete wavelet transform to transform the blurred frame sequence after the channel number expansion to separate the high-frequency part and the low-frequency part; specifically, use the first filter G1, the second filter G2, the third filter G3 and the fourth filter G4 to filter the input blurred image respectively to obtain the features of the four parts, which are X LL , X LH , X HL and X HH , where X LL Represents the low-frequency characteristics, X LH , X HL and X HH Represents the high-frequency characteristics; Step 3: Extract the adjacent frame features from the low-frequency part features through the adjacent frame feature extraction module. The processing process includes: Step 31: Low-frequency feature sequence Perform layer normalization and 3D convolution to obtain low-frequency feature sequences Step 32: Sequence the low-frequency features The features are divided into the first feature sequence along the channel dimension The second characteristic sequence and the third characteristic sequence Three parts; Step 33: Use continuous pooling operations and 2D convolution operations on the first feature sequence Perform downsampling operation and output the first downsampling feature map K (1) ; Step 34: The first downsampled feature map K (1) As the convolution kernel for the first feature sequence Perform convolution operation on each frame in and output the first convolution feature sequence Step 35: Use continuous pooling operations and 2D convolution operations to the third feature sequence Perform downsampling operation and output the second downsampling feature map K (2) ; Step 36: The second downsampled feature map K (2) As the convolution kernel for the third feature sequence Perform convolution operation on each frame in to obtain the second convolution feature sequence Step 37: Use element-wise multiplication to transform the first convolution feature sequence The second characteristic sequence and the second convolutional feature sequence Multiply, that is, multiply each element corresponding to each frame to obtain an enhanced feature sequence Step 38: Use 3D convolution to enhance the feature sequence Channel dimension reduction is performed, the number of channels becomes 64, and then the LeakyReLU activation function is used for activation, and then the enhanced feature sequence is enhanced using element addition. and low-frequency feature sequences Addition, that is, adding each corresponding element in each frame to obtain the feature sequence of adjacent frames Step 4: Use the feature extraction module of the current frame to extract the features of the current frame. The feature extraction module contains three identical submodules, namely the channel gating module, which represents the input feature sequence as the initial feature The channel gating module is used to extract the features of the current frame. The feature extraction module of the current frame is shown in Figure 5. The processing process includes: Step 41: Initial feature X of the input i Perform layer normalization, 3×3 convolution and 1×1 convolution to obtain convolution features The convolution feature The number of channels increases from 64 to 128; Step 42: Convolutional features Divided into two parts along the channel dimension, the fourth feature and the fifth characteristic Step 43: The fourth feature and the fifth characteristic 2D convolution is performed separately, and the size of the convolution kernel is 5×5 and 3×3 respectively. The convolution kernels of different sizes are used to extract features at different spatial scales. Then the Sigmoid activation function is used for activation to obtain the fourth spatial feature. and the fifth spatial feature Step 44: Add the fourth spatial feature and the fifth spatial feature Use element-wise multiplication, then perform 2D convolution with a convolution kernel size of 3×3 to obtain the enhanced features of the current frame Step 45: Convolutional features Perform 3×3 convolution and then use adaptive average pooling for downsampling to obtain the downsampled feature W i ; Step 46: Downsample the feature W i and the enhanced features of the current frame Multiply by channel multiplication to get the first fusion feature Step 47: First fusion feature Perform a 3×3 convolution and then use the LeakyReLU activation function to activate and get the first current frame feature Y i ; Step 48: Sequence the low-frequency features Segmentation is performed in the horizontal and vertical directions respectively to obtain the low-frequency horizontal features after segmentation. and low frequency vertical features Step 48: Segment the low-frequency level features and low frequency vertical features Perform feature extraction according to steps 41 to 47 to obtain the current frame level features and the vertical features of the current frame Step 49: Set the current frame horizontal features and the vertical features of the current frame Add element by element, then use two residual modules to fuse, and then add the first current frame feature Y i Add, and then use two residual modules to fuse to get the final current frame features Step 5: Forward fusion of feature sequences, feature sequences of adjacent frames Every two frames of images and the current frame sequence A frame of image in is fused through the intra-frame and inter-frame feature fusion modules. The processing process specifically includes: Step 51: Combine adjacent frame features Current frame features and adjacent frame features Splice from the channel dimension, then perform 3×3 convolution and use the LeakyReLu function for activation to obtain the second fusion feature At this time, the second fusion feature The number of channels has been changed to 128; Step 52: Combine the second fusion features along the channel dimension Divided into two parts, the first channel characteristics and the second channel characteristics First channel characteristics and the second channel characteristics The number of channels is 64, and then 3×3 convolution and LeakyReLU activation functions are used for activation, and the obtained features are combined with the first channel features. Second channel characteristics Perform element multiplication and then add to get channel fusion features Step 53: Channel fusion features Use dynamic convolution operation to perform downsampling, first use a series of operations to get the third downsampling feature The third downsampling feature The width and height are both 5, using the third downsampling feature As convolution kernel and channel fusion features Perform convolution operation; Step 54: Input the result of the convolution in step 53 into 15 consecutive residual modules to obtain the third fusion feature Step 6: Since only two frames can be fused at a time, forward fusion is performed in a cyclic manner to obtain the features after forward fusion. Step 7: Forward fusion of the feature sequence in step 6 Perform backward fusion to obtain backward fusion features Step 8: For the high-frequency feature sequence and Perform 3×3 convolution, LeakyReLu and 3×3 convolution operations in sequence to obtain the first high-frequency feature Second high frequency feature and the third high frequency feature Step 9: Use inverse wavelet transform to transform the three high-frequency features obtained in step 8 and the backward fusion features Transform back and then change the number of channels to 3 to get the final deblurred frame.