Video deblurring method based on deformable space-time sparse converter
Through a video deblurring method based on a deformable spatiotemporal sparse transformer, the problems of high computational complexity and low efficiency in the existing technology are solved by utilizing multi-scale blur map fusion and sparse attention mechanism, and an efficient video deblurring effect is achieved.
Patent Information
- Application Number
- CN202510919763.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-10
AI Technical Summary
Existing video deblurring technology has high computational complexity and low computational efficiency when processing high-resolution videos, and it is difficult to effectively utilize the spatiotemporal correlation information in video sequences, especially the insufficient feature matching capability in non-rigid motion and occluded areas.
A method based on deformable spatiotemporal sparse transformer is adopted. Through the collaborative design of bidirectional feature propagation module and deformable spatiotemporal sparse transformer module, multi-scale fuzzy graph fusion and sparse attention mechanism are utilized to achieve efficient deblurring.
It effectively reduces the computational complexity, improves the modeling accuracy and deblurring performance in complex motion scenes, and improves the accuracy and efficiency of video frame reconstruction.
Smart Images

Figure CN120765504A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video deblurring in computer vision, and in particular to a video deblurring method based on a deformable spatiotemporal sparse transformer. Background Art
[0002] With the rapid development of portable camera devices like action cameras and smartphones, handheld filming is becoming increasingly common. However, during handheld filming, rapid relative motion between the device and the subject, or device jitter, often results in motion blur in video frames. This blur not only degrades visual quality but also significantly impacts the accuracy of subsequent advanced vision tasks like object tracking and detection.
[0003] Video deblurring technology aims to recover potential high-quality clear frames from blurred video sequences. Since the blur kernel is unknown and the motion trajectory is complex in real scenes, this problem is mathematically a highly ill-posed problem. The core challenge lies in how to effectively utilize the spatiotemporal correlation information in the video sequence to reconstruct the potential clear content of the current blurred frame. Early deep learning-based methods usually adopted a multi-frame feature stacking strategy and directly learned the blur-to-clear mapping relationship through a convolutional neural network (CNN). However, such methods rely on pre-calculated optical flow or deformable convolution for feature alignment, and have insufficient feature matching capabilities for non-rigid motion and occluded areas. In addition, the local receptive field characteristics of CNN limit its ability to model long-range spatiotemporal dependencies.
[0004] To overcome the local limitations of CNNs, researchers have attempted to introduce Transformer architectures based on global self-attention mechanisms. While these approaches can capture non-local similarity features through spatiotemporal attention, their computational complexity grows quadratically with the spatiotemporal dimensions of the video, resulting in significant memory overhead and slow inference speed when processing high-resolution video. Furthermore, the global attention mechanism's intensive interaction across all spatial locations generates a significant amount of redundant computation, reducing the efficiency of modeling critical motion regions and hindering convergence in model training. Therefore, the field of video deblurring urgently needs a new approach that balances the ability to model long-range spatiotemporal dependencies with computational efficiency. Summary of the Invention
[0005] The present invention aims to address the shortcomings of existing technologies by providing a video deblurring method based on a deformable spatiotemporal sparse transformer. This method, through the collaborative design of a bidirectional feature propagation module and a deformable spatiotemporal sparse transformer module, effectively utilizes long-range video frame information to achieve efficient deblurring while reducing computational complexity.
[0006] In order to solve the above technical problems and improve the deblurring performance on the video deblurring dataset, the present invention adopts the following technical solution: a video deblurring method based on a deformable spatiotemporal sparse transformer, comprising the following steps:
[0007] Step S1: For the input n video frames to be deblurred, use the encoder module to extract features And use the optical flow estimation network to obtain the forward optical flow between adjacent frames With backward optical flow Then, the blur map of each frame is estimated based on the forward and backward optical flows;
[0008] Step S2, multi-level downsampling of the fuzzy image and fusion to generate a multi-scale fuzzy image, and the extracted features The forward optical flow, backward optical flow, and multi-scale fuzzy map are input into the bidirectional feature propagation module, and the sequence frame features are recursively updated through two forward propagation branches and two backward propagation branches. The features output by the four propagation branches of the bidirectional feature propagation module are then combined with the features extracted by the encoder. Perform splicing and input into the multi-source spatiotemporal aggregation module to generate aggregation features
[0009] Step S3: Use the deformable spatiotemporal sparse Transformer module to aggregate features Processing to obtain refined features
[0010] Step S4, refine the features The input is processed by the decoder module and connected with the original input residual to obtain the final deblurred result.
[0011] Furthermore, in step S1, the blurred video of consecutive n frames is used as the input sequence unit, and the encoder module is used to extract features frame by frame to obtain the features of each frame. The expression is as follows:
[0012] f t =Encoder(x t ),t∈[0,n-1]
[0013] Among them, f t Represents the features of the t-th frame, Encoder represents the encoder, x t represents the t-th frame input;
[0014] Use the pre-trained RAFT network as the optical flow estimation network to obtain the forward optical flow between adjacent frames With backward optical flow And calculate the fuzzy map corresponding to each frame of the input sequence unit The expression is as follows:
[0015]
[0016] For fuzzy graph Normalize and get the normalized fuzzy image By fuzzy graph Calculate the clear image corresponding to each frame of the input sequence unit The expression is as follows:
[0017]
[0018] S t =1-M t ,t∈[0,n-1]
[0019] in, and Respectively represent the minimum and maximum values of the fuzzy graph in the input sequence unit.
[0020] Furthermore, in step S2, a multi-scale fuzzy image fusion module is used to fuzzy image Perform two-level maximum pooling downsampling to generate the first and second level blur maps step by step, and fuse them with the original blur map across scales; finally, a multi-level integrated multi-scale blur map is obtained. The corresponding multi-scale clear image is
[0021] In the bidirectional feature propagation module, two backward propagation branches and two backward propagation branches are constructed to propagate features respectively; each propagation branch contains a deformable alignment module, which aligns the features of adjacent frames through the multi-scale blur map and the optical flow-guided deformable alignment module to reconstruct the current frame; the specific steps are as follows:
[0022] For forward propagation, the video sequence is processed frame by frame from the starting frame to the ending frame; at the t-th time step of the j-th propagation branch, the optical flow O from the current frame to the previous two frames is used. t→t-1 , O t→t-2 , for the features of the first two time steps Perform space warping to generate preliminary alignment features The expression is as follows:
[0023]
[0024] Among them, W represents the space-warping operation;
[0025] The specific steps of the multi-scale blur map and optical flow guided deformable alignment module are as follows:
[0026] Initially align features The optical flow O from the current frame to the previous two frames t→t-1, O t→t-2 , the clear picture of the first two frames and the features of the previous branch at the tth time step Splicing is performed; after processing through a convolutional layer consisting of four 2D convolutions, the residual offset and residual mask are obtained;
[0027] Using optical flow O t→t-1 , O t→t-2 Guide the initial offset generation of deformable convolution: take the pre-estimated optical flow fields of the first two frames as the base offset, predict the residual offset through the convolution layer, and then superimpose the base offset and the residual offset to obtain the initial offset;
[0028] Use clear pictures Guided deformable convolution initial mask generation: Introducing a clarity map-driven adaptive mask modulation mechanism, linearly mapping spatial clarity information to the mask space to generate a base mask. Using a convolutional network to predict the residual mask, the base mask and the residual mask are superimposed, and the offset weights of different regions are dynamically adjusted using a sigmoid function to obtain the initial mask.
[0029] The features of the first two frames are aligned through deformable convolution with parameters of initial offset and initial mask; thus, the features of the t-th time step of the previous branch are aligned. Splicing; processed by a convolution layer consisting of 4 2D convolutions to obtain the current time step aggregation feature
[0030] Backward propagation is performed backward along the time step, and its process is the same as the forward propagation process; the features output by the four propagation branches of the bidirectional feature propagation module are compared with the features extracted by the encoder. After splicing, input the multi-source spatiotemporal aggregation module to generate aggregate features
[0031] Furthermore, the multi-source spatiotemporal aggregation module includes a 3D convolution module and two Swin Transformer modules which are connected in sequence.
[0032] Furthermore, in step S3, two cascaded deformable spatiotemporal sparse Transformer modules are used to aggregate features generated by the multi-source spatiotemporal aggregation module. deal with.
[0033] Furthermore, the aggregation features Processing, including:
[0034] The soft segmentation module is used to aggregate features Expand the sliding window to obtain the block sequence feature F block ;
[0035] Block sequence feature F block After layer normalization, the sparse attention module is input to calculate the sparse attention A s ; Block sequence feature F block With sparse attention A s Residual connection obtains intermediate feature F mid , and then the intermediate feature F mid After normalization, the nonlinear feature F is obtained by using the feedforward network processing. ffn , thus the nonlinear feature F ffn With the intermediate feature F mid Perform residual connection to obtain refined block sequence features
[0036] For refined block sequence features The soft reorganization module is used to restore it to the refined features of the dimension size before the soft segmentation operation
[0037] Furthermore, sparse attention A s The calculation steps are as follows:
[0038] For the fuzzy graph M t Perform two maximum pooling downsamplings to generate a downsampled blur map M that adapts to the input size of the deformable spatiotemporal sparse Transformer module t ′ , and calculate the corresponding downsampled clear image S t ′ ; By setting a hard threshold, a binary spatial sparse mask is generated to distinguish high fuzzy areas; t ′ In descending order, select the first w% most blurred blur maps and multiply them element-wise with the spatial sparse mask to generate the query mask M q ; For clear image S t ′ In descending order, select the top w% of the clearest clear maps and multiply them element-wise with the spatial sparse mask to generate the key-value mask M kv ;
[0039] The normalized block sequence feature F block Perform linear projection to generate query vector, key vector, and value vector respectively, and divide them into local windows according to the window size to obtain local query vector Q and local key vector K L , local value vector V L ;
[0040] The deformable key vector K is predicted by the deformable sampling module D , deformable value vector V D ;
[0041] The normalized block sequence feature F block After the pooling layer, the pooled sequence feature F block Perform linear projection respectively to generate the global query vector K P , global key vector V P ; K L , K D , K P Splice to get the expanded key vector K, and then convert V L 、V D 、V P Perform concatenation to obtain the expanded value vector V;
[0042] Combine the local query vector Q with the query mask M q and key-value mask M kv , distinguish between masked windows and unmasked windows; use the expanded key vector K and the expanded value vector V to calculate the global-local attention of the masked window through the multi-head attention module; for the unmasked window, use the local key vector K L , local value vector V L The local attention is calculated through the multi-head attention module; the attention outputs of the two types of windows are merged to obtain the sparse attention A s .
[0043] The present invention also provides an electronic device, comprising a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned video deblurring method based on a deformable spatiotemporal sparse transformer.
[0044] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the video deblurring method based on a deformable spatiotemporal sparse transformer is implemented.
[0045] The present invention also provides a computer program product, comprising a computer program, which implements the above-mentioned video deblurring method based on a deformable spatiotemporal sparse transformer when executed by a processor.
[0046] Compared with the existing methods, the present invention has the following beneficial effects:
[0047] 1. The present invention adopts a multi-scale fuzzy graph fusion module to combine fuzzy information of different scales to guide the feature alignment of adjacent video frames during the bidirectional feature propagation process, suppress the error accumulation generated during the feature propagation process and improve the reconstruction accuracy of the current frame.
[0048] 2. The present invention adopts a deformable spatiotemporal sparse Transformer module to dynamically screen high-fuzzy area tokens (query space) and high-definition area tokens (key-value space) through fuzzy perception masks, and combines the spatiotemporal sparse attention mechanism of deformable sampling to suppress redundant token interactions and reduce computational complexity while capturing non-rigid motion features, thereby improving the modeling accuracy in complex motion scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0050] Figure 1 is a flow chart of the deblurring method of the present invention;
[0051] Figure 2 is a system framework diagram of the deblurring method in the present invention;
[0052] Figure 3 Schematic diagram of the multi-scale fuzzy image fusion module in the present invention;
[0053] Figure 4 Schematic diagram of the multi-scale blur map and optical flow guided deformable alignment module in the present invention;
[0054] Figure 5 Schematic diagram of the deformable spatiotemporal sparse Transformer module in the present invention;
[0055] Figure 6 Schematic diagram of the deblurring effect of the deblurring method of the present invention on some blurred video frames in the test set;
[0056] Figure 7 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0057] The present invention will be described in detail below with reference to the accompanying drawings. Unless there is any conflict, the features of the following embodiments and implementations may be combined with each other.
[0058] like Figure 2 As shown in the figure, the video deblurring method based on deformable spatiotemporal sparse transformer of the present invention is composed of a bidirectional feature propagation module, a multi-scale fuzzy map fusion module and a deformable spatiotemporal sparse transformer module. Figure 1 The specific steps are as follows:
[0059] Step 1, feature extraction and optical flow and blur map calculation: Take the blurred video of consecutive n frames as the input sequence unit. Use the encoder module to extract features frame by frame and obtain the features of each frame. The expression is as follows:
[0060] f t =Encoder(x t ),t∈[0,n-1]
[0061] Among them, f t Represents the features of the t-th frame, Encoder represents the encoder, x t Represents the t-th frame input.
[0062] Use the pre-trained RAFT network as the optical flow estimation network to obtain the forward optical flow between adjacent frames With backward optical flow And calculate the fuzzy map corresponding to each frame of the input sequence unit The expression is as follows:
[0063]
[0064] For fuzzy graph Normalize and get the normalized fuzzy image By fuzzy graph Further calculations are performed to obtain a clear image corresponding to each frame of the input sequence unit. The expression is as follows:
[0065]
[0066] S t =1-M t ,t∈[0,n-1]
[0067] in, and Respectively represent the minimum and maximum values of the fuzzy graph in the input sequence unit.
[0068] Step 2: Bidirectional feature propagation and multi-source feature fusion: The extracted video frame features The bidirectional feature propagation module is input, and the sequence frame features are recursively updated through two forward propagation branches and two backward propagation branches. Multi-source spatiotemporal context is aggregated in multiple rounds of iterative optimization to improve feature expression capabilities. The specific process is as follows:
[0069] like Figure 3 As shown, through the multi-scale fuzzy map fusion module, the fuzzy map Perform two-level maximum pooling downsampling to generate the first and second level blur maps step by step, and fuse them with the original blur map across scales. Finally, a multi-level integrated multi-scale blur map is obtained. The corresponding multi-scale clear image is The expression is as follows:
[0070]
[0071] The fused blur map captures both large-scale motion blur and localized detail blur, enhancing the model's ability to perceive blur patterns. Max-pooling downsampling effectively suppresses isolated noise points generated by optical flow estimation in texture-deficient areas and motion boundaries, while preserving the continuity of areas with significant motion.
[0072] In the bidirectional feature propagation module, the extracted video frame features are After the backward feature propagation branch 1 is processed, the backward branch 1 video frame feature is obtained The backward branch 1 video frame features After the forward feature propagation branch 1 is processed, the forward branch 1 video frame features are obtained The forward branch 1 video frame features After the backward feature propagation branch 2 is processed, the backward branch 2 video frame features are obtained The backward branch 2 video frame features After the forward feature propagation branch 2 is processed, the forward branch 2 video frame features are obtained Each propagation branch aligns features of adjacent frames through a multi-scale blur map and a deformable alignment module guided by optical flow. The specific steps are as follows:
[0073] For forward propagation, at the t-th time step of the j-th propagation branch, the optical flow O from the current frame to the previous two frames is used. t→t-1 , O t→t-2 , for the features of the first two time steps Perform space warping to generate preliminary alignment features The expression is as follows:
[0074]
[0075] Among them, W represents the space-based warping operation.
[0076] The structure of the multi-scale blur map and optical flow guided deformable alignment module is as follows: Figure 4 The specific steps are as follows:
[0077] Initially align features The optical flow O from the current frame to the previous two frames t→t-1 , O t→t-2 , the clear picture of the first two frames and the features of the previous branch at the tth time step After concatenation, the residual offset and residual mask are obtained through a convolutional layer consisting of four 2D convolutions.
[0078] Using optical flow O t→t-1 , O t→t-2 Guide the generation of the initial offset for deformable convolutions. Using the estimated optical flow fields of the first two frames as base offsets, an initial motion field is constructed to constrain the offset search space for deformable convolutions, reducing the difficulty of model learning. The residual offset is predicted through a convolutional network, and its range is rigidly constrained using a maximum residual amplitude threshold (10 pixels) to avoid training instability caused by offset overflow. The base offset is then superimposed on the residual offset to generate the initial offset, enhancing adaptability to complex motion patterns such as occlusion and non-rigid deformation.
[0079] Use clear pictures This method guides the generation of the initial mask for deformable convolution. A clarity map-driven adaptive mask modulation mechanism is introduced to linearly map spatial clarity information to the mask space to generate a base mask. A convolutional network is used to predict the residual mask. The base mask and the residual mask are then superimposed, and the offset weights for different regions are dynamically adjusted using a sigmoid function to obtain the initial mask. This method suppresses artifacts in highly blurred areas and enhances the feature contribution of clear areas, achieving spatially adaptive feature fusion and reducing the error accumulation caused by bidirectional feature propagation.
[0080] The features of the first two frames are aligned through deformable convolution with parameters of initial offset and initial mask. After processing through four 2D convolution layers, the aggregated features of the current time step are obtained.
[0081] Aggregate features at the current time step The calculation process expression is as follows:
[0082]
[0083] Among them, BFDA stands for Multi-scale Blur Map and Optical Flow Guided Deformable Alignment Module.
[0084] Backward propagation is performed backward along the time step, and its process is similar to the forward propagation process. The specific steps are as follows:
[0085] For backward propagation, at the t-th time step of the j-th propagation branch, the optical flow O from the current frame to the next two frames is used. t→t+1 , O t→t+2 , for the features of the two time steps before and after Perform space warping to generate preliminary alignment features The expression is as follows:
[0086]
[0087] Among them, W represents the space-based warping operation.
[0088] Then, using the preliminary alignment features The optical flow O from the current frame to the next two frames t→t+1 , O t→t+2 , the clear picture of the last two frames and the features of the previous branch at the tth time step The current time step aggregation feature is obtained through the multi-scale fuzzy map and the deformable alignment module guided by optical flow (the processing steps are exactly the same as those described in the forward propagation). The calculation process expression is as follows:
[0089]
[0090] Among them, BFDA stands for Multi-scale Blur Map and Optical Flow Guided Deformable Alignment Module.
[0091] The features output by the four propagation branches of the bidirectional feature propagation module With the original space characteristics Splicing is performed and the multi-source spatiotemporal aggregation module consisting of a 3D convolution module and two Swin Transformer modules is input to generate aggregated features The multi-source spatiotemporal aggregation module uses 3D convolution to reduce the channel dimension of the spliced features and align the spatial positions of different feature sources; it adopts the window self-attention mechanism of SwinTransformer to model spatiotemporal dependencies, enhance key motion cues, and suppress redundant information; it retains the original feature information through residual connections, avoids the loss of high-frequency details, and ensures the reconstruction quality.
[0092] Step 3: Deformable spatiotemporal sparse transformer module refines features: In this embodiment, two cascaded deformable spatiotemporal sparse transformer modules are used to refine the aggregated features generated by the multi-source spatiotemporal aggregation module. Further processing. The specific steps are as follows:
[0093] The soft segmentation module (see the literature FuseFormer: Fusing Fine-Grained Information in Transformers for Video Inpainting) uses an overlapping expansion operation with a kernel size of (5×5), a step size of (2×2), and a padding of (2×2) to aggregate the aggregate features output by the multi-source spatiotemporal aggregation module. Then, the expanded features are linearly embedded and mapped into a 512-dimensional low-dimensional space, and the feature sequence is reorganized into a five-dimensional tensor to obtain a block sequence feature F with a block size of (32×32). block Adjacent blocks overlap by 3 pixels in the spatial dimension, effectively preserving boundary continuity information.
[0094] By cascading two deformable spatiotemporal sparse Transformer modules, the modules with different sequence numbers are dynamically selected for odd and even frame inputs through the temporal mask to achieve temporal feature decoupling. The output of the previous module is used as the input of the next module. The structure of the deformable spatiotemporal sparse Transformer module is as follows: Figure 5 As shown in the figure, the module first performs block sequence feature F block Perform layer normalization and input the sparse attention module to calculate the sparse attention A s . The original input block sequence feature F block With sparse attention A s Residual connection obtains intermediate feature F mid , and then the intermediate feature F mid After normalization, the nonlinear feature F is obtained by using the feedforward network processing. ffn , thus the nonlinear feature F ffn With the intermediate feature F mid Perform residual connection to obtain refined block sequence features
[0095] For sparse attention A s , the calculation steps are as follows:
[0096] For the fuzzy graph M t Perform two maximum pooling downsamplings to generate a downsampled blur map M that adapts to the input size of the deformable spatiotemporal sparse Transformer module t ′ , further calculations are performed to obtain the corresponding downsampled clear image S t ′ By setting a hard threshold, that is, whether the pixel value of the blur image is greater than 0.3, a binary spatial sparse mask is generated to distinguish high blur areas. t ′ In descending order, select the top 50% of the most blurred blur maps and multiply them element-wise with the spatial sparse mask to generate the query mask M q . For clear image S t ′ In descending order, select the top 50% of the clearest clear images and multiply them element-wise with the spatial sparse mask to generate the key-value mask M kv .
[0097] The normalized block sequence feature F blockPerform linear projection to generate query vector, key vector and value vector respectively. Divide the above three vectors into local windows according to the window size (8×8), each window contains 64 tokens, thus obtaining local query vector Q and local key vector K L , local value vector V L The deformable key vector K is predicted by the deformable sampling module D , deformable value vector V D . The pooled sequence feature F block Perform linear projection respectively to generate the global query vector K P , global key vector V P . With K L 、V L The extended key vector K and the extended value vector V are obtained by concatenating along the sequence length dimension. The expressions are as follows:
[0098] K=C(K L ,K D ,K P )
[0099] V=C(V L ,V D ,V P )
[0100] Where C represents the concatenation operation.
[0101] Combine the local query vector Q with the query mask M q and key-value mask M kv , distinguishing between highly blurred areas that require cross-window interaction (masked windows) and relatively clear areas that do not require cross-window interaction (unmasked windows). The global-local attention of the masked window is calculated using the expanded key vector K and the expanded value vector V through the multi-head attention module. For the unmasked window that only needs to pay attention to the area within the window, the local key vector K is used. L , local value vector V L The local attention is calculated by the multi-head attention module. The attention outputs of the two types of windows are combined to obtain the sparse attention A s .
[0102] For refined block sequence features The soft composition module (see the literature FuseFormer: Fusing Fine-Grained Information in Transformers for Video Inpainting) is used to restore it to the refined features of the dimension size before the soft segmentation operation Specifically, through the linear inverse embedding mapping reduced to the original high-dimensional space. Subsequently, the high-dimensional space is subjected to feature reorganization by an inverse sliding window folding operation with the same parameters as the soft segmentation module. Finally, the local features are refined by a 3x3 convolution to obtain refined features
[0103] Step 4, Reconstructing the clear video: input the refined features to the decoder module for processing, and connect the residual error with the original input (n video frames to be deblurred) to obtain the final deblurring result The expression is as follows:
[0104] R t =Decoder(F t )+x t ,t∈[0,n-1]
[0105] where Decoder represents the decoder, x t represents the t-th frame input.
[0106] The method is implemented under the Windows operating system based on the PyTorch framework, and the hardware platform is configured as NVIDIA GeForce RTX 4060 Laptop GPU and Intel(R) Core(TM) i9-14900HX CPU. The experimental environment configuration is only used to verify the effectiveness of the method, and does not constitute a limitation on the deployment scenario.
[0107] The embodiment of the application adopts the commonly used GoPro dataset in the field of video deblurring for experiments, which consists of 3214 pairs of blurred and clear images with a resolution of 1280x720, covering 33 video scenes. Among them, 2103 image pairs are used for training, and 1111 image pairs are used for testing.
[0108] In the training stage, the number of training iterations is set to 50000, and the blurred-clear image pairs in the training dataset are randomly cropped to 256x256 pixels as model inputs for each iteration, and the batch size is set to 1, and each batch contains 6 video frames. The Adam optimizer is used to update the model parameters. During the training process, the cosine annealing scheduling method is used to dynamically adjust the learning rate, so that the learning rate value gradually decays from the configured initial value 4x10 -4 to 1x10 -7 .
[0109] The experimental results of the embodiment of the application are as follows:
[0110] Quantitative evaluation on the GoPro standard test set shows that its Peak Signal-to-Noise Ratio (PSNR) reaches 32.85dB and its Structural Similarity Index Measure (SSIM) is 0.938.
[0111] Figure 6 This is the result after deblurring some blurred video frames in the test set. Figure 5 (a), (c), (e), and (f) are continuous fuzzy input frames. Figure 6 (b), (d), (f), and (h) are the deblurred results of (a), (c), (e), and (f), respectively.
[0112] Figure 7 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 7 The electronic device provided in this embodiment includes: a memory and a processor, wherein the memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, a video deblurring method based on a deformable spatiotemporal sparse transformer of the present invention is implemented.
[0113] It should be noted that, in addition to Figure 7 In addition to the memory and processor shown, the electronic device may also include other hardware according to its actual functions, which will not be described in detail.
[0114] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the video deblurring method based on a deformable spatiotemporal sparse transformer is implemented.
[0115] The present invention also provides a computer program product, comprising a computer program, which implements the above-mentioned video deblurring method based on a deformable spatiotemporal sparse transformer when executed by a processor.
[0116] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0117] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0118] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0120] The above embodiments are intended only to illustrate the design concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. The scope of protection of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or modifications made based on the principles and design concepts disclosed in the present invention are within the scope of protection of the present invention.
Claims
1. A video deblurring method based on a deformable spatiotemporal sparse transformer, characterized in that: The following steps are involved: Step S1: For the input n video frames to be deblurred, use the encoder module to extract features And use the optical flow estimation network to obtain the forward optical flow between adjacent frames With backward optical flow Then, the blur map of each frame is estimated based on the forward and backward optical flows; Step S2, multi-level downsampling of the fuzzy image and fusion to generate a multi-scale fuzzy image, and the extracted features The forward optical flow, backward optical flow, and multi-scale fuzzy map are input into the bidirectional feature propagation module, and the sequence frame features are recursively updated through two forward propagation branches and two backward propagation branches. The features output by the four propagation branches of the bidirectional feature propagation module are then combined with the features extracted by the encoder. Perform splicing and input into the multi-source spatiotemporal aggregation module to generate aggregation features Step S3: Use the deformable spatiotemporal sparse Transformer module to aggregate features Processing to obtain refined features Step S4, refine the features The input is processed by the decoder module and connected with the original input residual to obtain the final deblurred result.
2. The video deblurring method based on a deformable spatiotemporal sparse transformer according to claim 1, wherein: In step S1, the blurred video of consecutive n frames is used as the input sequence unit, and the encoder module is used to extract features frame by frame to obtain the features of each frame. The expression is as follows: f t =Encoder(x tt ),t∈[0,n-1] Among them, f t Represents the features of the t-th frame, Encoder represents the encoder, x t represents the t-th frame input; Use the pre-trained RAFT network as the optical flow estimation network to obtain the forward optical flow between adjacent frames With backward optical flow And calculate the fuzzy map corresponding to each frame of the input sequence unit The expression is as follows: For fuzzy graph Normalize and get the normalized fuzzy image By fuzzy graph Calculate the clear image corresponding to each frame of the input sequence unit The expression is as follows: S t =1-M t ,t∈[0,n-1] in, and Respectively represent the minimum and maximum values of the fuzzy graph in the input sequence unit.
3. The video deblurring method based on a deformable spatiotemporal sparse transformer according to claim 1, wherein: In step S2, a multi-scale fuzzy image fusion module is used to fuzzy image Perform two-level maximum pooling downsampling to generate the first and second level blur maps step by step, and fuse them with the original blur map across scales; finally, a multi-level integrated multi-scale blur map is obtained. The corresponding multi-scale clear image is In the bidirectional feature propagation module, two backward propagation branches and two backward propagation branches are constructed to propagate features respectively; each propagation branch contains a deformable alignment module, which aligns the features of adjacent frames through the multi-scale blur map and the optical flow-guided deformable alignment module to reconstruct the current frame; the specific steps are as follows: For forward propagation, the video sequence is processed frame by frame from the starting frame to the ending frame; at the t-th time step of the j-th propagation branch, the optical flow O from the current frame to the previous two frames is used. t→t-1 , O t→t-2 , for the features of the first two time steps Perform space warping to generate preliminary alignment features The expression is as follows: Among them, W represents the space-warping operation; The specific steps of the multi-scale blur map and optical flow guided deformable alignment module are as follows: Initially align features The optical flow O from the current frame to the previous two frames t→t-1 , O t→t-2 , the clear picture of the first two frames and the features of the previous branch at the tth time step Splicing is performed; after processing through a convolutional layer consisting of four 2D convolutions, the residual offset and residual mask are obtained; Using optical flow O t→t-1 , O t→t-2 Guide the initial offset generation of deformable convolution: take the pre-estimated optical flow fields of the first two frames as the base offset, predict the residual offset through the convolution layer, and then superimpose the base offset and the residual offset to obtain the initial offset; Use clear pictures Guided deformable convolution initial mask generation: Introducing a clarity map-driven adaptive mask modulation mechanism, linearly mapping spatial clarity information to the mask space to generate a base mask. Using a convolutional network to predict the residual mask, the base mask and the residual mask are superimposed, and the offset weights of different regions are dynamically adjusted using a sigmoid function to obtain the initial mask. The features of the first two frames are aligned through deformable convolution with parameters of initial offset and initial mask; thus, the features of the t-th time step of the previous branch are aligned. Splicing; processed by a convolution layer consisting of 4 2D convolutions to obtain the current time step aggregation feature Backward propagation is performed backward along the time step, and its process is the same as the forward propagation process; the features output by the four propagation branches of the bidirectional feature propagation module are compared with the features extracted by the encoder. After splicing, input the multi-source spatiotemporal aggregation module to generate aggregate features 4. The video deblurring method based on a deformable spatiotemporal sparse transformer according to claim 1 or 3, characterized in that: The multi-source spatiotemporal aggregation module includes a 3D convolution module and two Swin Transformer modules which are connected in sequence.
5. The video deblurring method based on a deformable spatiotemporal sparse transformer according to claim 1, wherein: In step S3, two cascaded deformable spatiotemporal sparse Transformer modules are used to aggregate features generated by the multi-source spatiotemporal aggregation module. deal with.
6. The video deblurring method based on a deformable spatiotemporal sparse transformer according to claim 1 or 5, characterized in that: Aggregate features Processing, including: The soft segmentation module is used to aggregate features Expand the sliding window to obtain the block sequence feature F block ; Block sequence feature F block After layer normalization, the sparse attention module is input to calculate the sparse attention A s ; Block sequence feature F block With sparse attention A s Residual connection obtains intermediate feature F mid , and then the intermediate feature F mid After normalization, the nonlinear feature F is obtained by using the feedforward network processing. ffn , thus the nonlinear feature F ffn With the intermediate feature F mid Perform residual connection to obtain refined block sequence features For refined block sequence features The soft reorganization module is used to restore it to the refined features of the dimension size before the soft segmentation operation 7. The video deblurring method based on a deformable spatiotemporal sparse transformer according to claim 6, characterized in that: Sparse AttentionA s The calculation steps are as follows: For the fuzzy graph M t Perform two maximum pooling downsamplings to generate a downsampled blur map M that adapts to the input size of the deformable spatiotemporal sparse Transformer module t ′ , and calculate the corresponding downsampled clear image S t ′ ; By setting a hard threshold, a binary spatial sparse mask is generated to distinguish high fuzzy areas; t ′ In descending order, select the first w% most blurred blur maps and multiply them element-wise with the spatial sparse mask to generate the query mask M q ; For clear image S t ′ In descending order, select the top w% of the clearest clear maps and multiply them element-wise with the spatial sparse mask to generate the key-value mask M kv ; The normalized block sequence feature F block Perform linear projection to generate query vector, key vector, and value vector respectively, and divide them into local windows according to the window size to obtain local query vector Q and local key vector K L , local value vector V L ; The deformable key vector K is predicted by the deformable sampling module D , deformable value vector V D ; The normalized block sequence feature F block After the pooling layer, the pooled sequence feature F block Perform linear projection respectively to generate the global query vector K P , global key vector V P ; K L , K D , K P Splice to get the expanded key vector K, and then convert V L 、V D 、V P Perform concatenation to obtain the expanded value vector V; Combine the local query vector Q with the query mask M q and key-value mask M kv , distinguish between masked windows and unmasked windows; use the expanded key vector K and the expanded value vector V to calculate the global-local attention of the masked window through the multi-head attention module; for the unmasked window, use the local key vector K L , local value vector V L The local attention is calculated through the multi-head attention module; the attention outputs of the two types of windows are merged to obtain the sparse attention A s .
8. An electronic device comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the video deblurring method based on a deformable spatiotemporal sparse transformer as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a video deblurring method based on a deformable spatiotemporal sparse transformer as described in any one of claims 1 to 7 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the video deblurring method based on a deformable spatiotemporal sparse transformer according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Optical flow estimation method and system based on adaptive flow propagation
CN121904401A
An optical flow estimation method and system based on adaptive streaming
CN121904401B