Lightweight video deblurring method based on Shift operation

By introducing Shift operations and space-time shift blocks into the video defuzzing method, combining dynamic filtering and multi-frequency enhancement modules, the problems of insufficient inter-frame information utilization and high computational complexity in the prior art are solved, and efficient video defuzzing effect is achieved.

CN120147183APending Publication Date: 2025-06-13DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510212322.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing video defuzzing methods based on CNN, RNN and Transformer are difficult to effectively utilize interframe information in video sequences, and the frame alignment method based on dynamic filtering is very complex in computing when processing large motions of space-time variations.

Method used

The lightweight video defuzzing method based on Shift operation is adopted. Through the shift-guided dynamic filtering module and the space-time shift block, adjacent inter-frame information is used and intraframe information is strengthened, and the multi-frequency enhancement module and pixel difference convolution module are combined to realize video defuzzing.

Benefits of technology

It realizes better video defuzzing performance, while reducing model complexity and improving the processing ability of large motions of space-time changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147183A_ABST
    Figure CN120147183A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight video deblurring method based on Shift operation, and the method comprises the steps: carrying out the preprocessing of an input fuzzy video frame through a shallow feature extraction module, and obtaining a shallow extraction feature; frequency domain features in the shallow layer extraction features are extracted by using a multi-frequency enhancement module, and deep layer extraction features are obtained through an input frame extraction module; frame alignment is carried out on the deep extraction features through a shift-guided dynamic filtering module, and a space-time shift block uses adjacent inter-frame information and reinforces intra-frame information to obtain fusion features; the fusion feature, the deep layer extraction feature and the shallow layer extraction feature are spliced, a target frame reconstruction module reconstructs a target clear frame, a pixel difference convolution module is adopted to enhance texture details of the shallow layer extraction feature, the reconstructed target clear frame and an original fuzzy video frame are fused, and a final deblurring result is obtained. According to the method, the potential of shift operation for video deblurring is further explored, inter-frame information in a fuzzy video sequence is effectively utilized, and meanwhile, the model complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video deblurring in computer vision. In particular, it relates to a lightweight video deblurring method based on Shift operation. Background Art

[0002] Existing methods based on CNN, RNN, and Transformer cannot effectively utilize the inter-frame information in video sequences. Adjacent frames are generally utilized by stitching or frame alignment, which results in either limited network performance or high complexity. Therefore, inspired by the temporal shift operation, we simply perform cyclic temporal shifts on adjacent frames to achieve effective utilization.

[0003] Existing frame alignment methods based on dynamic filtering generally only perform simple operations on the fixed surrounding fields of target location features, which cannot effectively handle large motions with spatio-temporal variations. In addition, this method usually only applies filtering to the reference frame, which limits its ability to accurately utilize adjacent frame information. To fully utilize the motion information of adjacent frames, large-sized filters are required to capture large motions, resulting in very high computational complexity. Therefore, inspired by the spatial shift operation, a Shift-Guided Dynamic Filter method is designed to address the above problems. Summary of the Invention

[0004] According to the above-mentioned technical problems, a lightweight video deblurring method based on Shift operation is provided. The present invention improves the deficiencies of existing dynamic filtering methods. Compared with some existing methods, the proposed network is more lightweight and achieves better performance at the same time.

[0005] The technical means adopted by the present invention are as follows:

[0006] A lightweight video deblurring method based on Shift operation, comprising:

[0007] S1. Select a publicly available video deblurring dataset as the ground truth, and use a shallow feature extraction module to preprocess the input blurred video frame I i to obtain a shallow extracted feature x i ;

[0008] S2. Input the shallow extracted feature x i obtained in step S1 into a multi-frequency enhancement module to further extract frequency domain features, and then obtain a deep extracted feature f i through an input frame extraction module;

[0009] S3. Input the obtained deep extracted feature f iFirst, it is initially frame-aligned by a shift-guided dynamic filtering module, and then sent to a spatio-temporal shift block to further utilize the information between adjacent frames and strengthen the intra-frame information to obtain the fused feature F i ;

[0010] S4. Concatenate the fused feature F i , the deep extraction feature f i and the shallow extraction feature x i together, and input them into the target frame reconstruction module to reconstruct the target clear frame. Use a pixel difference convolution module to further enhance the texture details of the shallow extraction feature x i , and fuse the reconstructed target clear frame with the original blurred video frame I i to obtain the final de-blurring result R i ;

[0011] S5. Use the L1 Charbonnier loss and the L1 loss in the frequency domain to jointly train the entire network.

[0012] Further, the shallow extraction feature x i obtained in step S1 is specifically as follows:

[0013] x i = SFE(I i )

[0014] where I i represents the blurred input frame, x i represents the obtained shallow extraction feature, and SFE represents the shallow feature extraction module, which consists of a series of channel attention blocks.

[0015] Further, the deep extraction feature f i obtained in step S2 is specifically as follows:

[0016] f i = IFE(MFE(x i ))

[0017] where x i represents the shallow extraction feature, f i represents the obtained deep extraction feature, MFE represents the multi-frequency enhancement module, and IFE represents the input frame extraction module, which consists of a series of channel attention blocks.

[0018] Further, the multi-frequency enhancement module includes a spatial frequency domain enhancement module, an energy frequency domain enhancement module, and an Aux module composed of 7x1 and 1x7 depth convolutions; where:

[0019] The spatial frequency domain enhancement module includes a dynamic frequency domain filtering and a convolution branch;

[0020] The energy frequency domain enhancement module includes the parameter-free attention SimAM for extracting energy features and DCT filtering, where the parameter-free attention SimAM extracts energy features The principle is shown in the following formula:

[0021]

[0022] Among them, represents the mean of the input feature map of SimAM, represents the variance of the input feature map of SimAM, δ represents the hyperparameter, and e i,j represents the energy value at the target pixel t i,j while represents rescaling the energy value using the LeakyReLU activation function to limit too large energy values.

[0023] Furthermore, the fused feature F obtained in step S3 i is specifically as follows:

[0024] F i = STSB(SGDF(f i ))

[0025] Among them, f i represents the deeply extracted feature, F i represents the obtained fused feature, SGDF represents the shift-guided dynamic filtering module, and STSB represents the U-Net network composed of spatio-temporal shift blocks.

[0026] Furthermore, in step S3:

[0027] The shift-guided dynamic filtering module (SGDF) improves the traditional dynamic filtering method by introducing spatial shifts for preliminary mapping;

[0028] The spatio-temporal shift block includes a cyclic time shift block and dynamic local self-attention (DLSA), and a shift frequency domain self-attention module (XFSA), which are used as basic units to form a U-Net network for feature fusion; among them:

[0029] The cyclic time shift block can utilize the information between adjacent frames;

[0030] The dynamic local self-attention further introduces dynamic filtering and a gated feed-forward layer to enhance the intra-frame information;

[0031] The shift frequency domain self-attention module introduces adaptive frequency domain filtering and pixel shift attention to strengthen the spatial frequency domain features, which replaces the traditional self-attention operation.

[0032] Furthermore, the final deblurring result R obtained in step S4 iSpecifically as follows:

[0033] R i = TFR(Cat(PDC(x i ), f i , F i )) + I i

[0034] Among them, I i represents the blurred input frame, R i represents the finally restored de-blurred frame, x i represents the shallow feature extraction, f i represents the deep feature extraction, F i represents the obtained fused feature, TFR represents the target frame reconstruction module, which consists of a series of channel attention blocks; Cat represents the concatenation operation, PDC represents the pixel difference convolution module, which includes 4 difference convolution layers and 1 ordinary convolution layer, and the difference convolution layer includes central difference convolution, corner difference convolution, horizontal difference convolution and vertical difference convolution, which are used for further feature extraction to enhance texture details.

[0035] Furthermore, in step S5, the L1 Charbonnier loss and the L1 loss in the frequency domain are used to jointly train the entire network, specifically as follows:

[0036]

[0037] L fre = ||FFT(R t ) - FFT(G t )|| 1

[0038] L all = L char + L fre

[0039] Among them, R t represents the restored de-blurred frame, G t represents the corresponding real clear frame, ε is a constant for stabilizing training, and FFT represents the Fourier transform.

[0040] Compared with the prior art, the present invention has the following advantages:

[0041] 1. A lightweight video de-blurring method based on the Shift operation provided by the present invention further explores the potential of the shift operation for video de-blurring. In order to effectively utilize the inter-frame information in the blurred video sequence and at the same time reduce the model complexity, only a simple cyclic time shift operation is performed on adjacent frames, and a spatio-temporal shift block is further constructed.

[0042] 2. The present invention proposes a Shift-Guided Dynamic Filter method, which improves the deficiencies of existing dynamic filtering methods. Compared with some existing methods, the proposed network is more lightweight and achieves better performance at the same time.

[0043] For the above reasons, the present invention can be widely promoted in fields such as video deblurring in computer vision. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0045] Figure 1 It is a flowchart of the method of the present invention.

[0046] Figure 2 It is a schematic diagram of the shift-guided dynamic filtering module provided by the embodiment of the present invention.

[0047] Figure 3 It is a schematic diagram of the pixel difference convolution module (PDC) and the multi-frequency enhancement module (MFE) provided by the embodiment of the present invention.

[0048] Figure 4 It is a schematic diagram of the cyclic time shift module provided by the embodiment of the present invention.

[0049] Figure 5 It is a schematic diagram of the dynamic local self-attention (DLSA) and the shift frequency domain self-attention module (XFSA) provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0052] The present invention proposes a lightweight video deblurring method based on the shift operation, improves the existing frame alignment method based on dynamic filtering, and further explores the potential of the shift operation for video deblurring. As Figure 1 shown, the overall model follows the Encoder-Decoder architecture.

[0053] The overall process is as follows:

[0054] S1. Select a publicly available video deblurring dataset as the ground truth, and use the shallow feature extraction module to preprocess the input blurred video frame I i to obtain the shallow extraction feature x i ;

[0055] S2. Input the shallow extraction feature x i obtained in step S1 into the multi-frequency enhancement module (MFE) to further extract frequency domain features, and then obtain the deep extraction feature f i through the input frame extraction module (IFE);

[0056] S3. First, perform preliminary frame alignment on the obtained deep extraction feature f i through the shift-guided dynamic filtering module (SGDF), and then send it to the spatio-temporal shift block to further utilize the information between adjacent frames and strengthen the intra-frame information to obtain the fusion feature F i ;

[0057] S4. Concatenate the fusion feature F i , the deep extraction feature f i and the shallow extraction feature x i , and input them into the target frame reconstruction module (TFR) to reconstruct the target clear frame. Use the pixel difference convolution module (PDC) to further enhance the texture details of the shallow extraction feature x i , and fuse the reconstructed target clear frame with the original blurred video frame I i to obtain the final deblurring result Ri ;

[0058] S5. Jointly train the entire network using the L1 Charbonnier loss and the L1 loss in the frequency domain.

[0059] In specific implementation, as a preferred implementation manner of the present invention, the shallow extraction feature x obtained in step S1 i is specifically as follows:

[0060] x i = SFE(I i )

[0061] wherein, I i represents the blurred input frame, x i represents the obtained shallow extraction feature, and SFE (Shallow Feature Extractor) represents the shallow feature extraction module, which is composed of a series of channel attention blocks (CAB).

[0062] In specific implementation, as a preferred implementation manner of the present invention, the deep extraction feature f obtained in step S2 i is specifically as follows:

[0063] f i = IFE(MFE(x i ))

[0064] wherein, x i represents the shallow extraction feature, f i represents the obtained deep extraction feature, MFE represents the multi-frequency enhancement module, and IFE represents the input frame extraction module, which is composed of a series of channel attention blocks (CAB).

[0065] In specific implementation, as a preferred implementation manner of the present invention, as Figure 3 (b) shows, the multi-frequency enhancement module (MFE) includes a spatial frequency domain enhancement module (SFEM), an energy frequency domain enhancement module (EFEM), and an Aux module composed of 7x1 and 1x7 depth convolutions; wherein:

[0066] As Figure 3 (c) shows, the spatial frequency domain enhancement module (SFEM) includes a dynamic frequency domain filtering and convolution branch;

[0067] As Figure 3 (d) shows, the energy frequency domain enhancement module (EFEM) includes a parameter-free attention SimAM for extracting energy features and a DCT filter, wherein the parameter-free attention SimAM extracts energy features The principle is shown in the following formula:

[0068]

[0069] Among them, represents the mean of the input feature map of SimAM, represents the variance of the input feature map of SimAM, δ represents a hyperparameter, and e i,j represents the energy value at the target pixel t i,j while represents re-scaling the energy value using the LeakyReLU activation function to limit overly large energy values.

[0070] In specific implementation, as a preferred implementation manner of the present invention, the fused feature F i obtained in step S3 is specifically as follows:

[0071] F i = STSB(SGDF(f i ))

[0072] Among them, f i represents the deeply extracted feature, and F i represents the obtained fused feature. SGDF represents a shift-guided dynamic filtering module, and STSB represents a U-Net network composed of spatio-temporal shift blocks.

[0073] In specific implementation, as a preferred implementation manner of the present invention, in step S3:

[0074] As Figure 2 shown, the shift-guided dynamic filtering module (SGDF) improves the traditional dynamic filtering method by introducing spatial shift for preliminary mapping;

[0075] The spatio-temporal shift block includes a cyclic time shift block and dynamic local self-attention (DLSA), a shift frequency domain self-attention module (XFSA), which are used as basic units to form a U-Net network for feature fusion; among them:

[0076] As Figure 4 shown, the cyclic time shift block can utilize the information between adjacent frames;

[0077] As Figure 5 shown, the dynamic local self-attention (DLSA) further introduces dynamic filtering and a gated feed-forward layer to enhance the intra-frame information;

[0078] Continuing to refer to Figure 5 , the shift frequency domain self-attention module (XFSA) (X-shift and Frequency Self-Attention) introduces adaptive frequency domain filtering and pixel shift attention to strengthen the spatial frequency domain features, which replaces the traditional self-attention operation.

[0079] In specific implementation, as a preferred implementation manner of the present invention, the final deblurred result R obtained in step S4 i is as follows:

[0080] R i = TFR(Cat(PDC(x i ), f i , F i )) + I i

[0081] wherein, I i represents the blurred input frame, R i represents the finally restored deblurred frame, x i represents the shallow feature extraction, f i represents the deep feature extraction, F i represents the obtained fused feature, TFR represents the target frame reconstruction module, which is composed of a series of channel attention blocks; Cat represents the concatenation operation, and PDC represents the pixel difference convolution module, as Figure 3 (a) shown, which includes 4 difference convolution layers and 1 ordinary convolution layer, wherein the difference convolution layer includes the central difference convolution (CDC), the angular difference convolution (ADC), the horizontal difference convolution (HDC) and the vertical difference convolution (VDC), and is used for further feature extraction to enhance texture details.

[0082] In specific implementation, as a preferred implementation manner of the present invention, in step S5, the L1 Charbonnier loss and the L1 loss in the frequency domain are used to jointly train the entire network, which is as follows:

[0083]

[0084] L fre = ||FFT(R t ) - FFT(G t )|| 1

[0085] L all = L char + L fre

[0086] wherein, R t represents the restored deblurred frame, G t represents the corresponding real clear frame, ε is a constant used to stabilize the training, and FFT represents the Fourier transform.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A lightweight video deblurring method based on Shift operation, characterized in that: include: S1. Select a public video deblurring dataset as the true value and use the shallow feature extraction module to extract the input blurred video frame I i Preprocessing is performed to obtain shallow extraction features x i ; S2, extract the shallow feature x obtained in step S1 i The frequency domain features are further extracted in the input multi-frequency enhancement module, and then the deep extraction features f are obtained by the input frame extraction module. i ; S3, extract the deep features f i The initial frame alignment is first performed in the shift-guided dynamic filtering module, and then sent to the spatiotemporal shift block to further utilize the information between adjacent frames and enhance the intra-frame information to obtain the fusion feature F i ; S4, fusion feature F i , deep extraction features f i And shallowly extracted features x i The pixels are stitched together and input into the target frame reconstruction module to reconstruct the target clear frame. The pixel difference convolution module is used to further enhance the shallow layer extracted features x i The texture details of the target clear frame are reconstructed and compared with the original blurred video frame I i Fusion, get the final deblurred result R i ; S5. Use L1 Charbonnier loss and L1 loss in frequency domain to jointly train the entire network.

2. According to the lightweight video deblurring method based on Shift operation in claim 1, it is characterized in that: The shallow extraction feature x obtained in step S1 i The details are as follows: x i =SFE(I i ) Among them, I i represents the blurred input frame, x i Represents the shallow extracted features obtained, SFE represents the shallow feature extraction module, which is composed of a series of channel attention blocks.

3. The lightweight video deblurring method based on Shift operation according to claim 1, characterized in that: The deep extraction feature f obtained in step S2 i The details are as follows: f i =IF(MFE(x i )) Among them, x i represents shallow extraction features, f i represents the obtained deep extracted features, MFE represents the multi-frequency enhancement module, and IFE represents the input frame extraction module, which is composed of a series of channel attention blocks.

4. The lightweight video deblurring method based on Shift operation according to claim 3, characterized in that: The multi-frequency enhancement module includes a spatial frequency domain enhancement module, an energy frequency domain enhancement module and an Aux module composed of 7x1 and 1x7 depth convolutions; wherein: The spatial frequency domain enhancement module includes dynamic frequency domain filtering and convolution branches; The energy frequency domain enhancement module includes a parameter-free attention SimAM and a DCT filter for extracting energy features. The principle is as follows: in, represents the mean of the input feature map of SimAM, represents the variance of the input feature map of SimAM, δ represents the hyperparameter, e i,j represents the target pixel t i,j The energy value at This means that the LeakyReLU activation function is used to rescale the energy value to limit the energy value that is too large.

5. The lightweight video deblurring method based on Shift operation according to claim 1, characterized in that: The fusion feature F obtained in step S3 i The details are as follows: F i =STSB(SGDF(f i )) Among them, f i represents deep extraction features, F i represents the obtained fusion features, SGDF represents the shift-guided dynamic filtering module, and STSB represents the U-Net network composed of spatiotemporal shift blocks.

6. The lightweight video deblurring method based on Shift operation according to claim 5, characterized in that: In step S3: The shift-guided dynamic filtering module improves the traditional dynamic filtering method by introducing spatial shift for preliminary mapping; The spatiotemporal shift block includes a cyclic time shift block and a dynamic local self-attention and shifted frequency domain self-attention module, which serve as basic units to form a U-Net network for feature fusion; wherein: The cyclic time shift block can utilize the information between adjacent frames; Dynamic local self-attention further introduces dynamic filtering and gated feed-forward layers to enhance intra-frame information; The shifted frequency domain self-attention module introduces adaptive frequency domain filtering and pixel shift attention to enhance the spatial frequency domain features, which replaces the traditional self-attention operation.

7. The lightweight video deblurring method based on Shift operation according to claim 1, characterized in that: The final deblurred result R obtained in step S4 i The details are as follows: R i =TFR(Cat(PDC(x i ),f i ,F i ))+I i Among them, I i represents the blurred input frame, R i represents the final restored deblurred frame, x i represents shallow extraction features, f i represents deep extraction features, F i represents the obtained fusion features, TFR represents the target frame reconstruction module, which is composed of a series of channel attention blocks; Cat represents the splicing operation, PDC represents the pixel difference convolution module, which includes 4 difference convolution layers and 1 ordinary convolution layer, where the difference convolution layer includes center difference convolution, angle difference convolution, horizontal difference convolution and vertical difference convolution, which are used for further feature extraction to enhance texture details.

8. The lightweight video deblurring method based on Shift operation according to claim 1, characterized in that: In step S5, the L1 Charbonnier loss and the L1 loss in the frequency domain are used to jointly train the entire network, as follows: L fre =||FFT(R t )-FFT(G t )||1 L all =L char +L fre Among them, R t represents the restored deblurred frame, G t represents the corresponding real clear frame, ε is a constant used to stabilize training, and FFT represents Fourier transform.