Video deblurring method based on memory diffusion network

By using a memory diffusion network in the video defuzzing method, combined with diffusion prior features and memory database, the problem that existing methods are difficult to capture local motion characteristics between video frames is solved, and more efficient video defuzzing effect and time consistency are achieved.

CN120147182AActive Publication Date: 2025-06-13DALIAN MARITIME UNIVERSITY

Patent Information

Application Number
CN202510212321.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The existing video defuzzing method based on CNN, RNN and self-attention mechanisms is difficult to effectively capture local motion-related features between video frames, and depends on partial input frame information, so it is impossible to effectively process local features in space-time.

Method used

The video defuzzing method based on memory diffusion network is adopted, and the diffusion branches generated by the lightweight diffusion network are constructed by constructing a bidirectional cyclic structure, combining diffusion prior features and memory databases to achieve effective utilization of spatiotemporal relationships between video frames.

Benefits of technology

The effect of debuffing videos is improved, and the blur can be debuffed more granularly, ensuring the temporal consistency between frames, and improving the memory update mechanism to better utilize the spatio-temporal information between video frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147182A_ABST
    Figure CN120147182A_ABST
Patent Text Reader

Abstract

The invention provides a video deblurring method based on a memory diffusion network. The core of the method comprises a deblurring branch and a diffusion branch. The deblurring branch has a bidirectional circulation structure, maintains a simple and effective memory bank containing a specific updating mechanism, and can effectively utilize the space-time relationship between video frames to further ensure the time consistency between the frames in the final deblurred video. The diffusion branch comprises a lightweight diffusion network to generate multi-scale diffusion prior features. According to the method, the effect of an existing regression model acting on video deblurring is improved by combining the diffusion prior generated by the diffusion network. A simplified and effective memory bank is constructed to screen and store time frame information to accurately capture a complex space-time local relationship in a fuzzy video so as to realize deblurring with finer granularity, and a specific memory updating algorithm is designed to refresh the memory bank. The space-time information between continuous video frames is better utilized, and the time consistency between the frames is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a video deblurring method based on a memory diffusion network. Background Art

[0002] In reality, the blurring in videos may have various sources, such as object motion, camera shake, and scene depth change. Usually, these multiple factors are intertwined with each other, making it difficult to accurately capture the complex spatio-temporal relationships in blurred videos, thus posing challenges to subsequent deblurring.

[0003] Currently, a variety of deep learning-based video deblurring methods have been proposed to utilize the complex spatio-temporal information in the input frame sequence from different aspects. On the one hand, CNN-based models usually send the current frame and adjacent consecutive blurred frames to an encoder-decoder to directly restore the clear frame. On the other hand, RNN-based models maintain a hidden state between video frames to sequentially propagate features from the first frame to the last frame. However, this method only relies on the information of previous frames and cannot capture more complex motion relationships. Recently, the work on using self-attention mechanisms to handle the problem of video blurring complexity (such as spatio-temporal inhomogeneity of blurring) is also increasing. However, these networks usually simply utilize a rough attention mechanism and cannot effectively capture the local motion-related features between video frames.

[0004] Existing methods based on CNN, RNN, and self-attention mechanisms generally can only utilize the information of some input frames and cannot effectively process spatio-temporal local features. Adjacent frames are often utilized through simple splicing, frame alignment, etc., which results in either limited network performance or high computational complexity. In this regard, recently, some work has explored the potential of memory networks for video deblurring. It can not only more effectively capture spatio-temporal local features but also utilize the spatio-temporal information of all input frames, which is exactly what other methods lack. However, the existing memory network-based methods only simply maintain a certain number of memory frames and reset them when they expire, thus discarding all the previous memory information and failing to achieve the purpose of effective utilization. This is not an effective memory update mechanism.

[0005] Most of the recent methods (including the recent memory network-based methods) are based on regression models, which tend to be highly sensitive to degradation or have limited expressiveness because these networks essentially still rely on the information of the highly corrupted LQ input frames. The generative prior can alleviate this situation to a certain extent. For example, some image restoration methods further introduce the Ground Truth during training by combining diffusion networks instead of just the degraded LQ frames. The diffusion prior generated by adding and removing noise in the latent space effectively complements the adopted regression model, and the current video deblurring methods have not studied it deeply. Summary of the Invention

[0006] According to the above-mentioned technical problems, a video deblurring method based on a memory diffusion network is provided. The present invention further explores the potential of the diffusion prior generated by combining the existing regression model with the diffusion network for video deblurring. Its core includes two branches: a deblurring branch and a diffusion branch. The deblurring branch has a bidirectional cyclic structure and maintains a concise and effective memory bank with a specific update mechanism, which can effectively utilize the spatio-temporal relationship between video frames and further ensure the temporal consistency between frames in the final deblurred video. The diffusion branch contains a lightweight diffusion network to generate multi-scale diffusion prior features.

[0007] The technical means adopted by the present invention are as follows:

[0008] A video deblurring method based on a memory diffusion network, comprising:

[0009] S1. Select a publicly available video deblurring dataset as the ground truth, preprocess the input blurred video frames using a residual dense block, and downsample to obtain the initial downsampled feature x i ;

[0010] S2. Concatenate the blurred input frame and the corresponding Ground Truth frame, generate random noise through a diffusion forward noise addition process after encoding by a latent encoder, and then perform reverse denoising by a diffusion denoising network to generate multi-scale diffusion prior features

[0011] S3. Send the downsampled feature x obtained in step S1 i to a feature encoder for encoding to obtain multi-scale encoded features; during encoding, fuse the multi-scale diffusion prior features generated in step S2 into the encoder through a hierarchical integration module;

[0012] S4. Before decoding, construct a query key using the multi-scale encoded features obtained in step S3 to generate a query for retrieving the memory features of the memory bank, and send the retrieved memory features, the maintained instantaneous perceptual memory, and the encoded features of the third scale to the memory decoder to enhance the previous encoded features;

[0013] S5. Input the enhanced encoded features in step S4 into the feature decoder for decoding to obtain multi-scale decoded features;

[0014] S6. Input the decoded features of the first scale obtained, the features extracted by the previous cycle feature extraction module, and the initial downsampled features into the cycle feature extraction module to further enhance the decoded features;

[0015] S7. Input the previous encoded features, the enhanced decoded features, and the initial downsampled features into the target frame reconstruction module to reconstruct the target frame, then upsample through transposed convolution and fuse with the original blurred input to obtain the final deblurred result;

[0016] S8. Use the L1 Charbonnier loss and the L1 loss during the diffusion process to jointly train the entire network.

[0017] Further, in step S1, the initial downsampled feature x i , is as follows:

[0018] x i = RDB(I i ), i = 1, 2,..., N

[0019] where x i represents the downsampled feature, I i represents the blurred input frame I Blur , and RDB represents the residual dense block for downsampling.

[0020] Further, step S2 specifically includes:

[0021] S21. During the forward diffusion process, transform z into Gaussian noise through T iterations as shown in the following formula:

[0022]

[0023] where t = 1, 2,..., T; β t ∈(0, 1) represents the hyperparameter for controlling the noise variance; N represents the Gaussian distribution; I represents the standard normal distribution;

[0024] S22. Let α t = 1 - β t , Through iterative derivation of reparameterization, the formula in step S21 is rewritten as follows:

[0025]

[0026] where z 0 = z;

[0027] S23. During the reverse denoising process, DM samples a Gaussian random noise z T , and then gradually denoises z T back to z 0 as shown in the following formula:

[0028]

[0029] where

[0030] S24. Using the formula to replace z 0 then we get:

[0031]

[0032] Removing the variance estimation in we get:

[0033]

[0034] where the noise ε is the only uncertain variable.

[0035] Furthermore, step S3 specifically includes:

[0036]

[0037] where f i ej represents the multi-scale forward encoding features encoded by the encoder Encoder, represents the multi-scale backward encoding features encoded by the encoder Encoder, x i represents the downsampled features, represents the multi-scale diffusion prior features generated by the diffusion network, j = 1, 2, 3 represent 3 different scales, represents the encoding and decoding features obtained in the previous step during backward propagation, represents the encoding and decoding features obtained in the previous step during forward propagation; the cross-attention principle involved in the hierarchical integration module is shown in the following formula:

[0038]

[0039] where T represents the scaling factor and transpose operation; Q is obtained by linearly mapping the encoding or decoding features of different scales, and K and V are obtained by linearly mapping the diffusion priors of different scales.

[0040] Furthermore, step S4 specifically includes:

[0041]

[0042] Among them, f on the left side of the formula i e3 represents the third-scale forward encoding feature enhanced by the memory encoder MemoryDecoder, and on the left side of the formula represents the third-scale backward encoding feature enhanced by the memory encoder MemoryDecoder; f on the right side of the formula i e3 represents the third-scale forward feature encoded by the encoder Encoder, and on the right side of the formula represents the third-scale backward feature encoded by the encoder Encoder; represents the forward deblurring memory feature retrieved from the memory, represents the backward deblurring memory feature retrieved from the memory, represents the maintained forward instantaneous perceptual memory, represents the maintained backward instantaneous perceptual memory.

[0043] Furthermore, during the memory retrieval process in step S4, the encoding feature of the last scale output by the encoder is used to construct the query key, and key projection is further used to reduce the channel dimension to reduce the computational overhead. Then, the query key is regarded as query q, and the memory feature m retrieved from the memory can be calculated through the readout operation, as shown in the following formula:

[0044] m = vW(k, q)

[0045] Among them, k and v are the memory key-value pairs stored in the memory bank; W(k, q) is the affinity matrix of size N×HW, representing the readout operation controlled by the key and query; the readout operation maps each query element to all memory elements and aggregates their values v accordingly.

[0046] Furthermore, step S5 specifically includes:

[0047] During decoding, similar to encoding, the multi-scale diffusion prior features are fused into the decoder again through the hierarchical integration module. In addition, the instantaneous perceptual memory maintained in the memory bank is updated using the gated recurrent unit GRU mechanism, and the involved formula is as follows:

[0048]

[0049] Among them, f i ej and respectively represent the encoded features enhanced by the memory decoder, and f i dj represents the multi-scale forward decoding features decoded by the decoder Decoder, represents the multi-scale backward decoding features decoded by the decoder Decoder; the cross-attention principle involved in the hierarchical integration module is shown in the following formula:

[0050]

[0051] Among them, T represents the scaling factor and transpose operation; Q is obtained by linearly mapping the encoded or decoded features of different scales, and K and V are obtained by linearly mapping the diffusion priors of different scales.

[0052] Furthermore, step S6 specifically includes:

[0053] Design a discriminative feature fusion module to preprocess relevant features before the cyclic feature extraction module; in the backward propagation process, the enhanced decoding features obtained are encoded by the value encoder Value Encoder to obtain the memory values of the backward memory bank, and at the same time, the previous query keys are reused as the memory keys of the backward memory bank. The involved formulas are as follows:

[0054]

[0055] Among them, x i represents the downsampled feature, and f i d1 on the right side of the formula represents the first-scale forward decoding features decoded by the decoder Decoder, and on the right side of the formula represents the first-scale backward decoding features decoded by the decoder Decoder; f i d1 and on the left side of the formula both represent the enhanced decoding features obtained, represents the features extracted by the previous cyclic feature extraction module in the forward propagation, represents the features extracted by the previous cyclic feature extraction module in the backward propagation; DFF represents the discriminative feature fusion module; RFE represents the cyclic feature extraction module, which has a cyclic structure, and the currently extracted features depend on the previously extracted features.

[0056] Furthermore, step S7 specifically includes:

[0057] During the forward propagation process, the final deblurred features are encoded by the Value Encoder to obtain the memory values of the forward memory bank. The query keys from the previous stage are also reused as the memory keys of the forward memory bank. The relevant formulas are as follows:

[0058]

[0059] Among them, I i and R i represent the blurred input frame I Blur and the finally restored deblurred frame I DB respectively. x i represents the downsampled features, and f i ej represents the multi-scale forward encoded features encoded by the previous encoder. represents the multi-scale backward encoded features encoded by the previous encoder; f i dj represents the enhanced multi-scale forward decoded features. represents the enhanced multi-scale backward decoded features; U represents the upsampling module, and TFR represents the target frame reconstruction module.

[0060] Furthermore, step S8 specifically includes:

[0061] Jointly train the entire network using the L1 Charbonnier loss and the L1 loss during the diffusion process, as shown in the following formula:

[0062]

[0063] L all = L char + L diff

[0064] Among them, in the Charbonnier loss, R t represents the restored deblurred frame, G t represents the corresponding real clear frame, and ε is a constant used to stabilize the training; in the diffusion loss, represents the prior features predicted by the diffusion network, and z represents the original prior features obtained by encoding the concatenation of the blurred input frame and the corresponding Ground Truth frame by the latent encoder.

[0065] Compared with the prior art, the present invention has the following advantages:

[0066] 1. A video deblurring method based on a memory diffusion network provided by the present invention improves the effect of the existing regression model on video deblurring by combining the diffusion priors generated by the diffusion network.

[0067] 2. The video deblurring method based on the memory diffusion network provided by the present invention constructs a concise and effective memory bank to screen and store temporal frame information, accurately capture the complex spatio-temporal local relationships in the blurred video to achieve finer-grained deblurring, and improves the existing work based on the memory network. A specific memory update algorithm is designed to refresh the memory bank, so as to better utilize the spatio-temporal information between consecutive video frames and ensure the temporal consistency between frames.

[0068] For the above reasons, the present invention can be widely promoted in the fields of image processing and the like. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0070] Figure 1 It is a flowchart of the method of the present invention.

[0071] Figure 2 It is the memory bank provided by the embodiment of the present invention.

[0072] Figure 3 It is the perceptual memory provided by the embodiment of the present invention.

[0073] Figure 4 It is the distinguishable feature fusion module provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0075] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0076] The present invention proposes a video deblurring method based on a memory diffusion network, and further explores the potential of the diffusion prior generated by combining the existing regression model with the diffusion network for video deblurring. As Figure 1 shown, its core includes two branches: a deblurring branch and a diffusion branch. The deblurring branch has a bidirectional cyclic structure and maintains a concise and effective memory bank containing a specific update mechanism, which can effectively utilize the spatio-temporal relationship between video frames and further ensure the temporal consistency between frames in the final deblurred video. The diffusion branch contains a lightweight diffusion network to generate multi-scale diffusion prior features.

[0077] The overall process is as follows:

[0078] S1. Select a publicly available video deblurring dataset as the ground truth, preprocess the input blurred video frames using a residual dense block (RDB), and downsample to obtain the initial downsampled feature x i ;

[0079] S2. Concatenate the blurred input frame and the corresponding Ground Truth frame, generate random noise through a diffusion forward denoising process after encoding by a latent encoder (LatentEncoder), and then perform reverse denoising by a diffusion denoising network to generate multi-scale diffusion prior features

[0080] S3. Send the downsampled feature x obtained in step S1 i to an encoder (Encoder) for encoding to obtain multi-scale encoded features; during encoding, fuse the multi-scale diffusion prior features generated in step S2 into the encoder through a hierarchical integration module (HIM);

[0081] S4. Before decoding, use the multi-scale encoded features obtained in step S3 to construct a query key (Key) to generate a Query for retrieving the memory features of the memory bank, and send the retrieved memory features, the maintained instantaneous perceptual memory, and the encoded features of the third scale to the Memory Decoder to enhance the previous encoded features;

[0082] S5. Input the enhanced encoded features in step S4 into the Decoder to decode and obtain multi-scale decoded features;

[0083] S6. Input the decoded features of the first scale obtained, the features extracted by the previous cycle feature extraction module, and the initial downsampled features into the Recurrent Feature Extractor (RFE) to further enhance the decoded features;

[0084] S7. Input the previous encoded features, the enhanced decoded features, and the initial downsampled features into the Target Frame Reconstruction module (TFR) to reconstruct the target frame, and then upsample through transposed convolution and fuse with the original blurred input to obtain the final deblurred result;

[0085] S8. Use the L1 Charbonnier loss and the L1 loss during the diffusion process to jointly train the entire network.

[0086] Specifically, as a preferred implementation manner of the present invention, in step S1, the initially obtained downsampled feature x i , is as follows:

[0087] x i = RDB(I i ), i = 1, 2,..., N

[0088] where, x i represents the downsampled feature, I i represents the blurred input frame I Blur , and RDB represents the residual dense block for downsampling.

[0089] Specifically, as a preferred implementation manner of the present invention, step S2 specifically includes:

[0090] S21. During the forward diffusion process, transform z into Gaussian noise through T iterations as shown in the following formula:

[0091]

[0092] where, t = 1, 2,..., T; β t ∈(0, 1) represents the hyperparameter for controlling the noise variance; N represents the Gaussian distribution; I represents the standard normal distribution;

[0093] S22. Let α t = 1 - β t , Through reparameterized iterative derivation, rewrite the formula in step S21 as follows:

[0094]

[0095] where z 0 = z;

[0096] S23. In the reverse denoising process, DM samples a Gaussian random noise z T , and then gradually denoises z T back to z 0 , as shown in the following formula:

[0097]

[0098] where,

[0099] S24. Replace z with the formula 0 to obtain:

[0100]

[0101] Remove the variance estimation in to obtain:

[0102]

[0103] where the noise ε is the only uncertain variable.

[0104] Following previous work, this embodiment uses a neural network to predict the noise ε given z t , c, t. In addition, following the general process of the diffusion process, during test inference, a randomly generated noise is directly input into the diffusion denoising network to generate multi-scale diffusion prior features. In this embodiment, as shown in Figure 1 (a), because the inherent randomness of the diffusion model may damage the spatial fidelity and even generate some additional visual content, the corresponding blurred input frame is encoded by the latent encoder and used as the condition c to be input into the diffusion denoising network to constrain the generated prior features.

[0105] Specifically in implementation, as a preferred implementation manner of the present invention, step S3 specifically includes:

[0106]

[0107] where f iej Represents the multi-scale forward encoding features encoded by the Encoder. Represents the multi-scale backward encoding features encoded by the Encoder, x i Represents the downsampled features. Represents the multi-scale diffusion prior features generated by the diffusion network. j = 1, 2, 3 represent 3 different scales. Represents the encoding and decoding features obtained in the previous step during backpropagation. Represents the encoding and decoding features obtained in the previous step during forward propagation. The cross-attention principle involved in the hierarchical integration module is shown as follows:

[0108]

[0109] Among them, T represents the scaling factor and transpose operation; Q is obtained by linearly mapping the encoding or decoding features of different scales, and K, V are obtained by linearly mapping the diffusion priors of different scales.

[0110] In specific implementation, as a preferred implementation manner of the present invention, in step S4, as Figure 2 shown, before decoding, the obtained encoding features are used to construct the query key Key to generate Query for memory feature retrieval in the memory bank, and then the retrieved memory features, the maintained instantaneous perceptual memory, and the encoding features of the third scale are sent together to the MemoryDecoder to enhance the previous encoding features. The formulas involved are as follows:

[0111]

[0112] Among them, f on the left side of the formula i e3 Represents the enhanced third-scale forward encoding features by the MemoryDecoder, and on the left side of the formula represents the enhanced third-scale backward encoding features by the MemoryDecoder; f on the right side of the formula i e3 Represents the third-scale forward features encoded by the Encoder, and on the right side of the formula represents the third-scale backward features encoded by the Encoder; Represents the forward deblurring memory features retrieved from the memory, Represents the backward deblurring memory features retrieved from the memory, Represents the maintained forward instantaneous perceptual memory, Represents the maintained backward instantaneous perceptual memory. In this embodiment, the MemoryDecoder is composed of residual blocks and CBAM blocks.

[0113] In specific implementation, as a preferred implementation manner of the present invention, during the memory retrieval process in step S4, the encoding feature of the last scale output by the encoder is used to construct the query key, and further, key projection (which consists of a convolution) is used to reduce the channel dimension to reduce the computational overhead. Then, the query key is regarded as query q, and the memory feature m retrieved from the memory can be calculated through a readout operation, as shown in the following formula:

[0114] m = vW(k,q)

[0115] where k and v are the memory key-value pairs stored in the memory bank; W(k,q) is an affinity matrix of size N×HW, representing the readout operation controlled by the key and the query; the readout operation maps each query element to all memory elements and aggregates their values v accordingly. In fact, this entire operation can be regarded as a spatio-temporal attention, and W can be regarded as an attention map because we are looking for the position most relevant to restoring the current query position on the memory frames to utilize its additional useful information.

[0116] In specific implementation, as a preferred implementation manner of the present invention, step S5 specifically includes:

[0117] During decoding, similar to encoding, the multi-scale diffusion prior features are fused into the decoder through the hierarchical integration module (HIM) again. In addition, the instantaneous perceptual memory maintained by the memory bank is updated using the gated recurrent unit GRU mechanism, as Figure 3 shown (for simplicity, the HIM block is omitted), and the involved formula is as follows:

[0118]

[0119] where f i ej and respectively represent the encoding features enhanced by the memory decoder, f i dj represents the multi-scale forward decoding features decoded by the decoder Decoder, represents the multi-scale backward decoding features decoded by the decoder Decoder; the cross-attention principle involved in the hierarchical integration module is as shown in the following formula:

[0120]

[0121] where, T represents the scaling factor and transpose operation; Q is obtained by linearly mapping the encoding or decoding features of different scales, and K, V are obtained by linearly mapping the diffusion priors of different scales.

[0122] In specific implementation, as a preferred implementation manner of the present invention, step S6 specifically includes:

[0123] Design a discriminative feature fusion module (DFF) to preprocess relevant features before the recurrent feature extraction module (RFE). During the backpropagation process, the obtained enhanced decoded features are encoded by the value encoder Value Encoder to obtain the memory values of the backward memory bank, and at the same time, the previous query keys are reused as the memory keys of the backward memory bank. The involved formulas are as follows:

[0124]

[0125] Among them, x i represents the downsampled feature, f on the right side of the formula i d1 represents the first-scale forward decoded feature decoded by the decoder Decoder, and on the right side of the formula represents the first-scale backward decoded feature decoded by the decoder Decoder; f on the left side of the formula i d1 and both represent the obtained enhanced decoded features, represents the feature extracted by the previous recurrent feature extraction module in the forward propagation, represents the feature extracted by the previous recurrent feature extraction module in the backpropagation; DFF represents the discriminative feature fusion module; RFE represents the recurrent feature extraction module, which has a recurrent structure, and the currently extracted feature depends on the previously extracted feature. In addition, in this embodiment, the previous downsampled feature x i-1 / x i+1 is input into the recurrent feature extraction module RFE because this has been proven to be effective. In this embodiment, the recurrent feature extraction module RFE is composed of a series of channel attention blocks (CAB). The structure of the discriminative feature fusion module DFF is as Figure 4 shown, and its core lies in further introducing a gated attention mechanism and dynamic filtering (Dynamic Filter) to strengthen feature fusion.

[0126] In specific implementation, as a preferred implementation manner of the present invention, step S7 specifically includes:

[0127] During the forward propagation process, the final deblurred feature is encoded by the value encoder Value Encoder to obtain the memory values of the forward memory bank, and the previous query keys are also reused as the memory keys of the forward memory bank. The involved formulas are as follows:

[0128]

[0129] Among them, I i and R i respectively represent the blurred input frame I Blur and the finally restored de-blurred frame I DB , x i represents the downsampled feature, and f i ej represents the multi-scale forward encoding feature encoded by the previous encoder, represents the multi-scale backward encoding feature encoded by the previous encoder; f i dj represents the enhanced multi-scale forward decoding feature, represents the enhanced multi-scale backward decoding feature; U represents the upsampling module, and TFR represents the target frame reconstruction module. In this embodiment, the target frame reconstruction module TFR is composed of a series of channel attention blocks (CAB). The value encoder is mainly composed of a residual block and the first layer of the pre-trained ResNet50 network, and is used for memory value encoding.

[0130] Specifically, as a preferred implementation manner of the present invention, the memory key-value pairs obtained in steps S6 and S7 are appended to the memory bank. For simplicity, in this embodiment, they are simply concatenated together. At the same time, in order to construct a concise and effective memory bank and avoid simply resetting and re-memorizing (discarding all previous memory frames) during training as in existing work and being unable to effectively utilize memory information, this embodiment explores a better memory forgetting mechanism. In order to process longer videos, inspired by the least-frequently-used (LFU) algorithm, this embodiment selects the ones with high usage rates for retention. Specifically, this embodiment introduces the usage amount "Usage", which is obtained by cumulative addition row by row based on the affinity matrix W and is normalized by the duration of each memory frame in the memory bank. Finally, the k frames with low usage amounts are filtered out through top-k, which also ensures the balance between training and inference to a certain extent because the same strategy is adopted.

[0131] Specifically, as a preferred implementation manner of the present invention, step S8 specifically includes:

[0132] The entire network is jointly trained using the L1 Charbonnier loss and the L1 loss during the diffusion process, as shown in the following formula:

[0133]

[0134] L all = L char + L diff

[0135] Among them, in the Charbonnier loss, R t represents the restored deblurred frame, G t represents the corresponding real clear frame, and ε is a constant used to stabilize the training; in the diffusion loss, represents the prior feature predicted by the diffusion network, and z represents the original prior feature obtained by encoding the concatenation of the blurred input frame and the corresponding Ground Truth frame by the Latent Encoder.

[0136] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video deblurring method based on memory diffusion network, characterized in that: include: S1. Select a public video deblurring dataset as the true value, use the residual dense block to preprocess the input blurred video frame, and downsample to obtain the initial downsampled feature x i ; S2: The blurred input frame and the corresponding ground truth frame are concatenated together, encoded by the latent encoder, and then subjected to a diffusion forward denoising process to generate random noise, which is then reversely denoised by the diffusion denoising network to generate multi-scale diffusion prior features. S3, the downsampled feature x obtained in step S1 i Send it to the feature encoder for encoding to obtain multi-scale encoding features; during encoding, the multi-scale diffusion prior features generated in step S2 are Fused into the encoder via a hierarchical integration module; S4, before decoding, construct a query key using the multi-scale coding features obtained in step S3 to generate a query to retrieve memory features from the memory bank, and send the retrieved memory features, the maintained instantaneous perceptual memory, and the coding features of the third scale to the memory decoder to enhance the previous coding features; S5, inputting the enhanced coding features in step S4 into the feature decoder for decoding to obtain multi-scale decoding features; S6, inputting the obtained decoding features of the first scale, the features extracted by the previous step cyclic feature extraction module and the initial down-sampling features into the cyclic feature extraction module to further enhance the decoding features; S7, inputting the previous encoding features, enhanced decoding features and initial down-sampling features into the target frame reconstruction module to reconstruct the target frame, and then up-sampling and fusing with the original blurred input through transposed convolution to obtain the final deblurring result; S8. Use L1 Charbonnier loss and L1 loss in the diffusion process to jointly train the entire network.

2. The video deblurring method based on memory diffusion network according to claim 1, characterized in that: In step S1, the initial downsampled feature x is obtained i ,as follows: x i =RDB(I i ),i=1,2,...,N Among them, x i represents the down-sampled feature, I i Represents the blurred input frame I Blur ,RDB represents the residual dense block for downsampling.

3. The video deblurring method based on memory diffusion network according to claim 1, characterized in that: Step S2 specifically includes: S21. In the forward diffusion process, z is transformed into Gaussian noise through T iterations, as shown in the following formula: Where, t=1,2,...,T; β t ∈(0,1) represents a hyperparameter used to control the noise variance; N represents Gaussian distribution; I represents standard normal distribution; S22, let α t =1-β t , Through iterative derivation of reparameterization, the formula in step S21 is rewritten as follows: Where z0=z; S23, in the reverse denoising process, DM samples a Gaussian random noise z T , and then gradually increase z T Denoising back to z0 is as follows: in, S24. Use formula Substituting z0 we get: Remove The variance in is estimated as: Among them, the noise ε is the only uncertain variable.

4. The video deblurring method based on memory diffusion network according to claim 1, characterized in that: Step S3 specifically includes: Among them, f i ej Represents the multi-scale forward encoding features encoded by the encoder. represents the multi-scale backward encoding features encoded by the encoder, x i represents the down-sampled features, Represents the multi-scale diffusion prior features generated by the diffusion network, j=1,2,3 represents three different scales, represents the encoding and decoding features obtained in the previous step of backward propagation, represents the encoding and decoding features obtained in the previous step of forward propagation; the cross-attention principle involved in the hierarchical integration module is shown in the following formula: in, T represents the scaling factor and transposition operation; Q is obtained by mapping the encoding or decoding features of different scales through a linear layer, and K and V are obtained by mapping the diffusion priors of different scales through a linear layer.

5. The video deblurring method based on memory diffusion network according to claim 1, characterized in that: Step S4 specifically includes: Among them, f on the left side of the formula i e3 represents the third-scale forward coding feature enhanced by the memory encoder MemoryDecoder. The left side of the formula represents the third-scale backward encoding feature enhanced by the memory encoder MemoryDecoder; f on the right side of the formula i e3 Represents the third-scale forward feature encoded by the encoder, and the right side of the formula Represents the third-scale backward features encoded by the encoder; represents the forward defuzzified memory features retrieved from memory, represents the backward defuzzified memory features retrieved from memory, represents the maintained forward transient perceptual memory, Represents maintained backward transient perceptual memory.

6. The video deblurring method based on memory diffusion network according to claim 1, characterized in that: In the memory retrieval process in step S4, the query key is constructed using the encoded features of the last scale output by the encoder, and key projection is further used to reduce the channel dimension to reduce the computational overhead. Then, the query key is regarded as query q, and the memory feature m retrieved from the memory can be calculated through the readout operation, as shown in the following formula: m=vW(k,q) Where k and v are memory key-value pairs stored in the memory bank; W(k,q) is an affinity matrix of size N×HW, representing the read operation controlled by key and query; the read operation maps each query element to all memory elements and aggregates their values ​​v accordingly.

7. The video deblurring method based on memory diffusion network according to claim 1, characterized in that: Step S5 specifically includes: During decoding, as with encoding, the multi-scale diffusion prior features are again fused into the decoder through a hierarchical integration module. In addition, the instantaneous perceptual memory maintained by the memory bank is updated using the gated recurrent unit GRU mechanism, where the formula involved is as follows: Among them, f i ej and They represent the encoding features enhanced by the memory decoder, f i dj Represents the multi-scale forward decoding features decoded by the decoder, Represents the multi-scale backward decoding features decoded by the decoder; the cross-attention principle involved in the hierarchical integration module is shown in the following formula: in, T represents the scaling factor and transposition operation; Q is obtained by mapping the encoding or decoding features of different scales through a linear layer, and K and V are obtained by mapping the diffusion priors of different scales through a linear layer.

8. The video deblurring method based on memory diffusion network according to claim 1, characterized in that: Step S6 specifically includes: A distinguishable feature fusion module is designed to preprocess relevant features before the cyclic feature extraction module. In the backward propagation process, the enhanced decoding features obtained are encoded by the value encoder to obtain the memory value of the backward memory bank, and the previous query key is reused as the memory key of the backward memory bank. The formula involved is as follows: Among them, x i Represents the downsampling feature, f on the right side of the formula i d1 Decoder represents the first scale forward decoding feature of the decoder. The right side of the formula Represents the first-scale backward decoding feature of the decoder; f on the left side of the formula i d1 and They all represent the enhanced decoding features obtained, represents the features extracted by the previous cycle feature extraction module obtained in the forward propagation, It represents the features extracted by the previous cyclic feature extraction module obtained in the backward propagation; DFF represents the distinguishable feature fusion module; RFE represents the cyclic feature extraction module, which has a cyclic structure, and the currently extracted features depend on the previously extracted features.

9. The video deblurring method based on memory diffusion network according to claim 1, characterized in that: Step S7 specifically includes: In the forward propagation process, the final defuzzified features are encoded by the value encoder to obtain the memory value of the forward memory bank. The previous query key is also reused as the memory key of the forward memory bank. The formula involved is as follows: Among them, I i and R i Represent the blurred input frame I Blur and the final restored deblurred frame I DB , x i represents the down-sampled feature, f i ej represents the multi-scale forward encoding features encoded by the previous encoder, represents the multi-scale backward encoding features encoded by the previous encoder; f i dj represents the enhanced multi-scale forward decoding features, represents the enhanced multi-scale backward decoding features; U represents the upsampling module, and TFR represents the target frame reconstruction module.

10. The video deblurring method based on memory diffusion network according to claim 1, characterized in that: Step S8 specifically includes: The entire network is jointly trained using the L1 Charbonnier loss and the L1 loss in the diffusion process, as shown in the following formula: L all =L char +L diff Among them, in Charbonnier loss, R t represents the restored deblurred frame, G t represents the corresponding real clear frame, ε is a constant used to stabilize the training; in the diffusion loss, represents the prior features predicted by the diffusion network, and z represents the original prior features obtained by concatenating the blurred input frame and the corresponding Ground Truth frame and encoding them through the latent encoder.

Citation Information

Patent Citations

  • Video defogging model based on memory network fused phase features and method thereof

    CN115471418A

  • Image deblurring method based on diffusion model

    CN116645287A

  • Medical image reconstruction method and system based on orthogonal texture perception memory bank

    CN117333570A

  • Deep network model, method, device and equipment for blind deblurring of remote sensing image and medium

    CN118411312A

  • Method and apparattus for generative model with arbitrary resolution and scale using diffusion model and implicit neural network

    KR102689642B1

Cited By

  • Video restoration method and device, electronic equipment and storage medium

    CN121563818A

  • Video restoration method and device, electronic equipment and storage medium

    CN121563818B

  • Real scene video deblurring system and method based on single-step video diffusion model

    CN122265096A