Video deblurring method based on position attention mechanism, computer equipment and storage medium
By introducing position attention mechanism and Transformer model in the video defuzzing method, the problem of difficulty in dealing with non-uniform blur in dynamic scenarios is solved in the existing technology, and more efficient image detail recovery and system robustness are achieved.
Patent Information
- Application Number
- CN202510150393.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video defuzzing methods are difficult to accurately identify and restore the details of non-uniform blur areas under dynamic scenes and complex motions, and traditional methods cannot effectively capture the blur characteristics and dynamic changes between different frames.
The video defuzzing method based on the position attention mechanism is used to extract spatiotemporal features through Position-Optimize Attention Model Block (POAB), and process it using the encoder-decoder network and the pre-trained Transformer model to generate clear video frames.
It significantly improves the ability to restore images details, enhances the robustness and adaptability of the system, and can effectively deal with complex blur problems, especially in dynamically changing blur scenes, improving the image de-blurry effect.
Smart Images

Figure CN120070254A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video deblurring, and more specifically, to a video deblurring method, a computer device, and a storage medium based on a position attention mechanism. Background Art
[0002] Video deblurring is an important research direction in computer vision. Its core task is to remove the blur caused by factors such as camera movement, object movement, and lens shake in a video, thereby restoring clear images. This problem is of great significance in many practical applications, especially in fields such as drone aerial photography, autonomous driving, dynamic surveillance, and augmented reality. In these applications, due to the movement of the device and the dynamic changes in the environment, videos often appear blurred. Blurred images not only affect the visual effect but may also affect subsequent computer vision tasks, such as object detection, object tracking, and pose estimation.
[0003] The video deblurring task faces multiple technical challenges, especially in dynamic scenes and complex motions. Traditional deblurring methods are mainly based on image reconstruction techniques. However, when the movement speed of the scene is relatively fast or the blur distribution is uneven, the deblurring effect is often unsatisfactory. In real-world video scenes, the blur is usually not evenly distributed. In some dynamic scenes, the degree of blur changes significantly with factors such as the movement speed of objects, the camera movement trajectory, and environmental illumination changes. Especially when there are multiple motion sources in a video frame, such as fast-moving objects and severe camera shakes, the blur presents a non-uniform distribution, resulting in a large difference in the degree of blur in different regions. This non-uniformity of the blur makes it difficult for existing deblurring methods to accurately identify and restore the details of each local region.
[0004] The blur in a video is not only a spatial problem but also involves dynamic changes in the time dimension. The blurred regions usually change between different frames, and these changes continuously occur as the video frame sequence progresses. Traditional deblurring methods often fail to effectively capture the blur characteristics and dynamic changes between different frames, resulting in the inability to fully utilize cross-frame information during the restoration process, thereby affecting the final deblurring effect. In complex dynamic scenes, the temporal relationship between different frames and the degree of blur in spatial positions need to be accurately modeled to effectively restore the clarity of the image. The accurate localization of the blurred region is a key issue in video deblurring. The edges of the blurred region are usually relatively blurred and difficult to identify. Especially under high-speed motion or extreme blur conditions, existing methods rely on complex image segmentation and detection algorithms to locate the blurred region, but these methods usually perform poorly in complex backgrounds, especially when the background is complex or the foreground overlaps with the background. Summary of the Invention
[0005] To solve the deficiencies of the prior art, the present invention provides a video deblurring method, a computer device, and a storage medium based on a position attention mechanism, which improve the removal effect of the blurred regions of the video and enhance the robustness and adaptability of the system.
[0006] The technical solution adopted by the present invention to solve the above technical problems is as follows: A video deblurring method based on a position attention mechanism, comprising the following steps:
[0007] Obtain consecutive original video frames
[0008] Use a pre - POAB (Position - Optimize Attention Model Block) to process the original video frames to extract spatio - temporal features;
[0009] Input the spatio - temporal features into an encoder - decoder network for detection to obtain a blurred region feature map. The encoder - decoder network includes multiple encoding layers and multiple decoding layers. The multiple encoding layers are used to process the spatio - temporal features in sequence, and the encoding layer includes a first POAB. The multiple decoding layers are used to process the output result of the last encoding layer in sequence and perform skip connections with the output results of the corresponding encoding layers to obtain a blurred region feature map, and the encoding layer includes a second POAB;
[0010] Use a post - POAB to process the blurred region feature map to obtain a fused feature map;
[0011] Input the fused feature map into a pre - trained Transformer model for processing and obtain a clear video frame.
[0012] As a further optimization of the video deblurring method based on a position attention mechanism of the invention: The method of using a pre - POAB to process the original video frames to extract spatio - temporal features includes:
[0013] Process the original video frames through a convolutional layer and an activation function to obtain processed data;
[0014] Adopt global average pooling and global max pooling to encode each channel of the processed data in the horizontal and vertical directions to obtain pooled data;
[0015] Concatenate the pooled data, and generate intermediate features through convolution and activation function processing;
[0016] Divide the intermediate features in the height and width directions to obtain segmented data;
[0017] Expand the segmented data to a tensor with the same number of channels as the original video frame through convolution and the Sigmod activation function, and generate weights;
[0018] Multiply the weights element-wise with the X of the original video frame to obtain spatio-temporal features.
[0019] As a further optimization of the invention of a video deblurring method based on a position attention mechanism: The method of encoding each channel of the processed data in the horizontal and vertical directions by using global average pooling and global maximum pooling includes:
[0020] Perform global average pooling and global maximum pooling on the processed data in the vertical direction. The outputs of the c-th channel at height h are respectively:
[0021]
[0022]
[0023] Perform global average pooling and global maximum pooling on the processed data in the horizontal direction. The outputs of the c-th channel at width w are respectively:
[0024]
[0025]
[0026] As a further optimization of the invention of a video deblurring method based on a position attention mechanism: The method of concatenating the pooled data and generating intermediate features through convolution and activation function processing includes:
[0027]
[0028] where [·,·] represents the concatenation operation along the spatial dimension, φ is the activation function, f is the intermediate feature, and there is
[0029] As a further optimization of the invention of a video deblurring method based on a position attention mechanism: The method of splitting the intermediate feature in the height and width directions to obtain segmented data includes:
[0030] [x h ,x w = split(f, [h, w], dim = 2),
[0031] where
[0032] As a further optimization of the invention of a video deblurring method based on a position attention mechanism: The segmentation data is processed through convolution and the Sigmod activation function and extended to a tensor with the same number of channels as the original video frame The method for obtaining the weight includes:
[0033] σ h =(Sigmoid(Conv(x h )))expand(-1,-1,h,w),
[0034] σ w =(Sigmoid(Conv(x w )))expand(-1,-1,h,w),
[0035] Wherein,
[0036] Multiply the weight and the X of the original video frame element-wise to obtain the spatio-temporal feature, and the method is:
[0037]
[0038] Wherein, represents the element-wise multiplication operation, Y represents the spatio-temporal feature,
[0039] As a further optimization of the invention of a video deblurring method based on a position attention mechanism: The pre-trained Transformer model includes a self-attention module. The method for inputting the fused feature map into the pre-trained Transformer model for processing and obtaining a clear video frame includes:
[0040] Transfer the fused feature map to the self-attention module of the j-th layer, and calculate the query vector, key vector, and value vector according to different weight matrices;
[0041] Based on the query vector and the key vector, calculate the attention weight matrix through the dot product operation;
[0042] Use the attention weight matrix to perform weighted aggregation on the value vector to generate a single-head attention output;
[0043] Concatenate and linearly transform multiple single-head attention outputs to obtain a multi-head self-attention output;
[0044] Fuse the multi-head self-attention output with the fused feature map through residual connection to obtain the first fused sub-feature, and use a two-layer feed-forward network to extract features from the first fused sub-feature to obtain the second fused sub-feature;
[0045] The first fused sub - feature and the second fused sub - feature are subjected to residual connection to obtain a clear video frame.
[0046] The technical solution adopted by the present invention to solve the above - mentioned technical problems is: a computer device, including:
[0047] A memory for storing a computer program;
[0048] A processor for reading and executing the computer program to implement the above - mentioned video de - blurring method based on a position attention mechanism.
[0049] The technical solution adopted by the present invention to solve the above - mentioned technical problems is: a storage medium for storing a computer program, and when the computer program is executed, it implements the above - mentioned video de - blurring method based on a position attention mechanism.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] 1) The present invention combines a position - enhanced attention mechanism and an optimization module based on the Transformer architecture. For non - uniform dynamic blur scenes, it significantly improves the ability of image detail restoration through refined processing. This solution can effectively handle complex blur problems, especially in dynamically changing blur scenes, significantly improving the image de - blurring effect and enhancing the robustness and adaptability of the system;
[0052] 2) The present invention optimizes the use of computing resources through lightweight design and group attention mechanism. By adopting the strategy of reducing the number of channels, it reduces the computational overhead of the network while maintaining high image restoration quality, effectively improving the inference speed and processing efficiency, and is particularly suitable for resource - constrained embedded devices and real - time processing tasks.
[0053] 3) The present invention introduces a multi - head self - attention mechanism, which can balance the capture of global and local information, improve the detail performance of image restoration. Through multi - level feature optimization and residual connection, it further improves the quality of image de - blurring, ensuring that the system can stably and efficiently execute tasks in complex and dynamically changing scenes, and enhancing the overall performance of the de - blurring system. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is the network structure diagram of the position - optimized attention module of the present invention;
[0055] Figure 2 is the network structure diagram of the encoder - decoder network embedding POAB of the present invention;
[0056] Figure 3 is the schematic diagram of the actual processing process. DETAILED DESCRIPTION OF THE INVENTION
[0057] The technical solution of the present invention will be further elaborated in detail below in conjunction with specific embodiments. For parts that are not detailedly recorded and disclosed in the following embodiments of the present invention, they should all be understood as the prior art known or should be known to those skilled in the art, such as how the Transformer model is pre-trained, etc.
[0058] A video deblurring method based on a position attention mechanism, as Figures 1 to 3 shown, includes the following steps:
[0059] Obtain consecutive original video frames
[0060] Use the pre-positioned POAB (Position-Optimize Attention Model Block) to process the original video frames to extract spatio-temporal features, perform multi-scale feature modeling and image restoration;
[0061] The method of using the pre-positioned POAB to process the original video frames to extract spatio-temporal features includes:
[0062] Process the original video frames through a convolutional layer and an activation function to obtain processed data;
[0063] Adopt global average pooling and global max pooling to encode each channel of the processed data in the horizontal and vertical directions to obtain pooled data;
[0064] The method of using global average pooling and global max pooling to encode each channel of the processed data in the horizontal and vertical directions includes:
[0065] Perform global average pooling and global max pooling on the processed data in the vertical direction. The outputs of the c-th channel at height h are respectively:
[0066]
[0067]
[0068] Perform global average pooling and global max pooling on the processed data in the horizontal direction. The outputs of the c-th channel at width w are respectively:
[0069]
[0070]
[0071] The pooled data are spliced, and the intermediate features are generated through convolution and activation function processing, where the intermediate features are position-aware intermediate features; the method of splicing the pooled data and generating the intermediate features through convolution and activation function processing includes:
[0072]
[0073] Among them, [·,·] represents the concatenation operation along the spatial dimension, φ is the activation function, f is the intermediate feature, and there is
[0074] The intermediate feature is segmented in the height and width directions to obtain segmentation data; the method of segmenting the intermediate feature in the height and width directions to obtain segmentation data includes:
[0075] [x h ,x w ]=split(f,[h,w],dim=2);
[0076] in,
[0077] The segmented data is expanded to the original video frame through convolution and Sigmod activation function The tensor has the same number of channels and generates weights; the segmented data is expanded to the same size as the original video frame through convolution and Sigmod activation function. For tensors with the same number of channels, the methods for obtaining weights include:
[0078] σ h =(Sigmoid(Conv(x h )))expand(-1,-1,h,w);
[0079] σ w =(Sigmoid(Conv(x w )))expand(-1,-1,h,w);
[0080] in,
[0081] The weights are compared with the original video frame Multiply X element by element to get the spatiotemporal features; the weights are combined with the original video frame The method of multiplying X element by element to obtain the spatiotemporal features is:
[0082]
[0083] in, represents the element-by-element multiplication operation, Y represents the spatiotemporal features,
[0084] The spatio-temporal features are input into an encoder-decoder network for detection to obtain a blurred region feature map. The encoder-decoder network includes multiple encoding layers and multiple decoding layers, at least two encoding layers, and at least two decoding layers. The multiple encoding layers are used to process the spatio-temporal features sequentially, and an encoding layer includes a first POAB. Each encoding layer processes and downsamples the spatio-temporal features through the first POAB to obtain first encoded data. The next encoding layer then processes the first data of the previous layer through the first POAB of this layer and downsamples it to obtain second encoded data, and the subsequent encoding layers repeat this operation; the multiple decoding layers are used to process the output result of the last encoding layer sequentially and perform skip connections with the output results of the corresponding encoding layers to obtain a blurred region feature map. A decoding layer includes a second POAB. After the decoding layer upsamples the output result of the last encoding layer and performs a skip connection with the output result of the corresponding encoding layer, it is processed through the second POAB of this layer. The next decoding layer repeats this operation based on the output of the previous layer until the last decoding layer. The specific method is as follows:
[0085] Let the output of the encoding layer at the i-th layer be Then the operation of the encoder can be expressed as:
[0086]
[0087] where the initial input feature ε i (·) represents the downsampling operation of the i-th layer of the encoder, and F POAB is the feature map output by the first POAB. The encoding layer gradually extracts spatio-temporal features and reduces the spatial resolution through consecutive convolution operations and pooling operations. The size of the downsampled feature map is: C i = f(i),
[0088] During the recovery process of the decoding layer, it not only depends on the output of the previous layer of the decoding layer but also directly fuses the features of the corresponding encoding layer into the decoding layer through skip connections to retain high-resolution detail information;
[0089] The output of the l-th layer of the decoding layer can be defined as:
[0090]
[0091] where σ l (·) represents the upsampling operation of the l-th layer of the decoding layer. After the decoding layer is gradually refined and recovered, it outputs a feature map with the same resolution as the original: Through the above jump connection mechanism, the decoding layer can make full use of the detailed features extracted by the encoding layer during the recovery process to improve the accuracy of image recovery; the post-POAB is used to process the feature map of the blurred area to obtain a fused feature map; the post-POAB is the l-th decoding layer, and the fused feature map is represented by representation.
[0092] Through the above encoder-decoder network, multi-scale feature modeling, and jump connection fusion, the method realizes the precise compensation of complex motion blurred areas, significantly improving the recovery accuracy and detail consistency.
[0093] The fused feature map is input into a pre-trained Transformer model for processing to obtain a clear video frame.
[0094] The pre-trained Transformer model includes a self-attention module. The method of inputting the fused feature map into the pre-trained Transformer model for processing to obtain a clear video frame includes:
[0095] The fused feature map is passed to the self-attention module of the j-th layer, and query vectors, key vectors, and value vectors are calculated according to different weight matrices; the query vector is represented by Query, the key vector is represented by Key, and the value vector is represented by Value.
[0096] The formula is as follows:
[0097]
[0098]
[0099] Among them, and are the weight matrices of the corresponding layers, which are used to extract query vectors, key vectors, and value vectors from the fused feature map respectively. The query vector represents the key area of the target feature, and the key vector and value vector are used to construct feature matching and representation respectively.
[0100] Based on the query vector and the key vector, the attention weight matrix is calculated through a dot product operation; the formula is as follows:
[0101]
[0102] where d k is the dimension of the key vector;
[0103] Using the attention weight matrix, the value vectors are weighted and aggregated to generate a single-head attention output; the formula is as follows:
[0104]
[0105] Concatenate and linearly transform multiple single-head attention outputs to obtain the multi-head self-attention output; the formula is as follows:
[0106]
[0107] where h is the number of attention heads, is the linear transformation matrix for each layer;
[0108] Fuse the multi-head self-attention output with the fused feature map through residual connection to obtain the first fused sub-feature; the formula is as follows:
[0109]
[0110] Extract features from the first fused sub-feature using a two-layer feed-forward network to obtain the second fused sub-feature; the formula is as follows:
[0111]
[0112] where W 1 , W 2 and b 1 , b 2 are the weights and biases of the fully connected layer;
[0113] Obtain the clear video frame by passing the first fused sub-feature and the second fused sub-feature through a residual connection; the formula is as follows:
[0114]
[0115] To reduce the computational cost, a factor r for reducing the number of channels and a grouped attention mechanism are adopted. After reducing the number of feature channels from C to C / r, attention calculation is performed:
[0116]
[0117]
[0118]
[0119] where, are the channel compression weight matrices for the query vector, key vector, and value vector respectively, and the dimensions are all Q t , K r , V r become N×(C / r) in dimension, where N is the number of spatial positions in the feature map;
[0120] Finally, obtain the clear video frame through the mapping matrix for restoring the number of channels:
[0121]
[0122] The present invention combines a position-enhanced attention mechanism with a lightweight optimization module based on the Transformer architecture, effectively improving the network's ability to recover complex dynamic blurs while reducing the number of parameters and computational costs. On the premise of ensuring high-precision image restoration, this method further optimizes the inference efficiency, has good real-time performance and robustness, and is more suitable for video de-blurring scenarios that require efficient processing, such as multimedia communication and mobile devices.
[0123] A computer device, comprising:
[0124] A memory for storing a computer program;
[0125] A processor for reading and executing the computer program to implement the above-mentioned video de-blurring method based on a position attention mechanism.
[0126] A storage medium for storing a computer program, which when executed implements the above-mentioned video de-blurring method based on a position attention mechanism.
[0127] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A video deblurring method based on position attention mechanism, characterized in that: The steps include: Get continuous raw video frames Use the pre-POAB (Position-Optimize AttentionModel Block) to optimize the original video frame Processing is performed to extract spatiotemporal features; The spatiotemporal features are input into an encoder-decoder network for detection to obtain a fuzzy area feature map, the encoder-decoder network includes multiple encoding layers and multiple decoding layers, the multiple encoding layers are used to process the spatiotemporal features in sequence, and the encoding layer includes a first POAB, the multiple decoding layers are used to process the output result of the last encoding layer in sequence and perform jump connection with the output result of the corresponding encoding layer to obtain the fuzzy area feature map, and the encoding layer includes a second POAB; The fuzzy area feature map is processed by post-POAB to obtain a fused feature map; The fused feature map is input into the pre-trained Transformer model for processing to obtain a clear video frame.
2. The video deblurring method based on position attention mechanism as claimed in claim 1, characterized in that: Use pre-POAB to process the original video frame Methods for processing and extracting spatiotemporal features include: The original video frame is processed through convolutional layers and activation functions. Processing is performed to obtain processed data; Global average pooling and global maximum pooling are used to encode each channel of the processed data in the horizontal and vertical directions to obtain pooled data; The pooled data is concatenated and processed by convolution and activation functions to generate intermediate features; Segment the intermediate features in the height and width directions to obtain segmentation data; The segmented data is expanded to the original video frame through convolution and Sigmod activation function. Tensors with the same number of channels and weights are generated; The weights are compared with the original video frame Multiply X element by element to get the spatiotemporal features.
3. The video deblurring method based on position attention mechanism as claimed in claim 2, characterized in that: Methods for encoding each channel in the horizontal and vertical directions of the processed data using global average pooling and global maximum pooling include: Perform global average pooling and global maximum pooling on the processed data in the vertical direction, and the outputs of the cth channel at height h are: Perform global average pooling and global maximum pooling on the processed data in the horizontal direction, and the outputs of the cth channel at width w are:
4. The video deblurring method based on position attention mechanism as claimed in claim 2, characterized in that: The methods of concatenating the pooled data and generating intermediate features through convolution and activation functions include: Among them, [··] represents the concatenation operation along the spatial dimension, φ is the activation function, f is the intermediate feature, and there is 5. The video deblurring method based on position attention mechanism as claimed in claim 2, characterized in that: Methods for segmenting the intermediate features in height and width directions to obtain segmentation data include: [x h ,x w ]=split(f,[h,w],dim=2), in, 6. The video deblurring method based on position attention mechanism as claimed in claim 2, characterized in that: The segmented data is expanded to the original video frame through convolution and Sigmod activation function. For tensors with the same number of channels, the methods for obtaining weights include: σ h =(Sigmoid(Conv(x h )))expand(-1,-1,h,w), σ w =(Sigmoid(Conv(x w )))expand(-1,-1,h,w), in, The weights are compared with the original video frame The method of multiplying X element by element to obtain the spatiotemporal features is: in, represents the element-by-element multiplication operation, Y represents the spatiotemporal features, 7. The video deblurring method based on position attention mechanism as claimed in claim 1, characterized in that: The pre-trained Transformer model includes a self-attention module. The method of inputting the fused feature map into the pre-trained Transformer model for processing and obtaining a clear video frame includes: The fused feature map is passed to the self-attention module of the jth layer, and the query vector, key vector and value vector are calculated according to different weight matrices; Based on the query vector and the key vector, the attention weight matrix is calculated through the dot product operation; Using the attention weight matrix, the value vector is weightedly aggregated to generate a single-head attention output; Multiple single-head attention outputs are concatenated and linearly transformed to obtain multi-head self-attention outputs; The multi-head self-attention output is fused with the fusion feature map through residual connection to obtain the first fusion sub-feature, and the first fusion sub-feature is extracted using a two-layer feedforward network to obtain the second fusion sub-feature; The first fused sub-feature and the second fused sub-feature are connected through residual to obtain a clear video frame.
8. Computer device, characterized in that include: Memory for storing computer programs; A processor, configured to read and execute the computer program to implement a video deblurring method based on a position attention mechanism as described in any one of claims 1-8.
9. A storage medium, characterized in that Used to store a computer program, which, when executed, implements a video deblurring method based on a position attention mechanism as described in any one of claims 1-8.