A Video Super-Resolution Model and Method Combining Multiple Attention Mechanisms with Optical Flow

By using multiple attention combined with optical flow methods in the video super-resolution model, using two-stage feature alignment and deformable convolutional LSTM to process motion information, the high computational cost and jitter problems of video super-resolution method are solved, and efficient video super-resolution and timing consistency are achieved.

CN112734644BActive Publication Date: 2025-05-27ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110067283.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-19
Publication Date
2025-05-27
Estimated Expiration
2041-01-19

AI Technical Summary

Technical Problem

The existing video super-resolution methods have problems such as high computational costs or jitter prone to video recovery.

Method used

The video super-resolution model with multiple attention combined with optical flow is adopted. Through the idea of ​​dual-stage feature alignment, the optical flow network is used to process small motion information, and the deformable convolutional LSTM is used to process large motion information, reducing the deviation between the target frame and the reference frame, and ensuring the consistency of the video timing.

Benefits of technology

It effectively reduces the calculation cost, reduces the jitter phenomenon after video recovery, ensures the timing consistency of the video, makes full use of all layered feature information, and enhances the dependence and adaptability of the channel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112734644B_ABST
    Figure CN112734644B_ABST
Patent Text Reader

Abstract

A video super-resolution model and method combining multiple attentions with optical flow provided by the present invention belong to the technical fields of pattern recognition and computer vision. The model of the present invention includes a feature extraction part, a feature processing part, a deformable convolution part, and a video reconstruction part. The method of the present invention uses a two-stage idea to perform feature alignment on micro-motions and large motions respectively, processes the information of micro-motions and large motions separately, reduces the deviation between the target frame and the reference frame, makes full use of the feature information of all layers, uses multiple attentions to make the video spatial information not easily lost, retains the spatial information, enhances the channel dependence and self-adaptability, and can capture long-range dependencies to achieve global learning. And a deformable convolutional long short-term memory network (DLSTM) is used for video frame fusion, preventing phenomena such as jitter and flicker artifacts from appearing in the restored video and ensuring the consistency of the video time sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pattern recognition and computer vision, and more specifically, to a video super-resolution model and method that combines multiple attentions with optical flow. Background Art

[0002] At present, deep learning methods based on convolutional neural networks are widely used in the field of computer vision. The super-resolution technology in low-level vision has always been a very challenging and popular computer vision task. According to the data type classification, the current super-resolution work is divided into image super-resolution and video super-resolution. The differences between video super-resolution and image super-resolution mainly include two points, including: video frame alignment and video frame fusion. Among them, video frame alignment is because there is various motion information in the video, so there is a deviation between the reference frame and the target frame. In super-resolution, it is generally necessary to use adjacent frames and the reference frame for alignment. And there are problems of motion blur and scene switching in the video. Effective fusion of video frames can remove interference information.

[0003] For the above two points, the existing methods are as follows: one is to use three-dimensional convolution, directly using the function of 3D convolution to capture temporal features and directly perform inter-frame fusion; the second is to use a cyclic structure to extract inter-frame relationships and fuse the information of the target frame and the reference frame; the third is to use the inter-frame information after fusion to predict the filter parameters, and then perform super-resolution through filtering to obtain an adaptive filtering effect. And the overall framework of the current video super-resolution generally has two ideas. One is to use three-dimensional convolution, but using three-dimensional convolution will increase more parameters due to the introduction of an additional dimension, resulting in an increase in computational cost. The second is to process the video into one frame of image at a time, and then process it according to the method of image super-resolution. In this way, it is difficult to maintain the temporal consistency of the video, and the restored video is prone to jitter.

[0004] After retrieval, the Chinese patent application number is: ZL201911203785.1, the application date is: November 29, 2019, and the invention name is: A video super-resolution reconstruction method based on a deep dual attention network. This application case realizes accurate video super-resolution reconstruction by loading a cascaded motion compensation network model and a reconstruction network model, and fully utilizes spatio-temporal information features; the motion compensation network model therein can gradually learn the optical flow representation from rough to detailed to synthesize multi-scale motion information of adjacent frames; in the reconstruction network model, a dual attention mechanism is used, and a residual attention unit is formed to focus on intermediate information features, which can better restore image details. However, this application case still processes the video into one frame of image at a time for super-resolution processing, and the restored video still has jitter phenomenon. Summary of the Invention

[0005] 1. Technical Problems to be Solved by the Invention

[0006] In view of the problems that existing video super-resolution methods have high computational costs or are prone to jitter after video restoration, the present invention provides a video super-resolution model and method that combines multiple attention mechanisms with optical flow. Using the idea of two-stage feature alignment, an optical flow network is used to process tiny motion information, and a deformable convolutional LSTM is used to process large motion information. While reducing the deviation between the target frame and the reference frame, the temporal consistency of the restored video is ensured.

[0007] 2. Technical Solution

[0008] To achieve the above objectives, the technical solution provided by the present invention is as follows:

[0009] A video super-resolution model that combines multiple attention mechanisms with optical flow according to the present invention includes a feature extraction part, a feature processing part, a deformable convolutional part, and a video reconstruction part; video frames pass through the four parts in sequence to achieve super-resolution; the feature processing part includes a multi-attention branch and an attention optical flow estimation branch; the multi-attention branch includes a spatial attention module, a self-attention module, a convolutional module, and an upsampling module; the attention optical flow estimation branch includes a spatial attention module, a channel attention module, an optical flow estimation network module, a convolutional module, and an upsampling module.

[0010] Furthermore, the modules in the multi-attention branch are, in the order of video frames passing through, a spatial attention module - a self-attention module - a convolutional module - an upsampling module - a convolutional module; the module order of the attention optical flow estimation branch is that video frames pass through the spatial attention module and the channel attention module simultaneously, and then enter the optical flow estimation network module - a convolutional module - an upsampling module - a convolutional module.

[0011] Furthermore, the feature extraction part includes two convolutional modules and three residual dense modules. The input video frames pass through the two convolutional modules and the three residual dense modules in sequence, and then enter the feature processing module; the video reconstruction module is a convolutional module.

[0012] A video super-resolution method that combines multiple attention mechanisms with optical flow using the above model according to the present invention includes the following steps:

[0013] Step 1: Input 2n + 1 consecutive low-resolution video frames;

[0014] Step 2: Input the video frames into the feature extraction part of the model to extract video frame features F;

[0015] Step 3: Send the extracted features F into the multi-attention branch and the attention optical flow estimation branch respectively, and two branch outputs F a and F f ;

[0016] Step 4. After upsampling F a and F f , input them into the deformable convolutional network DLSTM and a convolutional model to obtain the video super-resolution feature F d .

[0017] Furthermore, in the above Step 1, the input video frames are 2n + 1 LR frames in MAFnet, and their sequence is The input size of MAFnet is (M L × N L ), where is the output HR frame, denoted as I SR , with a size of (M H × N H ), and M H > M L , N H > N L .

[0018] Furthermore, in the above Step 2, the video frames are subjected to two convolutional operations and residual dense block operations to obtain the feature F

[0019] F = H rdb (H c1 (H c0 (I LR ))) (1)

[0020] where I LR represents the input low-resolution frame, H rdb (·) represents the residual dense block operation, and H c (·) represents the convolutional operation.

[0021] Furthermore, in the above Step 3, the extracted feature F is fed into the multi-attention branch and the attention optical flow estimation branch to obtain the output of the multi-attention branch, Equation (2), and the output of the attention optical flow estimation branch, Equation (3),

[0022] F a = H se (H sa (H ca (F))) (2)

[0023] F f = H f (H sa (F), H ca (F)) (3)

[0024] where H se (·) is the self-attention module function, H sa (·) is the spatial attention module function, and H ca(·) is the channel attention module function, H f (·) is the optical flow module function.

[0025] Furthermore, in the channel attention, after the features are input, they are respectively subjected to adaptive average pooling and adaptive max pooling, then the number of channels is reduced through convolution and passed through the activation function ReLU, and then the number of channels is restored through convolution; the two obtained features are added and passed through the Sigmoid function to obtain the attention feature map, and then the attention feature map and the input features are multiplied matrix-wise to obtain the output features;

[0026] In the spatial attention, after the features are input, they first pass through convolution and the activation function LReLU, then pass through a pooling layer composed of average pooling, max pooling and concatenation operations, then pass through convolution and LReLU to obtain Feature 1. After that, through the repeated convolution, LReLU and pooling layer structure, and through two convolution and LReLU structures, and then interpolation operation to obtain Feature 2. After adding Feature 1 and Feature 2, through convolution, LReLU and interpolation operation, the features are successively sent into two convolution and LReLU to obtain Feature 3, and the attention feature map is obtained by using the Sigmoid function. The attention feature map and the input features are multiplied matrix-wise, and the result is added to Feature 3 to obtain the output features;

[0027] In the self-attention, after the features are input, they pass through three convolutional channels respectively to obtain Feature 1, Feature 2 and Feature 3. Feature 1 and Feature 2 are multiplied matrix-wise and passed through the softmax function to obtain the attention map, and then multiplied matrix-wise with Feature 3 to obtain the output features.

[0028] Furthermore, in the optical flow estimation network, given any two adjacent frames I i , I i+1 , then the optical flow calculation formula can be expressed as

[0029] f i→i+1 = N f (I i , I i+1 ) (4)

[0030] where, N f represents the optical flow estimation network.

[0031] Furthermore, in the fourth step, upsampling is performed on F a and F f y

[0032] y 1 = H c4 (↑(H c2 (F a ))) (5)

[0033] y2 = H c5 (↑(H c3 (F f ))) (6)

[0034] Among them, ↑ represents upsampling; y 1 , y 2 are fed into the DLSTM, and then the final output is obtained through one layer of convolution

[0035] F d = H c6 (DLSTM(y 1 , y 2 ) (7)

[0036] Among them, F d represents the features obtained through the DLSTM and the final reconstruction convolution; the entire network is finally expressed as

[0037] I SR = H MAFnet (I LR ) (8).

[0038] 3. Beneficial effects

[0039] Adopting the technical solution provided by the present invention, compared with the existing well-known technologies, the following remarkable effects are achieved:

[0040] (1) In view of the problems that the existing video super-resolution methods have high computational costs or are prone to jitter after video restoration, a video super-resolution method combining multiple attentions with optical flow of the present invention provides an idea of two-stage feature alignment, processes the information of small motions and large motions respectively, reduces the deviation between the target frame and the reference frame, makes full use of the feature information of all layers, makes the video spatial information not easily lost by using multiple attentions, retains the spatial information, enhances the channel dependence and self-adaptability, and can capture long-range dependencies to achieve global learning.

[0041] (2) A video super-resolution method combining multiple attentions with optical flow of the present invention uses an optical flow network for the first-stage feature alignment to process small motion information, and uses an LSTM with deformable convolution added to process large motion information, improves the resolution ability, reduces the jitter phenomenon, and ensures the consistency of video timing.

[0042] (3) A video super-resolution model that combines multiple attentions with optical flow in the present invention. In the spatial attention module, LReLU is selected as the activation function, which alleviates the problem of neuron death during training, better preserves spatial information, and solves the problem that ReLU causes neuron death during training and cannot further update the parameter gradient. In the model, deformable convolution is added to the traditional LSTM, which can adjust the displacement of spatial position information, retain the original advantages of LSTM, and at the same time enhance the ability of video frames to align in time series, effectively utilize context information to process large motion information in the video, and ensure the continuity of the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is the overall process block diagram of the model of the present invention;

[0044] Figure 2 It is the structural diagram of the channel attention model in the present invention;

[0045] Figure 3 It is the structural diagram of the spatial attention model in the present invention;

[0046] Figure 4 It is the structural diagram of the self-attention model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To further understand the content of the present invention, the present invention will be described in detail in combination with the drawings and embodiments.

[0048] In the prior art, traditional video super-resolution methods use 3D convolution to extract spatial information to retain the spatial features of the video. However, once 3D convolution is introduced, it means that a new dimension is introduced, which will not only bring more parameters, increase the computational cost, but also limit the depth of the network and affect the super-resolution performance. Some other solutions choose to process the video frame by frame and then perform super-resolution according to the image super-resolution method. However, this method is difficult to ensure the coherence of the video, especially for videos with large motion amplitudes, and local features and global dependencies cannot be well integrated. In addition, a recurrent neural network can be selected to maintain the coherence of the video, but this method has poor effect in retaining spatial information. A video super-resolution model and method that combines multiple attentions with optical flow in the present invention provides an idea of dual-stage feature alignment, processes the information of small motion and large motion respectively, reduces the deviation between the target frame and the reference frame, fully utilizes the feature information of all layers, makes the video spatial information not easy to be lost by using multiple attentions, retains the spatial information, enhances the channel dependence and self-adaptability, and can capture long-distance dependencies to achieve global learning.

[0049] Embodiment 1

[0050] Combined with Figure 1, A video super-resolution model that combines multiple attentions with optical flow in this embodiment includes a feature extraction part, a feature processing part, a deformable convolution part, and a video reconstruction part; video frames sequentially pass through the four parts to achieve super-resolution; the feature processing part includes a multi-attention branch and an attention optical flow estimation branch; the multi-attention branch includes a spatial attention module, a self-attention module, a convolution module, and an upsampling module; the attention optical flow estimation branch includes a spatial attention module, a channel attention module, an optical flow estimation network module, a convolution module, and an upsampling module. The modules in the multi-attention branch are, in the order of the video frames passing through, the spatial attention module - self-attention module - convolution module - upsampling module - convolution module; the module order of the attention optical flow estimation branch is that the video frames pass through the spatial attention module and the channel attention module simultaneously, and then enter the optical flow estimation network module - convolution module - upsampling module - convolution module. The feature extraction part includes two convolution modules and three residual dense modules, and the input video frames sequentially pass through the two convolution modules and the three residual dense modules, and then enter the feature processing module; the video reconstruction module is a convolution module.

[0051] A video super-resolution method that combines multiple attentions with optical flow using the above model of the present invention, the steps are as follows:

[0052] Step 1, input 2n + 1 consecutive low-resolution video frames: Represent the input size of the multi-attention optical flow network (MAFnet) as (M L ×N L ), and the input LR frames are a sequence of 2n + 1 LR frames where is the output HR frame, denoted as I SR , with a size of (M H ×N H ), and M H >M L , N H >N L .

[0053] Step 2, input the video frames into the feature extraction part of the model to extract the video frame feature F: Send the input video sequence to the first part for feature extraction, and the video frames obtain the feature F through two convolution operations and residual dense block operations:

[0054] F = H rdb (H c1 (H c0 (I LR ))) (1)

[0055] where, I LR represents the input low-resolution frame, and H rdb (·) represents the residual dense block operation, and Hc (·) represents the convolution operation.

[0056] Step 3: Send the extracted feature F into the multi-attention branch and the attention optical flow estimation branch respectively, and the outputs of the two branches, F a and F f : The extracted feature F is sent into the multi-attention branch and the attention optical flow estimation branch, and the outputs of the multi-attention branch, Equation (2), and the attention optical flow estimation branch, Equation (3), are obtained respectively.

[0057] F a = H se (H sa (H ca (F))) (2)

[0058] F f = H f (H sa (F), H ca (F)) (3)

[0059] where H se (·) is the self-attention module function, H sa (·) is the spatial attention module function, H ca (·) is the channel attention module function, H f (·) is the optical flow module function. The optical flow estimation network is the first stage of two-stage feature alignment, mainly dealing with small motions. The feature sent into the multi-attention branch passes through spatial attention and self-attention, aiming to enhance channel dependence and self-adaptability, retain spatial information, and achieve global learning.

[0060] Combined with Figures 2 - 4 , the specific structures and processes of each attention and optical flow estimation network are as follows.

[0061] (1) Channel attention

[0062] Channel attention considers the mutual dependence features between feature channels and adaptively adjusts the channel features. The feature obtained after the first feature extraction part is used as the input feature. At this time, the feature size is H×W×C. After passing through adaptive average pooling and adaptive max pooling respectively, the feature size becomes 1×1×C. Then, a convolution with a kernel size of 3 is used to change the feature size to r is the channel reduction ratio, which is set to 16 in this embodiment; then, the ReLU activation function is passed through. Subsequently, the two features after pooling both pass through a 3x3 convolution to restore the channels, with a size of 1×1×C. The obtained features are added and passed through a Sigmoid function to obtain the attention feature map, and then the attention feature map is multiplied by the input feature in matrix form to obtain the output feature, with a size of H×W×C.

[0063] (2) Spatial attention

[0064] Spatial attention can assign weights to each spatial position, make more effective use of cross-channel and spatial information, capture the spatial dependence between any positions in the feature map, and expose as much spatial information as possible. The output feature of channel attention is used as the input feature of spatial attention, which first passes through a 1x1 convolution and the activation function LReLU. The reason for choosing LReLU instead of ReLU in spatial attention is that ReLU may cause neuron death during training and cannot further update the parameter gradient. Using LReLU can alleviate this problem and better retain spatial information. After passing through the pooling layer, which is composed of average pooling, max pooling, and concatenation operations, the feature obtained after passing through a 1x1 convolution and LReLU is denoted as Feature 1. Then, through repeated 1x1 convolution, LReLU, and pooling layer structures, followed by a 3x3 convolution and LReLU, and repeating this structure once, and performing interpolation operations, the obtained feature is denoted as Feature 2. After adding Feature 1 and Feature 2, passing through a 1x1 convolution, LReLU, and interpolation operations, the feature is successively fed into a 3x3 convolution, a 1x1 convolution, and LReLU to obtain a feature denoted as Feature 3. Using the Sigmoid function to obtain the attention feature map, multiplying the attention feature map and the input feature matrix, and adding the result to Feature 3 to obtain the output feature. Spatial attention provides an effective and reliable guarantee for using two-dimensional convolution to implement feature processing in the spatio-temporal domain.

[0065] (3) Self-attention

[0066] The prototype of self-attention comes from the non-local operation network and can be inserted as an effective component into any existing network. In addition to expanding the receptive field, it can also calculate the distance relationship between any two positions in space, replace the skip connection, and achieve the function of global learning. After the feature is input, it passes through three convolutional channels respectively to obtain Feature 1, Feature 2, and Feature 3. Feature 1 and Feature 2 are multiplied matrix-wise and passed through the softmax function to obtain the attention map, and then multiplied matrix-wise with Feature 3 to obtain the output feature. The convolutional kernel size in this structure is all 1x1.

[0067] (4) Optical flow estimation

[0068] Traditional motion compensation methods have problems of high computational complexity and low accuracy. In this embodiment, a method of combining attention and optical flow is used to process the motion information of small moving objects while retaining object-related information to achieve feature alignment in the first stage. The features extracted from the first part are respectively passed through channel attention and spatial attention, and the outputs of both are fed into the optical flow estimation network to obtain the output of this branch.

[0069] Given any two adjacent frames I i , I i+1 , then the optical flow calculation formula can be expressed as

[0070] f i→i+1 = N f (I i , I i+1 ) (4)

[0071] where N f represents the optical flow estimation network.

[0072] Step 4: After upsampling F a and F f , input them into a deformable LSTM (i.e., DLSTM) and a convolutional model to obtain the video super-resolution feature F d : Upsample F a and F f

[0073] y 1 = H c4 (↑(H c2 (F a ))) (5)

[0074] y 2 = H c5 (↑(H c3 (F f ))) (6)

[0075] where ↑ represents upsampling; input y 1 , y 2 into DLSTM, and then obtain the final output through one layer of convolution

[0076] Fd = H c6 (DLSTM(y 1 , y 2 ) (7)

[0077] where F d represents the feature obtained through DLSTM and the final reconstruction convolution; the entire network is finally expressed as

[0078] I SR = H MAFnet (I LR ) (8).

[0079] ​Among them, DLSTM is adding deformable convolution to the traditional LSTM. Compared with the traditional convolution, deformable convolution can adjust the displacement of spatial position information, and is less likely to introduce grid artifacts compared with dilated convolution. It retains the original advantages of LSTM, while enhancing the ability of video frames to be aligned in time series, effectively using context information to process large motion information in the video. It ensures the continuity of the video.

[0080] The above schematically describes the present invention and its implementation manners. This description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. A video super-resolution model that combines multiple attentions with optical flow, characterized in that: The model includes a feature extraction part, a feature processing part, a deformable convolution part, and a video reconstruction part; Video frames sequentially pass through the four parts to achieve super-resolution; the feature processing part includes a multi-attention branch and an attention optical flow estimation branch; The multi-attention branch includes a spatial attention module, a self-attention module, a convolution module, and an upsampling module; The attention optical flow estimation branch includes a spatial attention module, a channel attention module, an optical flow estimation network module, a convolution module, and an upsampling module; the modules in the multi-attention branch are, in the order of video frames passing through, spatial attention module - self-attention module - convolution module - upsampling module - convolution module; The modules in the attention optical flow estimation branch are, in the order of video frames passing through: passing through the spatial attention module and the channel attention module simultaneously, and then entering the optical flow estimation network module - convolution module - upsampling module - convolution module.

2. A video super-resolution model that combines multiple attentions with optical flow according to claim 1, characterized in that: The feature extraction part includes two convolution modules and three residual dense modules. The input video frames sequentially pass through the two convolution modules and the three residual dense modules, and then enter the feature processing module; the video reconstruction part is a convolution module.

3. A method for video super-resolution that combines multiple attentions with optical flow using the model of claim 2, characterized in that, its steps are: Step 1: Input 2n + 1 consecutive low-resolution video frames; Step 2: Input the video frames into the feature extraction part of the model to extract the video frame feature F; Step 3: Send the extracted feature F into the multi-attention branch and the attention optical flow estimation branch respectively, and the outputs of the two branches, F a and F f ; Step 4. After upsampling F a and F f , input them into a DLSTM and a convolutional model to obtain the video super-resolution feature F d .

4. A method for video super-resolution that combines multiple attentions with optical flow according to claim 3, characterized in that: In the first step described above, the input video frames are 2n + 1 LR frames in MAFnet, and their sequence is The input size of MAFnet is (M L ×N L ), where is the output HR frame, denoted as I SR , with a size of (M H ×N H ), and M H > M L , N H > N L .

5. A method for video super-resolution that combines multiple attentions with optical flow according to claim 4, characterized in that: In the said step 2, the video frames obtain the feature F through two convolution operations and residual dense block operations, F = H rdb (H c1 (H c0 (I LR ))) (1) Among them, I LR represents the input low-resolution frame, H rdb (·) represents the residual dense block operation, H c (·) represents the convolution operation.

6. A method for video super-resolution that combines multiple attentions with optical flow according to claim 5, characterized in that: In the said step 3, the extracted feature F is sent to the multi-attention branch and the attention optical flow estimation branch, and the outputs of the multi-attention branch, formula (2), and the output of the attention optical flow estimation branch, formula (3), are respectively obtained, F a = H se (H sa (H ca (F))) (2) F f = H f (H sa (F), H ca (F)) (3) Among them, H se (·) is the self-attention module function, H sa (·) is the spatial attention module function, H ca (·) is the channel attention module function, H f (·) is the optical flow module function.

7. A method for video super-resolution that combines multiple attentions with optical flow according to claim 6, characterized in that: In the said channel attention module, after the feature is input, it is respectively subjected to adaptive average pooling and adaptive max pooling, then the convolution channels are reduced and passed through the activation function ReLU, and then the channels are restored through convolution; the two obtained features are added and passed through the Sigmoid function to obtain the attention feature map, and then the attention feature map is multiplied by the input feature in matrix form to obtain the output feature; In the described spatial attention module, after the feature is input, it first passes through convolution and the activation function LReLU, then through a pooling layer composed of average pooling, max pooling, and concatenation operations, and then through convolution and LReLU to obtain Feature 1. After that, it passes through a repeated structure of convolution, LReLU, and pooling layer, and through two convolution and LReLU structures, and then interpolation operation is performed to obtain Feature 2. After adding Feature 1 and Feature 2, convolution, LReLU, and interpolation operations are carried out. The feature is successively fed into two convolution and LReLU operations to obtain Feature 3. The attention feature map is obtained by using the Sigmoid function. The attention feature map and the input feature are multiplied matrix-wise, and the result is added to Feature 3 to obtain the output feature; In the described self-attention module, after the feature is input, it passes through three convolution channels respectively to obtain Feature 1, Feature 2, and Feature 3. Feature 1 and Feature 2 are multiplied matrix-wise and passed through the softmax function to obtain the attention map, and then multiplied matrix-wise with Feature 3 to obtain the output feature.

8. A video super-resolution method combining multiple attentions and optical flow according to claim 7, characterized in that: In the optical flow estimation branch described above, given any two adjacent frames I i , I i+1 , the optical flow calculation formula can be expressed as f i→i+1 = N f (I i , I i+1 )(4) Among them, N f represents an optical flow estimation network.

9. A video super-resolution method combining multiple attentions and optical flow according to claim 7, characterized in that: In the fourth step described above, for F a and F f perform upsampling y 1 = H c4 (↑(H c2 (F a ))) (5) y 2 = H c5 (↑(H c3 (F f ))) (6) Among them, ↑ represents upsampling; y 1 , y 2 is fed into the DLSTM, and then the final output is obtained through one layer of convolution F d = H c6 (DLSTM(y 1 , y 2 )(7) Among them, F d represents the features obtained through DLSTM and the final reconstruction convolution; the entire network is finally represented as I SR = H MAFnet (I LR ) (8).

Citation Information

Patent Citations

  • Video high-temporal-spatial-resolution signal processing method combining optical flow method and deep network

    CN110634105A

  • Video super-resolution reconstruction method based on deep dual attention network

    CN110969577A