Video Super-Resolution Reconstruction Method and System Based on Multi-Scale Local Self-Attention

By using multi-scale local self-attention and optical flow alignment modules in video super-resolution reconstruction technology, the problems of high computing complexity and large memory usage in the prior art are solved, and the efficient and low-complexity video super-resolution reconstruction effect is achieved.

CN115082308BActive Publication Date: 2025-06-10SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210564009.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-23
Publication Date
2025-06-10
Estimated Expiration
2042-05-23

AI Technical Summary

Technical Problem

The existing video super-resolution reconstruction technology has problems such as high computational complexity, large memory usage, many training iterations and high hardware requirements, and the global self-attention mechanism cannot effectively utilize local strong relevant information of image data.

Method used

Using a video super-resolution reconstruction method based on multi-scale local self-attention, the multi-scale deep feature extraction module and optical flow alignment module are constructed to reduce the network computation and limit the self-attention mechanism to local areas to accelerate model convergence and improve performance.

Benefits of technology

It reduces the computing volume and memory usage of the network, improves the convergence speed and performance of the model, avoids the computing complexity and noise interference caused by the global self-attention mechanism, and realizes efficient video super-resolution reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115082308B_ABST
    Figure CN115082308B_ABST
Patent Text Reader

Abstract

The present invention discloses a video super-resolution reconstruction method and system based on multi-scale local self-attention. The method includes: S1: constructing a low-resolution video frame sequence dataset; S2: predicting bidirectional optical flow information between adjacent frames in the input of the low-resolution video frame sequence through an optical flow prediction network; S3: constructing a video super-resolution reconstruction network, which includes a feature extraction module, a multi-scale deep feature extraction module, and an upsampling reconstruction module; S4: training the video super-resolution reconstruction network based on the dataset and the bidirectional optical flow information; S5: inputting the video sequence that needs super-resolution reconstruction into the trained video super-resolution reconstruction network, and the super-resolution reconstructed video sequence can be obtained. The present invention can reduce the overall computational amount of the network and strengthen information fusion through the optical flow prediction network, and has a good reconstruction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer vision technology, and particularly relates to a video super-resolution reconstruction method and system thereof. Background Art

[0002] Video super-resolution reconstruction technology is widely used in many scenarios such as video live broadcast, security monitoring, satellite remote sensing, etc., and has great practical significance. With the continuous improvement of the resolution of terminal display devices and the rapid development of video transmission requirements, it is urgent to find a low-cost and high-efficiency reconstruction solution for the existing massive video data, in order to obtain better display effects on high-resolution display devices.

[0003] The key to the video super-resolution reconstruction task is the utilization of redundant information between video frames. The dense sampling of video capture devices in time series can capture the sub-pixel displacement of objects, providing necessary information for super-resolution. There are many solutions for video super-resolution reconstruction. The current mainstream solutions are mostly based on deep learning. The main process is to use a deep learning model to extract video features, align redundant information between frames, and reconstruct a high-resolution video. The processing ideas are roughly divided into the sliding window method and the loop method. The sliding window method divides the video reconstruction task into multiple sub-tasks of multi-frame reconstruction, and uses multiple low-resolution images to reconstruct a high-resolution image; the loop method generally only needs to input one image, and then refers to the output results of the previous reconstructed images. The former has redundant calculations, but the advantage is that the sub-tasks do not affect each other. The latter is more efficient, but there is a common problem of cumulative errors in the loop structure, and the performance drops significantly when dealing with complex videos in real environments.

[0004] At present, the Transformer structure in deep learning has achieved great success in the field of natural language processing and has also begun to emerge in the field of image processing and analysis. It is worth noting that the self-attention mechanism of the Transformer structure can also well meet the need to fuse similar patterns in the video super-resolution task. The Transformer structure can aggregate the information of the feature map over a long distance. In theory, compared with the convolutional neural network, the receptive field of the Transformer structure is larger, can see more information, and has better effects. However, this comes at the cost of quadratic computational complexity and extremely high memory occupancy. Therefore, in the field of images, image patches are used as the smallest unit (token) for calculating self-attention instead of pixels. However, the movement of object pixels in the video does not necessarily coincide with the image patches they are in, resulting in the inability to achieve fine fusion through patch-level self-attention fusion. On the other hand, the global self-attention mechanism adopted to "see more information" discards the prior information of strong local correlation in the image data, so it additionally requires a long training time and a large number of parameters to "re-learn" this information. Cao J et al. designed a video super-resolution reconstruction network based on global self-attention, called VSR-Transformer, in "Video super-resolution transformer[J]. arXiv preprint arXiv:2106.06847, 2021". On the one hand, this network adopts global self-attention, which occupies a huge amount of resources, so there are strict constraints on the resolution of the input video frames to be processed. Before calculating the global self-attention, the video frames need to be first segmented into the maximum resolution that the network can handle, and the global self-attention is calculated separately for the multiple segmented video frames under this resolution constraint, and finally the respective results are stitched together. To prevent the grid effect caused by stitching, there also needs to be partial overlap when segmenting the video frames, which leads to a large amount of computational redundancy. When this network calculates the global self-attention, it will also segment the video frames that meet the resolution constraint again, and use the segmented small patches as the smallest unit of self-attention. On the other hand, this network maintains the spatial resolution of the feature map unchanged during processing, which is not conducive to dealing with large-scale optical flow changes and has a high computational demand. Although this super-resolution network can achieve good results, it has a large number of parameters and a large amount of computation, requires too many training iterations, has high hardware requirements, and has insufficient operability. Summary of the Invention

[0005] In view of the above deficiencies of the prior art, the present invention proposes a video super-resolution reconstruction method based on multi-scale local self-attention. This method constructs a multi-scale deep feature extraction module to reduce the overall computational amount of the network, and realizes inter-frame alignment based on an optical flow prediction network, strengthening local information fusion. At the same time, the present invention confines the self-attention of the Transformer structure from global to local, enabling it to focus more on local regions with higher information correlation and excluding noise interference.

[0006] To achieve the object of the present invention, the present invention provides a video super-resolution reconstruction method based on multi-scale local self-attention, including the following steps:

[0007] S1: Downsample the high-resolution video data to obtain the corresponding low-resolution video frame sequence, and divide the frame sequence to form a training set and a test set;

[0008] S2: Predict the bidirectional optical flow information between adjacent frames in the input of the low-resolution video frame sequence through an optical flow prediction network.

[0009] S3: Construct a video super-resolution reconstruction network, which includes a feature extraction module, a multi-scale deep feature extraction module, and an upsampling reconstruction module. Among them, the feature extraction module is used to extract the shallow features of the video frames from the input low-resolution video frame sequence, the multi-scale deep feature extraction module is used to obtain deep feature maps based on the shallow features, and the upsampling reconstruction module is used to reconstruct the low-resolution video sequence to obtain a high-resolution video sequence;

[0010] S4: Train the video super-resolution reconstruction network based on the data set and the bidirectional optical flow information;

[0011] S5: Input the video sequence that needs super-resolution reconstruction into the video super-resolution reconstruction network obtained after training, and a super-resolution reconstructed video sequence can be obtained.

[0012] In one implementation manner of step S1, downsample the high-resolution video frames and input them in units of 5 consecutive frames.

[0013] In one implementation manner of step S2, a pre-trained optical flow prediction network is used to extract the optical flow change between adjacent frames.

[0014] Further, in step S2, the low-resolution video frame sequence is input into the optical flow prediction network in the forward and reverse directions respectively to obtain the bidirectional optical flow information flow forward ,flow backward ,flow forward represents the optical flow information from the future moment to the past moment in the sequence, and flow backwardIt represents the optical flow information from the past to the future and outputs optical flow information at multiple scales through downsampling.

[0015] Furthermore, the multi-scale deep feature extraction module includes a cascade of multiple encoders and multiple decoders equal in number to the encoders. The encoders gradually downsample to obtain feature maps at multiple scales, and then the decoders gradually upsample to restore the size of the feature maps.

[0016] Furthermore, both the encoder and the decoder are composed of a cascade of a local self-attention module (Local Self-Attention, LSA) and an optical flow alignment module (Flow Alignment, FA). The LSA module first divides the input feature map into multiple patches, and then constrains the attention range of the self-attention mechanism of the Transformer to the local area, fusing similar patches within the local area; the FA module uses the bidirectional optical flow information of adjacent frames extracted by the aforementioned optical flow prediction network. First, it warps the video frames forward and backward along two branches respectively to achieve alignment of adjacent frames, and then inputs the aligned results into two parallel residual networks for processing respectively. Finally, it is fused through convolution operations.

[0017] Each encoder and decoder includes a local self-attention module and an optical flow alignment module. The operation steps in the local self-attention module include:

[0018] Divide the shallow feature map of the input video frame into small image patches with a resolution of p H ×p W without overlap, obtaining:

[0019]

[0020] where x unfold represents the tensor composed of small image patches at different times and spaces after segmentation, H and W represent height and width respectively, B represents the parallel processing batch size, T represents the length of a single input video sequence, and C represents the number of channels;

[0021] Divide adjacent small patches into several non-overlapping local windows, obtaining:

[0022]

[0023] where x local represents the tensor composed of the segmented local windows, L T represents the window range in the time dimension, L H and L W represent the height and width of the spatial window respectively;

[0024] For the tensor x localInput three independent linear layers Query, Key, and Value respectively to obtain three feature maps:

[0025]

[0026] Wherein, and respectively represent the feature maps after linear transformation by the corresponding linear layers, represents the batch size of the feature map after linear transformation, N' = L T ×L H ×L W represents the number of small patches within the local window, C' = C × p H ×p W represents the number of channels of the feature map after linear transformation;

[0027] Calculate the self-attention of small patches in the local area and fuse similar patches:

[0028]

[0029] Wherein, x sa represents the feature map obtained after local self-attention fusion.

[0030] Recombine and splice the fused feature map x sa into the original resolution to obtain the feature map x fold .

[0031] Furthermore, the optical flow alignment module includes a forward alignment module and a backward alignment module, and the operations in the optical flow alignment module include:

[0032] Use the bidirectional optical flow information flow forward , flow backward to perform adjacent frame alignment operations on the feature map x fold after restoring the resolution respectively;

[0033] Process the aligned results through the residual module respectively;

[0034] Fuse the results of the forward alignment module and the backward alignment module. The alignment operation warp is implemented through backward alignment.

[0035] Furthermore, step S4 includes the following sub-steps:

[0036] Step S41: Extract multiple groups of low-resolution video sequence samples and corresponding original high-resolution video sequence samples from the training set as a single training data;

[0037] Step S42: Input the low-resolution video sequence sampled from the training set and the bidirectional optical flow information into the video super-resolution reconstruction network for training, calculate the difference between the reconstructed high-resolution video frames and the corresponding real video frame samples using a loss function, and adjust the network parameters according to the difference until the video super-resolution reconstruction network converges.

[0038] The loss function in S4 is as follows:

[0039]

[0040] where N represents the number of samples sampled at each step of training, I(x, y, c) represents the intensity value of the corresponding high-resolution image, represents the intensity value of the x-th row, y-th column, and c-th channel in the reconstructed image, H represents the height of the video frame, W represents the width of the video frame, and ε is a small constant to prevent the calculation result from being 0.

[0041] Furthermore, the peak signal-to-noise ratio (PSNR) is used to evaluate the reconstruction effect. The peak signal-to-noise ratio represents the ratio between the maximum possible power of the signal and the noise power, and is commonly used as an objective evaluation index for the quality of signal reconstruction.

[0042] The larger the value of the peak signal-to-noise ratio, the better the reconstruction effect.

[0043]

[0044] where n represents the number of bits representing the intensity of the color channel. PSNR represents the value of the peak signal-to-noise ratio, and the unit is dB.

[0045] The present invention also provides a video super-resolution reconstruction system based on multi-scale local self-attention, including:

[0046] An optical flow information prediction module, configured to predict the bidirectional optical flow information between adjacent frames in the input low-resolution video frame sequence through an optical flow prediction network;

[0047] A video super-resolution reconstruction network training module, configured to train a video super-resolution reconstruction network based on a data set and the bidirectional optical flow information;

[0048] A reconstruction module, configured to input the video sequence to be super-resolved into the trained video super-resolution reconstruction network, and a super-resolved video sequence can be obtained.

[0049] The above technical solution conceived by the present invention can at least achieve the following beneficial effects:

[0050] 1. In the multi-scale deep feature extraction module composed of the encoder and decoder mentioned in the present invention, the process of gradually downsampling also gradually increases the receptive field of the network, strengthens the network's ability to capture optical flow changes of different amplitudes between frames, and improves the performance of the network. At the same time, the decrease in the size of the feature map also greatly reduces the computational complexity of the network.

[0051] 2. The method of restricting self-attention to the local area adopted in the present invention can, on the one hand, add the prior information of strong local correlation existing in the image data to the network, accelerate the convergence speed of the network model, and improve the network performance; on the other hand, it can avoid the large amount of computational complexity and memory occupation brought by the quadratic complexity of the self-attention mechanism itself. Adopting local self-attention can achieve the overall reconstruction of video frames, that is, the network does not need to strictly constrain the resolution of the input video sequence to be processed, and does not need to segment the video frames before calculating self-attention. This operation avoids the potential grid effect brought by video frame segmentation processing and the redundant calculation introduced to make up for it.

[0052] 3. The pixel-level optical flow alignment module adopted in the present invention can achieve frame-to-frame alignment in a more fine-grained manner, thus solving the stitching problem caused by using image patches as the smallest unit of self-attention. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic flowchart provided by an embodiment of the present invention;

[0054] Figure 2 is a schematic diagram of the overall network structure provided by an embodiment of the present invention;

[0055] Figure 3 is a schematic diagram of the local self-attention module in an embodiment of the present invention;

[0056] Figure 4 is a schematic diagram of the optical flow alignment (FA) module in an embodiment of the present invention;

[0057] Figure 5 is a schematic diagram comparing the video reconstruction results of the method of the present invention with the prior art. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] Please refer to Figure 1, a video super-resolution reconstruction method based on multi-scale local self-attention provided by the present invention includes the following steps:

[0060] Step S1: Make a super-resolution reconstruction dataset.

[0061] In the present invention, the specific steps of making the dataset include:

[0062] Step S11: Given high-resolution video data, intercept high-resolution video sequences from multiple original high-resolution videos;

[0063] Step S12: Downsample the high-resolution video sequence to obtain a low-resolution video sequence, reduce the original image by 4 times in space, then take 10% of the low-resolution video sequence and divide it into a test set, and the remaining 90% is divided into a training set.

[0064] In some embodiments of the present invention, the low-resolution video frame sequence is composed of a central frame and several adjacent front and rear auxiliary frames.

[0065] In some embodiments of the present invention, bicubic downsampling is specifically used to obtain the low-resolution video sequence. Of course, in other embodiments, other downsampling methods can also be used.

[0066] Step S2: Pre-acquire inter-frame optical flow information. Input the low-resolution video sequence into an optical flow prediction network, and predict the bidirectional optical flow information between adjacent frames in the low-resolution video frame sequence through the optical flow prediction network.

[0067] In some embodiments of the present invention, the optical flow prediction network SPyNet is used for prediction. The optical flow prediction network SPyNet is an existing network (Ranjan A, Black M J. Optical flow estimation using a spatial pyramid network[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 4161-4170). Initialize the pre-trained optical flow prediction network SPyNet. The main body of this network is a convolutional neural network, which can predict the pixel movement between any two adjacent video frames and output the result in the form of a feature map with 2 channels. These two channels represent the horizontal and vertical movements on the two-dimensional plane, representing the optical flow information between two adjacent frames. Input the low-resolution video frame sequence into the network in the forward and reverse directions respectively to obtain the bidirectional optical flow information flow forward , flow backward . flow forwardRepresents the optical flow information from future moments to past moments in the representative sequence, flow backward It represents the optical flow information from past to future. Then, downsampling changes the optical flow information, that is, the spatial resolution of the output feature map, to obtain optical flow information at multiple scales. The optical flow information obtained in this step is saved for subsequent training use.

[0068] Step S3: Construct a video super-resolution reconstruction network, which includes a feature extraction module, a multi-scale deep feature extraction module, and an upsampling reconstruction module. Among them, the feature extraction module is used to extract the shallow features of video frames from the input low-resolution video frame sequence; the multi-scale deep feature extraction module is used to obtain deep feature maps based on the shallow features; the upsampling reconstruction module is used to reconstruct the low-resolution video sequence to obtain a high-resolution video sequence. Specifically, the upsampling reconstruction module upsamples the central frame map of the original low-resolution video sequence without processing by bilinear interpolation ×4, changes the number of channels of the deep feature map to be consistent with the upsampled central frame through a convolutional layer, then upsamples by PixelShuffle ×4, and finally adds the central frame and the deep feature Figure 2 to obtain the reconstruction result.

[0069] In the present invention, the multi-scale deep feature extraction module includes a plurality of cascaded encoders and a plurality of decoders equal in number to the encoders. The encoders gradually downsample to obtain feature maps at multiple scales, and then the decoders gradually upsample to restore the size of the feature maps. At the same time, the decoders fuse the feature maps of the same scale output by the encoders, and after upsampling, they are used as the input of the next-level decoder. Among them, each encoder and decoder includes a local self-attention module and an optical flow alignment module.

[0070] Specifically, Figure 2 As shown, it is the overall structure of the video super-resolution reconstruction network involved in the present invention, including a feature extraction module, a multi-scale deep feature extraction module, and an upsampling reconstruction module.

[0071] Let the input low-resolution video frame sequence B represents the parallel processing batch size, T represents the length of the single input video sequence, C represents the number of channels, and H and W represent the height and width respectively. In this embodiment, T = 5, C = 3, H = 64, W = 64.

[0072] The feature extraction module includes a residual network, which is used to initially extract the feature information of the input low-resolution video sequence by the residual network and increase the number of input channels to 64 to obtain a shallow feature map. Then, the shallow feature map is input into the multi-scale deep feature extraction module, which first gradually downsamples and then gradually upsamples to restore the resolution.

[0073] In some embodiments of the present invention, the multi-scale deep feature extraction module includes 6 identical processing units composed of cascading local self-attention (LSA) modules and optical flow alignment (FA) modules, and each processing unit will add upsampling or downsampling operations according to its location. Among these 6 processing units, in the order of the feature map flow direction, the first 3 are all encoders, and the last 3 are all decoders. The encoder will also downsample the input feature map to achieve the effect of encoding. This measure can reduce the noise during reconstruction. The decoder will also upsample the input feature map to restore the resolution of the feature map to meet the requirements of the output resolution. As Figure 2 shown, the deep processing unit will first fuse the output results of the shallow processing unit, and then upsample it as the input of the next-level processing unit. Further, the above fusion operation is implemented through a 3×3 convolutional layer.

[0074] In some embodiments of the present invention, the upsampling in the multi-scale deep feature extraction module is implemented using the PixelShuffle [Shi W, Caballero J, Huszár F, et al. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2016:1874-188] method. This method first uses a convolutional layer to increase the number of channels of the feature map by 4 times, and then rearranges the information within the channels to the spatial dimension to achieve ×4 upsampling; the downsampling is implemented using a 3×3 convolutional layer with a stride of 2.

[0075] In the present invention, please refer to Figure 3 , the operations of the local self-attention (LSA) module include:

[0076] First, divide the shallow feature map of the input video frame into small image patches with a resolution of p H ×p W without overlap, and obtain:

[0077]

[0078] where x unfold represents the tensor composed of small image patches at different times and spaces after segmentation.

[0079] In some embodiments of the present invention, p H = p W= 8, and the splitting operation is implemented through the unfold function of Pytorch.

[0080] Secondly, use the same operation to divide adjacent small patches into several non-overlapping local windows, obtaining:

[0081]

[0082] where x local represents the tensor composed of the local windows after splitting, and L T represents the window range in the time dimension, and L H , L w represent the height and width of the spatial window respectively. In some embodiments of the present invention, L T = 3, L H = 4, L W = 4, indicating that 3 adjacent small patches are taken in the time dimension, and small patches within the range of 4×4 are taken in the spatial dimension to construct the window.

[0083] Thirdly, input the tensor x local into three independent linear layers Query, Key, and Value respectively, obtaining three feature maps:

[0084]

[0085] where and respectively represent the feature maps after linear transformation by the corresponding linear layers, represents the batch size of the feature maps after linear transformation, N' = L T ×L H ×L W represents the number of small patches within the local window, C' = C×p H ×p w represents the number of channels of the feature maps after linear transformation.

[0086] Subsequently, calculate the self-attention of the small patches within the local region for each feature map in the batch and fuse similar patches. As Figure 3 shown, in some embodiments of the present invention, a multi-head attention mechanism is also adopted, and this method calculates multiple groups of Q, K, V, and fuses the results of multiple attention heads through convolution.

[0087]

[0088] where x sa represents the feature map obtained after local self-attention fusion.

[0089] Subsequently, use the fold operation on x saRecombine and splice to the original resolution, and obtain x through a 3×3 convolution fold . The convolution operation here can smooth the problem of boundary inconsistency caused by splicing to a certain extent.

[0090] Among them represents the feature map after restoring the resolution.

[0091] To prevent the gradient from being unable to be transmitted normally due to excessive numerical differences during training, it is necessary to perform LayerNorm normalization operations on the H and W dimensions of the feature map. Finally, add the normalized feature map to the input to obtain the output of the local self-attention module.

[0092] Specifically, the optical flow alignment (FA) module is as Figure 4 shown. According to the direction of optical flow alignment, the optical flow alignment module includes a forward alignment module and a backward alignment module. First, use the bidirectional optical flow information flow forward , flow backward obtained in the above steps to perform adjacent frame alignment operations on the feature map x fold after restoring the resolution. In this embodiment, to avoid the hole effect, the alignment operation warp is actually implemented through backward alignment.

[0093] x forward_align = warp(x fold , flow forward )

[0094] x backward_align = warp(x fold , flow backward )

[0095] Example of the warp operation: The input optical flow information indicates the movement required for all pixels in frame A to be aligned to frame B. Generally, the pixel coordinates in frame A should be added with the pixel movement to obtain the result of aligning frame A to frame B. However, due to reasons such as occlusion or perspective transformation, multiple pixels in frame A may need to be aligned to the same position in frame B, resulting in some positions in frame B having no corresponding pixels. Therefore, in operation, frame B will actually be aligned backward to frame A to avoid holes.

[0096] x forward_align is obtained by the forward alignment module and represents the feature map aligned from the future moment to the past moment of the video sequence. x backward_align is then obtained by the backward alignment module and represents the feature map aligned from the past to the future.

[0097] Subsequently, to prevent the bidirectional optical flow information from being incorrect due to reasons such as occlusion of objects in the video, which may cause the alignment operation to fail, the aligned results x forward_align , x backward_alignIt also needs to be processed separately through a residual module composed of several residual layers to obtain the corrected forward and backward alignment results. Finally, the convolutional layer fuses the forward and backward alignment results. The fusion operation can integrate the information of the forward and backward directions and make the alignment results more accurate. Similarly, in order to avoid abnormal phenomena such as gradient disappearance during training, this module also adopts the same normalization operation as the LSA module.

[0098] Step S4: Train the video super-resolution reconstruction network based on the dataset and the bidirectional optical flow information.

[0099] In the present invention, step S4 includes the following sub-steps:

[0100] Step S41: Extract multiple groups of low-resolution video sequence samples and corresponding original high-resolution video sequence samples from the training set as single training data.

[0101] In some embodiments of the present invention, randomly extract an array of low-resolution video sequence samples in groups of 5 frames and corresponding original high-resolution video sequence samples from the training set as single training data. This embodiment also crops the training data to a resolution of 64×64 and performs random rotation and flipping operations.

[0102] Step S42: Input the low-resolution video sequence sampled from the training set and the bidirectional optical flow information pre-computed in step S2 into the video super-resolution reconstruction network for training, and use the loss function to calculate the difference between the reconstructed high-resolution video frame and the corresponding real video frame sample, and adjust the network parameters according to the difference until the video super-resolution reconstruction network converges.

[0103] In some embodiments of the present invention, the loss function is:

[0104]

[0105] where N represents the number of samples sampled at each step of training, I(x,y,c) represents the intensity value of the corresponding high-resolution image, represents the intensity value of the x-th row, y-th column, and c-th channel in the reconstructed image, H represents the height of the video frame, W represents the width of the video frame, and ε is a small constant to prevent the calculation result from being 0.

[0106] In some embodiments of the present invention, the peak signal-to-noise ratio (PSNR) is used to evaluate the reconstruction effect. The peak signal-to-noise ratio represents the ratio between the maximum possible power of the signal and the noise power, and is an objective evaluation index for the signal reconstruction quality of this application. The larger the value of the peak signal-to-noise ratio, the better the reconstruction effect. When the network converges, the rising amplitude of the peak signal-to-noise ratio PSNR tends to be stable.

[0107]

[0108] Where n represents the number of bits for representing the intensity of the color channel. PSNR represents the value of the peak signal-to-noise ratio, and the unit is dB.

[0109] Step S5: Use the trained video super-resolution reconstruction network to reconstruct the low-resolution video sequence. Input the video sequence to be super-resolved into the video super-resolution reconstruction network, and the super-resolved video sequence can be obtained.

[0110] The present invention also provides a system for implementing the foregoing method.

[0111] A video super-resolution reconstruction system based on multi-scale local self-attention, comprising:

[0112] An optical flow information prediction module for predicting the bidirectional optical flow information between adjacent frames in the input of the low-resolution video frame sequence through an optical flow prediction network;

[0113] A video super-resolution reconstruction network training module for training the video super-resolution reconstruction network based on a data set and the bidirectional optical flow information;

[0114] A reconstruction module for inputting the video sequence to be super-resolved into the trained video super-resolution reconstruction network, and the super-resolved video sequence can be obtained.

[0115] To verify the effectiveness of the method proposed in the embodiments of the present invention, Table 1 compares this embodiment with the VSR-Transformer method mentioned in the background art. As can be seen from Table 1, compared with the VSR-Transformer method, the peak signal-to-noise ratio of the method of the present invention only drops by 0.26 dB, but the required number of parameters is only 54.6% of the latter, and the amount of computation is 47.0% of the latter. Figure 5 Then a comparison is made between the two from a visual perspective. From a visual perspective, the method provided in the embodiments of the present invention can obtain good reconstruction effects.

[0116] Table 1. Comparison of the number of parameters, amount of computation, and performance metrics for the 4×SR task on the REDS4 dataset

[0117] VSR-Transformer Method The Method of the Present Invention Params(M) 32.6 17.8 FLOPs(G) 570 268 PSNR(dB) 31.19 30.93

[0118] Where Params represents the number of model parameters, M represents 10 6 ; FLOPs represents the number of floating-point operations, which is used to indicate the amount of computation, and G represents 10 9; PSNR stands for Peak Signal-to-Noise Ratio. Nah S et al. proposed the REDS4 dataset in "Ntire 2019 challenge on video deblurring and super-resolution: Dataset and study[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition Workshops. 2019:0-0."

[0119] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. Video super-resolution reconstruction method based on multi-scale local self-attention, characterized in that, it includes the following steps: S1: Construct a low-resolution video frame sequence dataset and divide it into a training set and a test set; S2: Predict the bidirectional optical flow information between adjacent frames in the input of the low-resolution video frame sequence through an optical flow prediction network; S3: Construct a video super-resolution reconstruction network, which includes a feature extraction module, a multi-scale deep feature extraction module, and an upsampling reconstruction module. Among them, the feature extraction module is used to extract the shallow features of video frames from the input low-resolution video frame sequence, the multi-scale deep feature extraction module is used to obtain deep feature maps based on the shallow features, and the upsampling reconstruction module is used to reconstruct the low-resolution video sequence to obtain a high-resolution video sequence; S4: Train the video super-resolution reconstruction network based on the dataset and the bidirectional optical flow information; S5: Input the video sequence that needs super-resolution reconstruction into the trained video super-resolution reconstruction network, and the super-resolution reconstructed video sequence can be obtained; In step S2, the low-resolution video frame sequence is input into the optical flow prediction network in the forward and reverse directions respectively to obtain the bidirectional optical flow information flow forward , flow backward , flow forward represents the optical flow information from future moments to past moments in the sequence, and flow backward represents the optical flow information from past to future, and multiple scales of optical flow information are output through downsampling; The multi-scale deep feature extraction module includes a cascade of multiple encoders and multiple decoders equal in number to the encoders. The encoders gradually downsample to obtain multi-scale feature maps, and then the decoders gradually upsample to restore the size of the feature maps; each encoder and decoder includes a local self-attention module and an optical flow alignment module. The operation steps in the local self-attention module include: Divide the shallow feature map of the input video frame into small image patches with a resolution of p H × p W × p that do not overlap, obtaining: where x unfold represents a tensor composed of small image patches in different time and space after segmentation, H and W respectively represent the height and width of the video frame, B represents the parallel processing batch size, T represents the length of a single input video sequence, and C represents the number of channels; Divide adjacent small patches into several non-overlapping local windows to obtain: where x local represents the tensor formed by the segmented local windows, and L T represents the window range in the time dimension, and L H , L W represent the height and width of the spatial window respectively; Input the tensor x local into three independent linear layers Query, Key, and Value respectively to obtain three feature maps: Wherein, and respectively represent the feature maps after linear transformation by the corresponding linear layers, represents the batch size of the feature maps after linear transformation, N ′ = L T × L H × L W represents the number of small patches within the local window, C ′ = C × p H × p W represents the number of channels of the feature maps after linear transformation; For each feature map in the batch Calculate the self-attention of small patches in the local area and fuse similar patches: where x sa represents the feature map obtained after local self-attention fusion; The feature map x obtained after fusion sa is recombined and spliced into the original resolution to obtain the feature map x after restoring the resolution fold ; The optical flow alignment module includes a forward alignment module and a backward alignment module. The operations in the optical flow alignment module include: Utilize bidirectional optical flow information flow forward , flow backward Perform adjacent frame alignment operations on the feature map x after restoring the resolution fold respectively; Process the aligned results through a residual module respectively; Fuse the results of the forward alignment module and the backward alignment module.

2. The video super-resolution reconstruction method based on multi-scale local self-attention according to claim 1, characterized in that, The alignment operation warp is implemented through backward alignment.

3. The video super-resolution reconstruction method based on multi-scale local self-attention according to any one of claims 1-2, characterized in that, Step S4 includes the following sub-steps: Step S41: Extract multiple groups of low-resolution video sequence samples and corresponding original high-resolution video sequence samples from the training set as single training data; Step S42: Input the low-resolution video sequence sampled from the training set and the bidirectional optical flow information into the video super-resolution reconstruction network for training, and use the loss function to calculate the difference between the reconstructed high-resolution video frames and the corresponding real video frame samples, and adjust the network parameters according to the difference until the video super-resolution reconstruction network converges.

4. The video super-resolution reconstruction method based on multi-scale local self-attention according to claim 3, characterized in that, The loss function is: Among them, N represents the number of samples sampled at each step of training, I(x, y, c) represents the intensity value of the corresponding high-resolution image, represents the intensity value of the x-th row, y-th column, and c-th channel in the reconstructed image, H represents the height of the video frame, W represents the width of the video frame, and ε is a small constant to prevent the calculation result from being 0.

5. The video super-resolution reconstruction method based on multi-scale local self-attention according to claim 3, characterized in that, Peak signal-to-noise ratio PSNR is used to evaluate the reconstruction effect: Where n represents the number of bits for representing the color channel intensity, PSNR represents the value of the peak signal-to-noise ratio, H represents the height of the video frame, W represents the width of the video frame, and I(x, y, c) represents the intensity value of the corresponding high-resolution image. represents the intensity of the c-th channel at the x-th row and y-th column in the reconstructed image.

6. Video super-resolution reconstruction system based on multi-scale local self-attention, characterized in that, For implementing the method according to any one of claims 1-5, the system includes: An optical flow information prediction module, configured to predict bidirectional optical flow information between adjacent frames in a low-resolution video frame sequence input through an optical flow prediction network; A video super-resolution reconstruction network training module, configured to train a video super-resolution reconstruction network based on a data set and the bidirectional optical flow information; A reconstruction module, configured to input a video sequence to be super-resolution reconstructed into the trained video super-resolution reconstruction network, and a super-resolution reconstructed video sequence can be obtained.

Citation Information

Patent Citations

  • Video super-resolution reconstruction method based on multi-frame fusion optical flow

    CN111311490A

  • Video super-resolution reconstruction method based on multi-scale feature fusion

    CN112070667A