An efficient spatio-temporal super-resolution video compression and restoration method based on deep learning
By downsampling in the video compression encoding stage and using the spatiotemporal super-resolution network SRFI for restoration in the decoding stage, the existing deep learning video compression methods have solved the problems of high computational complexity and large parameters, and efficient spatiotemporal super-resolution video compression recovery is achieved, improving transmission efficiency and visual quality.
Patent Information
- Application Number
- CN202211285099.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-10-20
AI Technical Summary
Due to the high computational complexity of the existing video compression method, the existing video compression and recovery method hinders its application in the design of video compression and recovery schemes, and the number of parameters of the content adaptive image super-segment network greatly increases the transmission amount of the bitstream, limiting its applicability in low bit rate scenarios.
A highly efficient spatiotemporal super-resolution video compression and restoration method based on deep learning is proposed. By downsampling of spatial and temporal dimensions in the encoding stage, and designing spatiotemporal super-resolution network SRFI in the decoding stage, the spatiotemporal super-segment restoration of video is realized.
This method has a high level in transmission efficiency and visual quality, the model inference speed exceeds the same type of deep learning methods, and shows better bit transmission efficiency in low bit rate scenarios.
Smart Images

Figure CN115689917B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video encoding, decoding and restoration, and particularly to an efficient spatio-temporal super-resolution video compression and restoration method based on deep learning. Background Art
[0002] In recent years, the transmission proportion of video content in Internet traffic has experienced an explosive growth. Exploring and designing new video compression and restoration methods with lower bandwidth consumption and smaller visual quality loss is a research topic with profound application value and demand.
[0003] Traditional video compression algorithms, such as encoding standards like H.264, H.265, AV1, etc., mostly rely on manually crafted operation modules. For example, a region-based motion estimator and discrete cosine transform (DCT) are used to remove redundancy in the video to complete compression. Although these modules are well-designed and efficient, they cannot jointly optimize the entire compression scheme in an end-to-end manner.
[0004] With the continuous development of deep learning technology and its success in a series of visual tasks, video compression and restoration schemes based on deep learning technology are receiving increasing attention. These deep learning-based video compression methods use learnable deep neural networks to replace the handcrafted modules in traditional video codecs, thus achieving end-to-end optimization on large-scale video datasets. However, although they have achieved compression and restoration effects superior to current traditional codec standards, deep learning-based methods have high computational complexity overhead, which is an important problem and challenge currently hindering the application of deep learning technology in the design of video compression and restoration schemes.
[0005] To overcome this challenge, non-patent literature 1 (Liu, et.al. "Overfitting the Data: Compact Neural Video Delivery via Content-aware Feature Modulation.", ICCV, 2021.) and non-patent literature 2 (Khani, et.al. "Efficient Video Compression via Content-Adaptive Super-Resolution.", ICCV, 2021.) managed to reduce the bitrate during the encoding and decoding process by adding a lightweight content-adaptive image super-resolution network on the decoder side. Although the inference time of video decoding is very fast, their methods require relatively long training time to encode videos. In addition, the parameters of the above-mentioned content-adaptive image super-resolution network (usually more than 5 megabytes) also need to be included in the bitstream, which increases the amount of bits that need to be transmitted in the end, hindering their applicability in low-bitrate scenarios.
[0006] How to further combine multiple video restoration tasks into the process of video decompression, transmission and quality enhancement to improve the overall transmission efficiency and performance has become an important issue that needs to be urgently solved in academia and industry and has broad application value.
[0007] Based on the above problems, the present invention provides an efficient spatiotemporal super-resolution video compression and restoration method based on deep learning. Summary of the invention
[0008] In view of the shortcomings of existing video compression, restoration and transmission methods in performance and efficiency, the present invention provides a high-efficiency spatiotemporal super-resolution video compression and restoration method based on deep learning. It is used in the compression and distribution process of Internet video content. The restored image quality and bit transmission efficiency are at a high level, and the inference speed of the model exceeds that of similar deep learning methods.
[0009] The present invention is achieved through the following technical solution: an efficient spatiotemporal super-resolution video compression and restoration method based on deep learning, which follows the currently widely used video distribution process and includes an encoding (Encoder) stage and a decoding (Decoder) stage.
[0010] Step 1: Encoder
[0011] The task of the Encoder stage is to convert the given source video (A total of N video frames, with height H, width W, and C channels) are encoded and compressed to minimize the bit overhead during transmission. Applying the encoding algorithms of conventional compression coding standards to the source video is a straightforward solution. These codecs use a series of efficient handcrafted modules to reduce redundancy in the video. In an actual system, ready-made conventional video compression coding algorithms such as H.264 and H.265 can be used in this step. Considering the recent progress in video restoration tasks in deep learning, before generating the bitstream, downsampling operations in both the spatial and temporal dimensions can be performed first to further reduce the number of transmitted bits without sacrificing quality. The downsampled video will be restored through spatio-temporal super-resolution operations during the decoding phase.
[0012] Step 1 includes:
[0013] Step 1.1, extracting frames from the source video;
[0014] Step 1.2, downsampling the source video;
[0015] Step 1.3, compressing the source video;
[0016] Step 2, decoding phase
[0017] The task of the decoding phase is that after receiving the video bitstream from the encoding phase, this module needs to generate the final restored video whose visual quality needs to be as close as possible to that of the source video and have the same picture size and video frame rate.
[0018] Step 2 includes:
[0019] Step 2.1, decompressing the video processed in Step 1;
[0020] Step 2.2, performing spatio-temporal super-resolution restoration on the decompressed video.
[0021] Furthermore, the frame extraction operation in Step 1.1 includes:
[0022] Performing a frame extraction operation on the source video in the time domain and setting a time-domain frame extraction coefficient K t , indicating that only every Kth t frame in the video is retained. For example, K t = 2 means deleting half of the frames to obtain a time-domain downsampled video
[0023] Furthermore, the downsampling operation in Step 1.2 includes:
[0024] Using a bicubic interpolation downsampling operation to reduce the height and width of each frame in by a factor of Ks times (K s is the downsampling coefficient in the spatial domain) to reduce the spatial resolution.
[0025] Furthermore, the compression operation in step 1.3 includes:
[0026] Use an off-the-shelf conventional video coding algorithm to compress the spatio-temporally downsampled video to further reduce its size and generate a bitstream. Since deep learning is not involved in this encoding process and there is no need to perform time-consuming model inference, the encoding stage here only introduces negligible computational overhead comparable to using the H.265 coding standard.
[0027] Furthermore, the decompression operation in step 2.1 includes:
[0028] In the decompression stage of step 2.1, use a decoding algorithm with the same standard as the encoding stage to decompress the video bitstream to obtain a low-quality (LQ) video The resolution and frame rate of this video have been downsampled, and it also has artifacts and defects introduced by the compression algorithm.
[0029] Furthermore, the spatio-temporal super-resolution restoration operation in step 2.2 includes:
[0030] In the spatio-temporal super-resolution restoration stage of step 2.2, in order to restore to obtain Three restoration tasks of video frame interpolation, video super-resolution, and video quality enhancement need to be carried out. Next, through a joint optimization modeling process, the restoration effects of these three tasks benefit each other. Specifically, in the decoding stage of step 2.2, a spatio-temporal super-resolution network SRFI with an emerging structure design is designed to simultaneously complete the spatio-temporal super-resolution quality enhancement task for the low-quality video : perform upsampling operations on it in both the temporal and spatial dimensions (including temporal video frame interpolation and spatial video super-resolution), and remove compression artifacts to obtain a restored high-quality (HQ) video The SRFI network in the decoding (Decoder) stage takes N in low-quality frames as input and restores N out high-quality frames.
[0031] The decoding stage of step 2.2 can be further divided into:
[0032] Propagation Sub - network: To utilize the potential complementary information between different temporal and spatial domains in video clips, a forward - backward propagation recurrent neural network structure is first applied to the input frames to extract frame feature maps that combine global information. The bidirectional network consists of a forward - propagation (FP) sub - network and a backward - propagation (BP) sub - network. They have the same structure, and the difference lies in the order of image frame feeding during input. Inside the sub - network, the input video frame sequence is first input into an optical flow predictor network SpyNet, which is from non - patent literature 3 (Ranjan, et.al. "Optical Flow Estimation using a Spatial Pyramid Network.", CVPR, 2017.), for optical flow prediction of the motion between images. Then, the image frames at each moment are warped and aligned with the obtained corresponding optical flow maps. After that, two groups of residual feature blocks perform the fusion of hidden features to obtain the output feature map representation of the sub - network at that moment. The outputs of the forward and backward sub - networks will be concatenated by channel and fused using a 1×1 convolutional layer, finally obtaining the image fusion feature maps at each moment output by the sub - network.
[0033] Frame Interpolation Sub - network: The image fusion feature maps at each moment obtained in the previous step are used as the input of the feature interpolation sub - network, which will synthesize the frame feature maps lost during compression and frame extraction. In this sub - network, a variable convolution operation with a pyramid structure is adopted to effectively capture the motion cues between frames. Variable convolution is an improvement and generalization of conventional convolution, with the ability to better capture large - range motion.
[0034] Define the 3×3 convolution kernel K as: K = {(-1,-1),(-1,0),…,(0,1),(1,1)}, and the calculation formula of the conventional convolution F(·) is as follows:
[0035]
[0036] where k is the position in K, w t (·) is the weight, f(·) is the feature of the input image, and p is the initial position;
[0037] Variable convolution F deform (·) adds an additional set of two - dimensional position offsets for each convolution position, increasing the network's motion capture ability and enhancing the network's robustness:
[0038]
[0039] In the formula, compared with conventional convolution, the additional parameter Δk of variable convolution is the offset layer pre - learned by the network, and the variable convolution operation is achieved through addition compensation and bilinear interpolation operations for the final variable convolution operation.
[0040] Features from two adjacent frames are first concatenated channel-wise, and a learnable offset layer for calculating two deformable convolutions is generated using a conventional convolutional layer. Then, the two adjacent frame feature maps are respectively input into the deformable convolution operations using the corresponding offsets to obtain the synthesized centered frame feature map. Finally, 1×1 convolution is used for the final fusion to obtain the synthesized feature result.
[0041] Supramolecular network: Pixel Shuffle is adopted to improve the spatial resolution of the N out frame feature maps generated by the first two sub-networks. For frame moments with available low-resolution original feature maps, a residual block structure is additionally introduced, and the final output
[0042] The SRFI network is trained on the MFQEv2 dataset, which contains a total of 160 lossless video sequences. The resolutions of these video sequences are 2K (2048×1080), 1080p (1920×1080), 360p (640×360), CIF (352×288), etc. During testing, the HEVC standard test dataset, which is widely used to evaluate video compression-related tasks, is used. It consists of 16 video sequences with different contents and resolutions.
[0043] The SRFI network uses Peak Signal-to-Noise Ratio (PSNR) and Bits Per Pixel (BPP) to evaluate the overall performance. These two metrics are calculated in the network output video frames obtained from training on different dataset categories and H.265 compressed videos with different Constant Rate Factors (CRF). At the same time, two video coding evaluation metrics, BD-Rate and BD-PSNR, are also calculated to compare the proposed SRFI network with other schemes.
[0044] The SRFI network is implemented using the PyTorch language framework, and the FFmpeg and libx265 toolkits are used to generate bitstreams. During network training, a pre-trained model of SPyNet is used to initialize the optical flow estimator. When preprocessing the data, crops are taken from the source video and the corresponding low-quality video, and pictures of size 128×128 are used as training sample pairs. Data augmentation is achieved through random horizontal flipping and 90° rotation. The learning rates of the optical flow estimator and other networks are initially set to 2.5×10 -5 and 2×10 -4 . The Adam optimizer with β 1 =0.9, β 2 =0.99 is used for parameter update, and its update strategy is:
[0045]
[0046] where η is the learning rate, θk is the weight parameter for the k-th iteration.
[0047] Adopt a cosine annealing scheme to periodically change the learning rate according to the training process. The learning rate for the k-th time is:
[0048]
[0049] where η 0 is the initial learning rate (set to 2×10 -4 ), and the total number of iterations T is 300000.
[0050] During training and evaluation, a total of four different constant rate factors (CRFs) were selected, namely 20, 25, 30, 35, and a separate model was trained for each CRF value for different compression settings.
[0051] The beneficial effects of the present invention are as follows:
[0052] 1) A novel video compression method is proposed, which combines the latest progress of spatio-temporal super-resolution tasks to improve the performance of traditional video codecs. This is the first practice of applying a deep neural network related to spatio-temporal super-resolution tasks to video compression.
[0053] 2) An efficient spatio-temporal super-resolution network SRFI is proposed to improve video decoding. The SRFI network can automatically capture deep motion features in low-resolution, low-frame-rate compressed video sequences and efficiently complete tasks such as feature alignment bending, frame feature interpolation, upsampling super-resolution, and quality enhancement. The comparison results with the prior art prove the advantages of this network model in terms of speed and accuracy.
[0054] 3) Experimental results show that the proposed STSR-VC (Space-Time Super-Resolution Video Compression) video compression and restoration method has both the high efficiency of traditional codecs (the computational overhead at the encoder end is approximately 0) and the strong robustness of learnable DNNs (superior to traditional H.265 codecs and other existing DNN-based codecs), achieving better bit transmission efficiency and being more adaptable to low-bit transmission scenario conditions. It has high value in helping the Internet content distribution industry save transmission bandwidth and improve efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is the implementation flowchart of an efficient spatio-temporal super-resolution video compression and restoration method based on deep learning according to an embodiment of the present invention;
[0056] Figure 2 is the structural diagram of the SRFI spatio-temporal super-resolution network in the decoding stage according to an embodiment of the present invention;
[0057] Figure 3 It is a structural comparison diagram of the SRFI network in an embodiment of the present invention and the existing FISR spatio-temporal super-resolution network;
[0058] Figure 4 It is a line graph of the BPP-PSNR index comparison of the SRFI in an embodiment of the present invention for spatio-temporal super-resolution compression and restoration tasks and other existing methods;
[0059] Figure 5 It is a comparison diagram of the detailed effect of the results of the SRFI in an embodiment of the present invention for spatio-temporal super-resolution compression and restoration tasks and other existing methods. Detailed implementation manners
[0060] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0061] As Figure 1 shown is the implementation process of an efficient spatio-temporal super-resolution video compression and restoration method based on deep learning provided by the present invention. This method follows the currently widely used video distribution process, including an encoding (Encoder) stage and a decoding (Decoder) stage.
[0062] The task of the encoding (Encoder) stage is to encode and compress a given source video (a total of N video frames, with height and width being H and W respectively, and the number of channels being C) so as to minimize the bit overhead during transmission. Applying the encoding algorithm of the conventional compression coding standard to the source video is a direct solution, and these codecs use a series of efficient handcrafted modules to reduce the redundancy in the video. Considering the recent progress in video restoration tasks in deep learning, before the bitstream generation, downsampling operations in the spatial and temporal dimensions can be first performed to further reduce the number of transmitted bits without affecting the quality, and the downsampled video will be restored by spatio-temporal super-resolution operations in the decoding stage.
[0063] In the encoding stage, first, frame extraction is performed on the source video in the time domain, and a time-domain frame extraction coefficient K t is set, indicating that only every K t th frame in the video is retained. For example, K t = 2 means deleting half of the frames to obtain a time-domain downsampled video Then, using bicubic interpolation downsampling operation, the height and width of each frame in are reduced by K s times (K s(is the spatial downsampling factor) to reduce the spatial resolution. Finally, use an off-the-shelf conventional video coding algorithm to compress the spatio-temporally downsampled video to further reduce its size and generate a bitstream. In an actual system, this step can use off-the-shelf conventional video compression coding algorithms such as H.264 and H.265. Since deep learning is not involved in this coding process and there is no need to perform time-consuming model inference, the coding stage only introduces negligible computational overhead comparable to using traditional coding standard algorithms.
[0064] The task of the decoding stage is that after receiving the video bitstream from the coding stage, this module needs to generate the final restored video whose visual quality needs to be as close as possible to the source video and have the same picture size and video frame rate. The decoding stage first uses a decoding algorithm of the same standard as the coding stage to decompress these bits to obtain a low-quality (LQ) video whose resolution and frame rate have been downsampled, and various artifacts and defects introduced by the H.265 compression algorithm are also attached. In order to restore three restoration tasks of video frame interpolation, video super-resolution, and video quality enhancement need to be performed respectively. However, considering the internal correlation of these three tasks, their restoration effects can benefit from each other through a joint optimization modeling process. Specifically, in the decoding stage, a spatio-temporal super-resolution network SRFI with an emerging structure design is designed to simultaneously complete the spatio-temporal super-resolution quality enhancement task for the low-quality video : perform upsampling operations on it in both the temporal and spatial dimensions (including temporal video frame interpolation and spatial video super-resolution), and remove compression artifacts to obtain the restored high-quality (HQ) video
[0065] As Figure 2 shown, the SRFI network of the decoding (Decoder) module takes N in low-quality frames as input and restores N out high-quality frames. In the implementation structure, it can be mainly divided into three processing sub-networks: "propagation", "frame interpolation", and "super-resolution".
[0066] Figure 2(b) The network structure shown is a schematic diagram of a single forward propagation part module in the "propagation" sub-network of the SRFI network. To utilize the potential complementary information between different temporal and spatial domains in video clips, a forward and backward propagation recurrent neural network structure is first applied to the input frames to extract frame feature maps that combine global information. The bidirectional network consists of a forward propagation (FP) sub-network and a backward propagation (BP) sub-network, which have the same structure, except for the order of image frame feeding during input. Inside the sub-network, the input video frame sequence is first input into a SpyNet network for optical flow prediction of the motion between images, and then the image frames at each moment are bent and aligned with the corresponding obtained optical flow maps. After that, two groups of residual feature blocks perform the fusion of hidden features to obtain the output feature map representation of the sub-network at that moment. The outputs of the forward and backward sub-networks will be concatenated by channel and fused using a 1×1 convolutional layer to finally obtain the image fusion feature maps at each moment output by the sub-network.
[0067] Figure 2 (c) The network structure shown is the "frame interpolation" sub-network in the SRFI network: The image fusion feature maps at each moment obtained in the previous step are used as the input to the feature interpolation sub-network, which will synthesize the frame feature maps lost during compression and frame extraction. In this sub-network, a variable convolution operation with a pyramid structure is adopted to effectively capture the motion cues between frames. Variable convolution is an improvement and generalization of conventional convolution, with the ability to better capture large-scale motion.
[0068] Define the 3×3 convolution kernel K as: K = {(-1, -1), (-1, 0), …, (0, 1), (1, 1)}, and the calculation formula of the conventional convolution F(·) is as follows:
[0069]
[0070] where k is the position in K, w t (·) is the weight, f(·) is the feature of the input image, and o is the initial position;
[0071] The variable convolution F deform (·) adds a set of additional two-dimensional position offsets to each convolution position, increasing the network's motion capture ability and enhancing the network's robustness:
[0072]
[0073] In the formula, compared with the conventional convolution, the additional parameter Δk of the variable convolution is the offset layer pre-learned by the network, and the final variable convolution operation is achieved through addition compensation and bilinear interpolation operations.
[0074] Features from two adjacent frames are first concatenated by channel, and a learnable offset layer for calculating two deformable convolutions is generated using a conventional convolutional layer. Then, the two adjacent frame feature maps are respectively input into the deformable convolution operations using the corresponding offsets to obtain a synthesized centered frame feature map. Finally, 1×1 convolution is used for the final fusion to obtain the synthesized feature result.
[0075] The "super-resolution" sub-network: adopts pixel recombination operation (Pixel Shuffle) to improve the spatial resolution of the N out frame feature maps generated by the first two sub-networks. For frame moments with available low-resolution original feature maps, a residual block structure is additionally introduced, and the final output
[0076] As Figure 3 shown, it is a comparison of the structural differences between the SRFI (Super-Resolution Frame-Interpolation) network and the previous traditional FISR (Frame-Interpolation Super-Resolution) architecture spatio-temporal super-resolution network. The previously existing STSR networks often follow the "FI-SR" architecture, that is, given N input frames, first perform frame feature interpolation to generate (n - 1) (assuming K t = 2) new feature representations, and then input the overall (2n - 1) feature maps into the super-resolution network, which includes a number of computationally intensive operations such as flow estimation, motion alignment, and feature propagation. Although this architecture is intuitive, it mainly has two limitations:
[0077] 1. Since the newly generated (n - 1) feature maps are from adjacent frames and do not contain new useful information, feeding them into the super-resolution sub-network may not be beneficial to the overall performance;
[0078] 2. The "FI-SR" architecture has an inherent complexity. The number of input feature maps in the super-resolution sub-network part is (2n - 1), which will lead to a significant increase in the computational overhead of the entire module and reduce the inference speed. At the same time, the excessive memory consumption also limits the upper limit of the number of input frames in the decoding stage.
[0079] To address these two limitations, a simple and effective "SR-FI" architecture can be adopted. First, video frames are fed into the propagation network. Then, intermediate features containing global and high-resolution information are used for feature interpolation to obtain (n - 1) new features. Here, the SR network only receives n frames to perform high-cost operations, saving nearly half of the computational cost compared to the "FI-SR" architecture that inputs (2n - 1) frames. This not only improves the network's inference speed but also allows the overall method to input more video frames each time. It is worth noting that there is no separate subnet for quality enhancement in the SRFI network because the SRFI network can automatically learn to remove compression artifacts and complete the quality enhancement task through end-to-end training.
[0080] During the training process of the proposed SRFI network, to generate input samples, first, the complete encoding stage processing is performed to generate a bitstream, and then the bitstream is decoded into low-quality video frames using the same decoding algorithm as in the encoding stage. When updating the neural network's backpropagation parameters, the restored reconstructed video frames and the corresponding source video frames adopt the Charbonnier penalty function as the loss function:
[0081]
[0082] where the value of ∈ is set to 1×10 during training -3 .
[0083] Comparing the SRFI network with other deep learning video compression and restoration methods in terms of metrics, Table 1 and Figure 4The BD-Rate / BD-PSNR quantization results and BPP-PSNR curves of different methods on the HEVC dataset are shown respectively. It can be seen from the table that taking the H.265 scheme of traditional coding and decoding as the baseline performance, since the disclosed inventive method combines the H.265 scheme and the SRFI network based on deep learning at the same time, it achieves a 10.14% reduction in BD-Rate and a 0.48dB gain in BD-PSNR compared with the H.265 baseline. This shows that the proposed emerging spatio-temporal super-resolution video compression and restoration method of the present invention has higher coding performance than the H.265 scheme. SRFI achieves results comparable to the baseline H.265 codec in high bitrate scenarios and is also superior to all similar schemes in low bitrate scenarios. In addition, the overall scheme only incurs negligible computational overhead at the encoding end and can run at a speed of 70 frames per second to decode test videos with a size of 352×288, achieving an inference speed improvement of nearly 2 times compared with Non-Patent Document 4 (Lu, et.al. "DVC: An End-To-End Deep Video Compression Framework.", CVPR, 2019.).
[0084]
[0085] Table 1 Comparison chart of BD-Rate (%) / BD-PSNR (PSNR) metrics for the spatio-temporal super-resolution compression and restoration task of SRFI in an embodiment of the present invention and other existing methods
[0086] Figure 5 The detailed comparison diagram of the restored video frames on the HEVC dataset is shown. It can be seen that the video frames compressed only by H.265 are still severely distorted by various artifacts after restoration (for example, blurring in the first row of images and ringing effects in the third row of images). Under similar BPP (Bit Per Pixel) metrics, although other existing deep learning video compression methods can reduce artifacts to a certain extent, the generated frames still have problems of excessive blurring and detail loss. Compared with these existing methods, the SRFI network restores more accurate details. For example, in the first row of images, the texture of the vent can be clearly identified, and there is no obvious blurring around it compared with the results of other methods; in the second row of images, the result obtained by the SRFI network is also the closest to the real result.
[0087] Those skilled in the art can easily make various changes and modifications based on the written description, drawings, and claims provided by the present invention without departing from the spirit and scope of the present invention defined by the claims. Any modification or equivalent change made to the above embodiments based on the technical spirit and essence of the present invention falls within the protection scope defined by the claims of the present invention.
Claims
1. An efficient spatio-temporal super-resolution video compression and restoration method based on deep learning, characterized in that, it includes: Step 1, encoding stage, encoding and compressing the given source video, including: Step 1.1, extracting frames from the source video; Step 1.2, downsampling the source video; Step 1.3, compressing the source video; Step 2, decoding stage, after receiving the video bitstream from the encoding stage, generating the final restored video, including: Step 2.1, decompressing the video processed in Step 1; Step 2.2, performing spatio-temporal super-resolution restoration on the decompressed video, and the spatio-temporal super-resolution restoration operation includes: Design a spatio-temporal super-resolution network SRFI, and complete the spatio-temporal super-resolution quality enhancement task for low-quality videos through this spatio-temporal super-resolution network SRFI The spatio-temporal super-resolution quality enhancement task includes: for low-quality videos Perform upsampling operations in both the time domain and the spatial domain. The upsampling operations include temporal video frame interpolation and spatial video super-resolution, and remove compression artifacts to obtain a restored high-quality HQ video The spatio-temporal super-resolution network SRFI takes N in low-quality frames as input and restores N out high-quality frames; The spatio-temporal super-resolution network SRFI includes a propagator network, an interpolation network, and a super-resolution network; The propagator network applies a forward and backward propagation recurrent neural network structure to the input frame to extract a frame feature map that combines global information; the bidirectional network consists of a forward propagation subnet and a backward propagation subnet, and the structures of the forward propagation subnet and the backward propagation subnet are the same, except for the order of image frame feeding during input. Inside the subnet, first, the input video frame sequence is input into an optical flow predictor network to predict the optical flow of the motion between images, and then the image frames at each moment are warped and aligned with the obtained corresponding optical flow map. After that, two groups of residual feature blocks perform the fusion of hidden features to obtain the output feature map representation of the propagator network at this moment; the outputs of the forward and backward subnets will be concatenated by channel and fused using a 1×1 convolutional layer to finally obtain the image fusion feature map of each moment output by the propagator network; The interpolation network uses the image fusion feature map of each moment obtained by the propagator network as the input of the feature interpolation subnet, and this network will synthetically output the frame feature map lost during compression and frame extraction; in this interpolation network, variable convolution operations with a pyramid structure are used to capture the motion cues between frames; Supramolecular network, which uses pixel recombination operations to improve the spatial resolution of the N-frame feature maps generated by the propagator network and the interpolation frame network. For frame moments with available low-resolution original feature maps, a residual block structure is additionally introduced to finally output high-quality HQ video out For frame moments with available low-resolution original feature maps, a residual block structure is additionally introduced to finally output high-quality HQ video 2. The efficient spatio-temporal super-resolution video compression and restoration method based on deep learning according to claim 1, characterized in that, The frame extraction operation in Step 1.1 includes: For a given source video where N, H, W, and C represent the batch number, frame height, frame width, and number of channels of the input source video respectively, set a temporal frame extraction coefficient K t , indicating that only every K t th frame in the video is retained to obtain a temporally downsampled video 3. The efficient spatio-temporal super-resolution video compression and restoration method based on deep learning according to claim 2, characterized in that, The downsampling operation in Step 1.2 includes: Use bicubic interpolation downsampling operation to reduce the height and width of each frame in by a factor of K, where K s is the spatial downsampling coefficient, to obtain the spatio-temporal downsampled video s 4. The efficient spatio-temporal super-resolution video compression and restoration method based on deep learning according to claim 3, characterized in that, The compression operation in Step 1.3 includes: Compress the spatio-temporal downsampled video using existing off-the-shelf conventional video coding algorithms to further reduce its size and generate a bitstream.
5. The efficient spatio-temporal super-resolution video compression and restoration method based on deep learning according to claim 4, characterized in that, The decompression operation in Step 2.1 includes: Decompress the video bitstream using the decoding algorithm with the same standard as in the encoding stage to obtain a low-quality video
Citation Information
Patent Citations
Video description method based on space-time super-resolution and electronic equipment
CN114549317A
End-to-end video compression method and system based on deep learning, and storage medium
WO2021164176A1