Compressed Video Super-Resolution Based on Deep Feature Fusion Network

Through the deep feature fusion network, the compressed video is processed using hybrid convolution and ordinary differential equation modules, which solves the problem of low reconstruction quality of compressed video and achieves high-quality super-resolution reconstruction effect.

CN115409695BActive Publication Date: 2025-07-04SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110579150.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-26
Publication Date
2025-07-04
Estimated Expiration
2041-05-26

AI Technical Summary

Technical Problem

The prior art is difficult to effectively restore high-resolution compressed video, especially when hardware cost, storage capacity and transmission bandwidth are limited. The compressed noise is highly correlated with the video content, and direct reconstruction is easy to amplify noise or lose information, reducing super-resolution performance.

Method used

Using a deep feature fusion network method, low-dimensional features are extracted using 2D/3D hybrid convolutional blocks and residual networks, combined with the ordinary differential equation module to reduce compression traces, and super-resolution reconstruction is carried out through adaptive channels and pixel attention modules to construct an effective super-resolution method for compressed video.

Benefits of technology

Improve the reconstruction quality of compressed videos, clear and natural edges and more details, the subjective and objective evaluation results show that they are better than traditional methods, and the PSNR and SSIM indicators are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115409695B_ABST
    Figure CN115409695B_ABST
Patent Text Reader

Abstract

The present invention discloses a super-resolution method for compressed videos based on a deep feature fusion network. The method mainly includes the following steps: for the input low-resolution compressed video sequence, five consecutive frames are used as the input of the network, and a hybrid convolution block and a residual block are used to extract low-dimensional feature information; the ordinary differential equation block in the restoration module is used to reduce compression artifacts and obtain high-dimensional feature information; different-dimensional feature maps are fused and input into the reconstruction module, and super-resolution reconstruction is completed by using an adaptive channel attention and pixel attention module and a sub-pixel convolutional layer for upsampling to obtain high-resolution target video frames; training samples are constructed in the video dataset, the network is trained, and the final model is obtained. The method of the present invention is used to reconstruct a low-resolution compressed video into a high-resolution video, and is an effective super-resolution reconstruction method for compressed videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of compressed video super-resolution reconstruction, and specifically to a method for compressed video super-resolution based on a deep feature fusion network, belonging to the field of image processing. Background Art

[0002] The goal of super-resolution is to recover high-resolution images or videos from observed low-resolution images or videos. It has wide applications in some fields with high requirements for image or video resolution and details, such as medical imaging, remote sensing imaging, and satellite detection. Currently, most commonly used video super-resolution algorithms are for degraded video frames after downsampling. However, due to limitations in aspects such as hardware cost, storage capacity, transmission bandwidth, and response time, security and traffic monitoring systems as well as Internet applications usually can only obtain low-resolution compressed videos, and the further degradation of video quality also increases the difficulty of restoration and reconstruction. In addition, the noise introduced by compression usually has a strong correlation with the content of the video frames themselves. If directly reconstructing video frames with two types of degradations (compression and downsampling) or simply removing compression artifacts before super-resolution, it often amplifies noise, loses important information, or reduces super-resolution performance. Summary of the Invention

[0003] The present invention uses a convolutional neural network to extract and fuse spatio-temporal information features and an ordinary differential equation network to reduce compression artifacts, and then constructs an effective method for compressed video super-resolution.

[0004] The compressed video super-resolution based on a deep feature fusion network proposed by the present invention mainly includes the following operation steps:

[0005] (1) For the input low-resolution compressed video sequence, take five consecutive frames as the input of the network, and then use a 2D / 3D hybrid convolutional block and a residual network to initially extract low-dimensional feature information;

[0006] (2) Input the obtained low-dimensional features into a restoration module, use an ordinary differential equation (ODE) module to reduce compression artifacts, and calculate the loss between the output after passing through a convolutional layer and the uncompressed low-resolution target video frame;

[0007] (3) Fuse the original feature map and the output feature maps of steps (1) and (2) with different dimensions together to obtain feature information;

[0008] (4) Input the output results of steps (3) and (2) into a reconstruction module, and use an adaptive channel attention and pixel attention module and a sub-pixel convolutional layer for upsampling to complete super-resolution reconstruction, obtaining the final high-resolution target video frame;

[0009] (5) Construct training sample pairs in the video dataset, train the network parameters, and complete the network training and obtain the final model when the loss function of the reconstructed high-resolution video frame calculation model is minimized. Description of the Drawings

[0010] Figure 1 is a block diagram of the compressed video super-resolution based on the deep feature fusion network of the present invention.

[0011] Figure 2 is a block diagram of the ordinary differential equation block in the network of the present invention.

[0012] Figure 3 is a comparison diagram of the reconstruction results of the test video "BQMall" by the present invention and six other methods, where (a) is the original high-resolution image, (b) is the reconstruction result of bicubic interpolation, (c) to (g) are the reconstruction results of methods 1 to 6, and (h) is the reconstruction result of the present invention.

[0013] Figure 4 is a comparison diagram of the reconstruction results of the test video "PartyScene" by the present invention and six other methods, where (a) is the original high-resolution image, (b) is the reconstruction result of bicubic interpolation, (c) to (g) are the reconstruction results of methods 1 to 6, and (h) is the reconstruction result of the present invention. Detailed Embodiment

[0014] The present invention will be further described below with reference to the accompanying drawings:

[0015] Figure 1 In, the compressed video super-resolution based on the deep feature fusion network can be specifically divided into the following five steps:

[0016] (1) For the input low-resolution compressed video sequence, take five consecutive frames as the input of the network, and then use the 2D / 3D hybrid convolution block and the residual network to initially extract low-dimensional feature information;

[0017] (2) Input the obtained low-dimensional features into the restoration module, use the ordinary differential equation (ODE) module to reduce the compression traces, and calculate the loss between the output after passing through a convolutional layer and the uncompressed low-resolution target video frame;

[0018] (3) Fuse the original feature map and the output feature maps of steps (1) and (2) with different dimensions together through the fusion block to obtain feature information;

[0019] (4) Input the output results of steps (3) and (2) into the reconstruction module, and use the adaptive channel attention and pixel attention modules and the sub-pixel convolutional layer for upsampling to complete the super-resolution reconstruction and obtain the final high-resolution target video frame;

[0020] (5) Train the network parameters. When the loss function of the reconstructed high-resolution video frame calculation model is minimized, complete the network training and obtain the final model;

[0021] Specifically, in the step (1), perform bicubic downsampling on the original video sequence to obtain a low-resolution video sequence. Use HM 16.0 to encode the low-resolution video sequence with quantization parameters (QP) of 32, 37, 42, and 47 to obtain a low-resolution compressed video sequence. Take five consecutive low-resolution compressed video frames as the input of the network, and then use a 2D / 3D hybrid convolution block and a residual network to initially extract low-dimensional feature information.

[0022] In the step (2), the constructed ordinary differential equation module structure is as Figure 2 shown.

[0023] Mathematically, the definition of an ordinary differential equation (ODE) is:

[0024] dy / dx = f(x, y)

[0025] where x and y are the independent variable and the dependent variable respectively. The mapping relationship of a dynamic system can be expressed by an ODE as:

[0026] Ψ(y0, x) = y(x:y0)

[0027] where Ψ is the mapping relationship and y0 is the initial state of the input feature. Assume that p(y0) is the distribution of the input feature y0 in the domain Ω. If the process of compensating for the high-frequency information of the video is regarded as a dynamic system, the solution is to minimize the following equation:

[0028] L = ∫ Ω / Ψ(y0, x) - y / dp(y0)

[0029] When the system is non-linear, it is often difficult to describe the mapping relationship with a simple formula in many cases. Therefore, when solving the problem, differential is usually approximated by difference, and the simplest method is the forward Euler method. Divide the interval [0, T] into N equal parts, h is called the step size, x n = n*h (n = 0, 1, 2,..., N) is called the node, and the approximate value of f(x, y) can be expressed as f(x, y) ≈ yn+1 - yn / h. Therefore, the formula definition of the forward Euler method is:

[0030] f(x, y) ≈ yn+1 - yn / h

[0031] When representing the input of the y n th residual block, and y n+1 represents the output, the residual block has a similar expression form:

[0032] yn+1 = y n + S(y n )

[0033] S(y n ) = h * f(x n , y n )

[0034] where S(·) represents the residual operation. The above forward Euler algorithm is a simple first-order numerical method, which is unstable and has low accuracy. Therefore, it is changed to the second-order Velocity Verlet algorithm, which is expressed as:

[0035] y n+2 = y n + h * (y' n + y' n+2 )

[0036] where h = 1.

[0037] In order to obtain a specific block structure and maintain its flexibility, the above second-order Velocity Verlet algorithm process is divided into three formulas to form a block structure, which can be expressed as:

[0038] y n+1 = y n + 2 * y' n

[0039] y n+2 = y n + 2 * y' n+1

[0040] y n+2 = y n + (y' n + y' n+2 )

[0041] where the derivative process is explained as passing through a parametric rectified linear unit PReLU and a 3×3 convolutional layer.

[0042] In step (3), the block structure of the fusion block built is as Figure 1 shown. The extraction of different-dimensional features simultaneously utilizes the intra-frame spatial information and the inter-frame temporal information. The original feature map, the low-dimensional feature map, and the high-dimensional feature Figure 3 feature maps of different depths are fused together through the fusion block, enhancing the spatio-temporal information and effectively preventing the loss of detailed information. Then, the fused feature map is concatenated with the output of the restoration module as the input of the reconstruction module.

[0043] In step (4), the channel attention mechanism rescales the extracted features according to the importance of channels, that is, different weights are assigned to different channels, which helps to pay more attention to important information, while the pixel attention mechanism generates attention coefficients for all pixels in the feature map. By using the adaptive channel attention and pixel attention modules, the intermediate information characteristics of high-frequency information reconstruction can be effectively obtained, and the quality of the reconstruction result can be improved.

[0044] In step (5), the input continuous video frame sequence is input into the network model trained in step (4) to obtain the super-resolution reconstruction result. To better illustrate the effectiveness of the present invention, the "BQMall" and "PartyScene" test sets are selected from common test videos. We perform bicubic downsampling on the original video sequence to obtain a low-resolution video sequence, and use HM 16.0 to encode the low-resolution video sequence at quantization parameters (QP) of 32, 37, 42, and 47 to obtain a low-resolution compressed video sequence. In the experiment, bicubic interpolation Bicubic and two "one-step" methods and three other "two-step" compressed video super-resolution methods are selected for comparison.

[0045] The selected algorithms are:

[0046] Algorithm 1: The method proposed by Guan et al., reference "MFQE 2.0: A new approach for multi-frame quality enhancement on compressed video. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2019".

[0047] Algorithm 2: The method proposed by Wang et al., reference "Deep video super-resolution using HR optical flow estimation. IEEE Transactions on Image Processing 29: 4323-4336, 2020".

[0048] Algorithm 3: The method proposed by Zhao et al., reference "Efficient image super-resolution using pixel attention. arXiv preprint arXiv:2010.01073, 2020."

[0049] Algorithm 4: The method proposed by Ho et al., reference "Down-sampling based video coding with degradation-aware restoration-reconstruction deep neural network. In: International Conference on Multimedia Modeling. Springer, Cham. 99-110, 2020".

[0050] Algorithm 5: The method proposed by Ho et al., reference "RR-DnCNN v2.0: Enhanced Restoration-Reconstruction Deep Neural Network for Down-Sampling-Based Video Coding. IEEE Transactions on Image Processing 30: 1702-1715, 2021".

[0051] The comparison compression video super-resolution reconstruction methods are as follows:

[0052] Method 1: Algorithm 1 + Bicubic interpolation

[0053] Method 2: Algorithm 1 + Algorithm 2

[0054] Method 3: Algorithm 1 + Algorithm 3

[0055] Method 4: Algorithm 4

[0056] Method 5: Algorithm 5

[0057] Experiment 1, use Bicubic interpolation, Methods 1 to 5, and the present invention respectively to perform 2-fold reconstruction on the low-resolution compressed test video obtained after degradation. The super-resolution reconstruction results are respectively shown by Figures 3 to 4 as shown. The objective evaluation results of the reconstruction results are shown in Table 1. PSNR (Peak Signal to Noise Ratio, unit: dB) and SSIM (Structure Similarity Index) are respectively used to evaluate the reconstruction effect. The higher the PSNR / SSIM value, the better the reconstruction effect.

[0058] As can be seen from Table 1, the present invention has obtained higher PSNR and SSIM. From Figure 3 and Figure 4It can be seen that the reconstruction result of the present invention has clear and natural edges, showing more details, while the reconstruction result of the contrast algorithm has certain artifacts and relatively blurred edges in the subjective visual effect. In summary, compared with the comparison method, the reconstruction result of the present invention has great advantages in both subjective and objective evaluations. Therefore, the present invention is an effective method for compressed video super-resolution reconstruction.

[0059] Table 1

[0060]

Claims

1. Compression video super-resolution method based on deep feature fusion network, characterized in that It includes the following steps: Step 1: Low-dimensional feature extraction; specifically, for the input low-resolution compressed video sequence, five consecutive frames are used as the input of the network, and then the 2D / 3D hybrid convolutional block and the residual network are used to initially extract low-dimensional feature information; Step 2: Restoration module; specifically, the obtained low-dimensional features are used as the input of the restoration network. The second-order Velocity Verlet algorithm is implemented through the ordinary differential equation (ODE) block for dynamic system modeling to reduce compression artifacts, and the loss between the high-dimensional feature output and the uncompressed low-resolution target video frame is calculated after passing through a convolutional layer; Step 3: Different-dimensional feature extraction; specifically, the original feature map, low-dimensional feature map, and high-dimensional feature map are fused together through a fusion block to obtain different-dimensional feature information; Step 4: Reconstruction module; specifically, the output results of Step 3 and Step 2 are combined as the input of the reconstruction module. The super-resolution reconstruction is completed by using the adaptive channel attention, pixel attention module, and sub-pixel convolutional layer upsampling to obtain the final high-resolution target video frame; Step 5: Construct training sample pairs in the video dataset, train the network parameters, and complete the network training and obtain the final model when the loss function of the reconstructed high-resolution video frame calculation model is minimized.

2. The compressed video super-resolution based on the deep feature fusion network according to claim 1, wherein In Step 2, an ordinary differential equation (ODE) block is used to reduce compression artifacts. Specifically, leveraging the characteristic of high similarity between the input and output of the video decompression process, the theory of ordinary differential equations is introduced from the perspective of dynamic systems. The ordinary differential equation is used to replace the residual block, the conventional first-order forward Euler algorithm is changed to the second-order Velocity Verlet algorithm, and the algorithm process is divided into three formulas to form a block structure. The process is expressed as: y n+1 = y n + 2 * y' n , y n+2 = y n + 2 * y' n+1 , y n+2 = y n +(y' n + y' n+2 ), where y n represents the input of the ODE block, y n+1 and the first y n+2 represent the intermediate process outputs. The derivative process is explained as passing through a parametric rectified linear unit (PReLU) and a 3×3 convolutional layer.

3. The compressed video super-resolution based on the deep feature fusion network according to claim 1, wherein In Step 3, the fusion block is used to fuse the original feature map and low-dimensional feature map and then fuse with the deep features to obtain different-dimensional feature information, prevent a large amount of detailed information from being lost, and improve the subsequent reconstruction quality.

Citation Information

Patent Citations

  • Rapid space-time residual attention video super-resolution reconstruction method

    CN111028150A

  • Compressed video quality enhancement method based on attention mechanism and time dependence

    CN111031315A