Method for improving quality of space-time domain compressed video by using dense network
Through dense network extraction and optimization of space-time domain features, the quality problem of HEVC compressed video is solved, especially in the case of violent movement or rapid scene change, and better video quality improvement effect is achieved.
Patent Information
- Application Number
- CN202410148262.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-02
- Publication Date
- 2025-08-05
AI Technical Summary
Existing video compression technologies such as HEVC, while increasing file size reduction, are often accompanied by problems such as loss of subtle details, blurred image and inaccurate motion estimation, which affects video quality, especially in video sequences with severe motion or rapid scene change.
The spatial information of the current frame is extracted using a dense network, and multi-frame video is aligned through deformable convolution, combining the space-time adaptive fusion module and dense connection structure, the space-time domain characteristics are optimized, and finally a video frame with improved quality is generated in the reconstruction module.
Effectively suppress the video compression effect and improve the video quality, especially in the case of violent movement or rapid scene change, to achieve better visual effects and PSNR improvement.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the technical problem of improving the quality of compressed videos in the field of video coding, and specifically to a method of extracting spatial information in the current frame using a dense network, optimizing spatial and temporal fusion features typically obtained by deformable convolution in a spatiotemporal adaptive fusion module, and then sending the optimized spatial and temporal features to a reconstruction module with a dense connection structure to obtain an enhanced compressed video. Background Art
[0002] In today's digital age, video has become an important medium for communication, entertainment, and information acquisition. However, with the widespread use of high-quality videos such as HD and UHD, the enormous size of video files has brought challenges in storage, transmission, and processing. To address these issues, video compression technology has emerged. However, while compression standards such as High Efficiency Video Coding (HEVC) have achieved significant success in reducing file sizes, they have also raised new issues, particularly video quality. HEVC compression is often associated with quality issues such as loss of subtle details, image blur, and inaccurate motion estimation, which can affect the viewing experience and information delivery in certain application scenarios. Therefore, there is an urgent need to research compressed video quality enhancement methods to improve video quality while maintaining a small file size. This research can explore innovative technologies in fields such as machine learning, deep learning, and image processing to compensate for the quality loss caused by compression.
[0003] Early research focused on single-frame video enhancement. These methods exploited inter-pixel correlations by modifying filter coefficients and adjusting filtering methods to modify pixel values, resulting in an enhanced image closer to the original. Foi et al. proposed an image filtering method based on shape-adaptive DCT transform (SA-DCT). Dong et al. pioneered a convolutional network-based approach for JPEG image deblocking. However, single-frame approaches fail to exploit temporal correlations between adjacent frames in a video sequence, and therefore often achieve limited results. Yang et al. proposed MFQE, which uses optical flow estimation to align different frames, achieving the earliest multi-frame video quality enhancement. Subsequently, Deng et al. utilized deformable convolutions to achieve multi-frame alignment, addressing the difficulty of training and inaccuracy of optical flow estimation. Due to the limitations of deformable convolution alignment, the performance of feature fusion is constrained by the accuracy of the alignment offset. Luo et al. proposed using 3D convolutions to obtain a larger receptive field, achieving higher performance. Furthermore, multi-frame quality enhancement often yields unsatisfactory results for video sequences with scene changes or intense motion, and further exploration is needed for the quality improvement of compressed video. Summary of the Invention
[0004] Aiming at the task of improving the quality of compressed videos, the present invention aims to optimize spatiotemporal features.
[0005] The basic idea of this invention is to further utilize spatial information based on the temporal and spatial information extracted by deformable convolution, a common method for multi-frame quality improvement, to improve the quality of reconstructed video. At the same time, it solves the problem that the multi-frame method has a poor improvement effect when the motion is intense. The method mainly includes four steps, namely:
[0006] Step 1: Input the low-quality video obtained by HEVC compression into the network, align the frame to be enhanced with the three adjacent frames before and after it through a general deformable convolution method, and extract the initial spatiotemporal fusion features;
[0007] Step 2: We use our dense network to extract more layers of spatial feature information from the frame to be enhanced, especially focusing on texture-rich areas such as object edges.
[0008] Step 3: The initial spatiotemporal fusion features and spatial features are input into the spatiotemporal information adaptive fusion module together, and the optimized spatiotemporal information is obtained through the introduction of a series of attention mechanisms and convolutional layers;
[0009] Step 4: The spatiotemporal information obtained above is sent to the reconstruction module, and the network obtains the best residual, which is added to the compressed frame to obtain the reconstructed video frame.
[0010] The specific process is as follows:
[0011] (1) Construct a training dataset, select the dataset used in the multi-frame quality enhancement (MFQE) method, select four QP values of 22, 27, 32, and 37 for compression through HEVC, obtain the original video and its corresponding compressed video sequences under the four QPs, and select random cropping to 128×128 blocks, flipping and rotation to improve generalization;
[0012] (2) The overall framework of the present invention is shown in the attached Figure 1 As shown in the figure, the current compressed frame, the three forward frames, and the three backward frames are first input into the deformable convolution module to obtain the initial spatiotemporal fusion features. Then, the spatial features are extracted from the current compressed frame. The features obtained from the two branches are passed through the fusion module in the network to obtain the optimized fusion features. Finally, the reconstruction module outputs the reconstructed frame with improved quality.
[0013] (3) During the training phase, the present invention selects the original video and a compressed video of a selected QP for training. The compressed videos at other QPs are also trained separately with the original video, resulting in a total of four quality improvement models at different QPs. The HEVC standard test sequence is used for testing. After passing the quality improvement network, the corresponding enhanced video sequence is obtained, and the peak signal-to-noise ratio (PSNR) indicator is used to verify the effectiveness of the method.
[0014] When extracting the initial spatio-temporal fusion features by the deformable convolution method in process (2), we choose dcnv2 which is widely used at present. Its structure is that a layer of convolution transfers multiple original images to be aligned to the feature domain, then the U-net is used to obtain the offset values of the convolution kernel in the horizontal and vertical directions. Finally, multiple original images and offset values pass through a layer of deformable convolution to obtain the initial spatio-temporal fusion features;
[0015] The structure of the spatial feature extraction module is as follows: First, a layer of convolution transfers the frame to be enhanced to the feature domain; then three consecutive feature propagation branches are connected, and the internal of each branch adopts the architecture of a dense block. Specifically, the dense block has five convolutional layers with the output channels set to 32. The number of channels of the initial convolutional layer is set to 32, and each subsequent convolutional layer adds the outputs of each previous convolutional layer, that is, the channel growth rate is set to 32. Each convolution except the last one is followed by a linear pooling layer. Then the number of input channels of the last convolutional layer is 160, and the number of output channels is 32. And a residual structure is introduced to add the input of the dense block and the output of the last convolutional layer, allowing more layers to be stacked without causing the problem of gradient disappearance; In addition to the residual structure inside each dense block, we also make a residual of the original input and the final output of the three dense blocks;
[0016] To obtain the optimized spatio-temporal fusion features, we first need to process the spatial features. We choose deformable spatio-temporal attention (DSTA), a spatial attention module used to judge the useful information in the extracted spatial information. Then we perform channel-wise fitting of the spatial features and the initial fusion features, and then introduce a channel attention mechanism to distinguish useful and useless channel information. Finally, a convolutional layer with a kernel size of 1×1 is used to reduce the additional increase in the number of channels caused by the fitting and reduce the network parameters;
[0017] For the reconstruction module, we choose to build our network in the way of connecting dense blocks in series. We choose a total of six dense blocks for feature reconstruction, and each dense block has a three-layer structure. Similar to the previous spatial feature extraction, our channel growth number is set to 32, but we do not introduce residuals anymore. Instead, we choose to place DSTA at the end of each dense block to ensure that the network pays more attention to the texture-rich differences, such as the edges of objects, etc.;
[0018] In process (3), is the sample block of the original frame, is the sample block of the corresponding encoded frame. F(·) represents the post-processing network for compressed video, and θ1 represents the parameters of the post-processing network. Thus, the loss function of the post-processing network for compressed video is expressed as:
[0019]
[0020] The advantages and beneficial technical effects of the present invention compared with the prior art are as follows:
[0021] (1) The dense block concatenation method proposed in this invention strengthens the feature transfer process. The convolution layer can access all subsequent layers and pass the information that needs to be saved. Compared with ordinary convolution, it achieves lower parameter count and better effect.
[0022] (2) The present invention proposes to reuse the spatial information of the frame to be enhanced, thereby improving the problem of inaccurate spatiotemporal features obtained by deformable convolution during alignment, and at the same time compensating for the problem of the effectiveness of the multi-frame quality enhancement method in the case of intense inter-frame motion or rapid scene changes;
[0023] (3) The method for improving the quality of space-time compressed videos using a dense network proposed in this invention is an end-to-end network, that is, the compressed video can directly obtain a quality-improved video sequence through the network. Our method is easier to train and test, and solves the problem of unified modeling of PQF and non-PQF. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a framework diagram of the method for improving the quality of space-time compressed videos using dense networks.
[0025] Figure 2 This is a structural diagram of the DB block used to extract spatial information in the present invention.
[0026] Figure 3 FIG. 4 is a structural diagram of the DQE block used for quality enhancement in the present invention.
[0027] Figure 4 The rate-distortion curves of the sequence Basketball Drill 832×480 based on the method of the present invention, HEVC, and the comparison method [2] under the LDP configuration are shown.
[0028] Figure 5 The rate-distortion curves of the sequence PartyScene832×480 based on the method of the present invention, HEVC and the comparison method [2] under the LDP configuration.
[0029] Figure 6For the Kimono1920×1080 sequence at QP = 37, it is a comparison chart of the subjective visual quality of the HEVC standard, the method of the present invention, and two comparison methods. Figure (a) is a frame of the sequence after compression by the HEVC standard, with a PSNR of 34.41 dB. Figure (b) is the same frame of the sequence after compression by the HEVC standard and then processed by comparison method [1], with a PSNR of 34.96 dB. Figure (c) is the same frame of the sequence after compression by the HEVC standard and then processed by comparison method [2], with a PSNR of 35.43 dB. Figure (d) is the same frame of the sequence after compression by the HEVC standard and then processed by the present invention, with a PSNR of 35.50 dB.
[0030] Figure 7 For the RaceHorses832×480 sequence at QP = 37, it is a comparison chart of the subjective visual quality of the HEVC standard, the method of the present invention, and two comparison methods. Figure (a) is a frame of the sequence after compression by the HEVC standard, with a PSNR of 30.09 dB. Figure (b) is the same frame of the sequence after compression by the HEVC standard and then processed by comparison method [1], with a PSNR of 30.48 dB. Figure (c) is the same frame of the sequence after compression by the HEVC standard and then processed by comparison method [2], with a PSNR of 30.57 dB. Figure (d) is the same frame of the sequence after compression by the HEVC standard and then processed by the present invention, with a PSNR of 30.76 dB. Detailed implementation manners
[0031] The present invention will be further described in detail below in conjunction with embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and cannot be understood as limiting the protection scope of the present invention. Those skilled in the art can make some non-essential improvements and adjustments to the present invention according to the above-mentioned invention content and conduct specific implementations, which should still fall within the protection scope of the present invention.
[0032] (1) The method proposed in the present invention is carried out on the HM-16.5 platform of the HEVC standard test code. Among them, under the LDP configuration, the configuration file selects encoder_lowdelay_p_main.cfg. All test sequences have fully tested all frames. The standard video test sequences are encoded and decoded under the quantization parameters QP of 22, 27, 32, and 37, and the bit rate and peak signal-to-noise ratio PSNR during standard HEVC video encoding are recorded;
[0033] (2) The objects to be encoded are 18 standard HEVC test videos, and their names and resolutions are: Traffic (2560×1600), PeopleOnStreet (2560×1600), Kimono (1920×1080), ParkScene (1920×1080), BQTerrace (1920×1080), Cactus (1920×1080), BasketballDrive (1920×1080), FourPeople (1280×720), Johnny (1280×720), KristenAndSara (1280×720), BasketballDrill (832×480), PartyScene (832×480), RaceHorses (832×480), BQMall (832×480), BasketballPass (416×240), BQSquare (416×240), BlowingBubbles (416×240), RaceHorses (416×240);
[0034] (3) Use the method of the present invention to sequentially improve the quality of the video sequences encoded under the above conditions;
[0035] (4) We use the peak signal-to-noise ratio (PSNR) to measure the reconstruction quality of the compressed videos of the method of the present invention relative to HEVC and other methods. The higher this index, the better the quality of the reconstructed video. All video sequences are measured only in the luminance (Y) channel;
[0036] (5) Table 1 shows the comparison of the reconstruction quality of compressed videos between HEVC and other methods and the present invention under the LDP configuration. When the QP is 22, compared with the compressed video before improvement, the average PSNR of the method of the present invention on 18 video sequences is increased by 0.89 dB. Compared with other methods, the method of the present invention also has a certain quality improvement. At the same time, under other QPs, the method of the present invention also achieves a certain performance improvement compared with the relevant comparison methods, which indicates that our method has achieved good results;
[0037] (6) From Figure 4 and Figure 5 it can be seen that after the quality of video sequences with different resolutions is improved by our network, the obtained rate-distortion curves are above the standard and comparison methods, which indicates that under the condition of the same bit rate, the PSNR of the reconstructed videos of the method we proposed is higher. That is to say, the quality improvement network we proposed has good effects in improving the quality of compressed videos;
[0038] (7) FromFigure 6 and Figure 7 As can be seen from Figure 7 , compared with the HEVC compressed video, the subjective effect of the reconstructed video of the method we proposed shows that the reconstructed images are closer to the real images in terms of edges and textures. Compared with several comparison methods, it also has certain advantages, indicating that the method we proposed can reduce the compression effect and better restore the detailed information of the video.
[0039] The comparison methods are as follows:
[0040] Method 1: The method proposed by Zhenyu G, Qunliang X, Mai X, etc., reference "MFQE 2.0: A New Approach for Multi-frame Quality Enhancement on Compressed Video. [J] / / IEEE transactions on pattern analysis and machine intelligence, 2019: 949 - 963."
[0041] Method 2: The method proposed by Minyi Z; Yi X; Shuigeng Z, etc., reference "Recursive Fusion and Deformable Spatiotemporal Attention for Video Compression Artifact Reduction. [C] / / Proceedings of the 29th ACM International Conference on Multimedia. 2021: 5646 - 5654."
[0042] Table 1 PSNR comparison of the present invention with HEVC and other methods under LDP configuration
[0043]
[0044]
Claims
1. A method for improving the quality of space-time compressed video using dense networks, characterized by The following steps are involved: (1) Construct a training dataset. The dataset used in the multi-frame quality enhancement (MFQE) method is selected and compressed using HEVC with four QP values of 22, 27, 32, and 37. The original video and its corresponding compressed video sequences under the four QPs are obtained. The dataset is then processed by randomly cropping it to 128×128 blocks, flipping it, and rotating it. (2) First, the current compressed frame, the three forward frames, and the three backward frames are input into the deformable convolution module to obtain the initial spatiotemporal fusion features. Then, the spatial features are extracted from the current compressed frame. The features obtained from the two branches are passed through the fusion module in the network to obtain the optimized fusion features. Finally, the reconstruction module outputs the reconstructed frame with improved quality. When extracting the initial spatiotemporal fusion features by deformable convolution, we choose the currently widely used DCNV2. Its structure is a layer of convolution to transfer multiple original images to be aligned to the feature domain, and then use Unet to obtain the offset values of the convolution kernel in the horizontal and vertical directions. Finally, multiple original images and offset values are passed through a layer of deformable convolution to obtain the initial spatiotemporal fusion features. The structure of the spatial feature extraction module is as follows: first, the frame to be enhanced is transferred to the feature domain through a layer of convolution; then three consecutive feature propagation branches are connected, and the internal structure of each branch adopts a dense block. Specifically, the dense block has five convolutional layers with an output channel set to 32. The number of channels of the initial convolutional layer is set to 32. Each subsequent convolutional layer will add the output of each previous convolutional layer, that is, the channel growth rate is set to 32. Each convolution layer except the last one is followed by a linear pooling layer, so the input channel number of the last convolutional layer is 160 and the output channel number is 32. In order to ensure the stability of training, a residual structure is introduced, and the input of the dense block and the output of the last convolutional layer are added together. In addition to the residual structure within each dense block, a residual is also performed on the original input and final output of the three dense blocks. To obtain optimized spatiotemporal fusion features, we first need to process the spatial features. We selected Deformable Spatiotemporal Attention (DSTA), a spatial attention module, to determine the useful information in the extracted spatial information. We then performed channel-wise alignment of the spatial features with the initial fusion features, and introduced a channel-wise attention mechanism to distinguish useful and useless channel information. Finally, we used a convolutional layer with a 1×1 kernel to reduce the additional number of channels caused by the alignment, thereby reducing the number of network parameters. In the reconstruction module, we constructed our network by connecting dense blocks in series. We used a total of six dense blocks for feature reconstruction, and each dense block contained three convolutional layers. Similar to the previous spatial feature extraction, our channel growth was set to 32. However, we did not introduce residuals. Instead, we chose to place DSTA at the end of each dense block to ensure that the network focused on texture-rich distinctions, such as object edges. (3) In the training phase, the present invention selects the original video and the compressed video of a selected QP for training. The compressed videos of other QPs are also trained separately with the original video, and a total of four quality improvement models under four QPs are obtained. The test uses the HEVC standard test sequence, obtains the corresponding enhanced video sequence after passing the quality improvement network, and uses the peak signal-to-noise ratio (PSNR) indicator to verify the effectiveness of the method.
Citation Information
Cited By
Rice multi-view image three-dimensional imaging method considering processing time and reconstruction precision
CN121392103A
A method for three-dimensional imaging of rice from multiple perspectives, balancing processing time and reconstruction accuracy.
CN121392103B