A method for enhancing compressed video quality based on omniscient network

Through the omniscient network learning of video space-time information and full-frequency domain information, the problem of degradation of compressed video quality is solved, and the video quality is significantly improved and the bit rate is reduced.

CN115496683BActive Publication Date: 2025-05-13GUANGZHOU WATER GUARD INFORMATION TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211132048.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-16
Publication Date
2025-05-13
Estimated Expiration
2042-09-16

AI Technical Summary

Technical Problem

The existing compressed video technology will lead to a decline in the video quality during the compression process, resulting in block effects and artifacts, affecting the subjective and objective quality of the video.

Method used

The omniscient network is adopted to learn the space-time information and full-frequency domain information of video through the space-time feature fusion module and the full-frequency adaptive enhancement block, and quality enhancement is carried out to reduce the quality fluctuations of compressed video.

Benefits of technology

It significantly improves the subjective and objective quality of compressed video, reduces the bit rate, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496683B_ABST
    Figure CN115496683B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for enhancing the quality of compressed videos based on an omniscient network, and belongs to the field of image processing technology. The present application includes: first, the features of the current frame are initialized by aggregating the spatiotemporal information of 2R+1 video frames using a spatiotemporal feature fusion module. Then, a network similar to a grid structure is used to realize the bidirectional propagation of features, and the spatiotemporal information of the entire video is utilized to the maximum extent. The features are repeatedly iterated and refined in the network similar to a grid structure, and a plurality of different hidden features are generated. Finally, all hidden features are fused, and put into a quality enhancement network to further learn the full frequency domain information within the frame, and generate an enhanced residual. The final enhanced frame consists of a compressed frame plus an enhanced residual. The present invention proposes a method for multi-frame quality enhancement of compressed videos by fully utilizing the spatiotemporal information and full frequency domain information of the entire video. The present invention significantly improves the subjective and objective quality of the video, while reducing the bit rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing, and in particular relates to a compressed video quality enhancement method based on an omniscient network. Background Art

[0002] Since the Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T and ISO / IEC officially promulgated the High Efficiency Video Coding (HEVC) standard in April 2013, it has received widespread attention. HEVC video compression usually consists of five steps: a) dividing the input frame into small blocks; b) performing intra-frame and inter-frame prediction; c) applying discrete cosine transform (DCT) to the prediction block; d) using quantization parameters to remove the DCT coefficients of each transform block and rounding the quantized value; e) using entropy coding to generate the bitstream of the compressed video. Since the transformation and quantization process ignores the correlation between blocks, the compressed video will show block effects. At the same time, the principle of quantization is to map a large data set to a small data set. Simply put, it is to round out and select parameters with large data volumes so that the information can be stored in a smaller memory. Since this process is irreversible, quantization will lead to information loss. All compressed videos will inevitably lose their original details and produce a large number of compression artifacts, resulting in a serious decline in the subjective and objective quality of the video.

[0003] With the development of deep learning, many methods based on deep learning to improve video quality have emerged, and they have achieved very good results. For example, in the Chinese patent application with publication number CN107481209A and application name "A method for enhancing image or video quality based on convolutional neural network", two convolutional neural networks for video (or image) quality enhancement are first designed, and the two networks have different computational complexities; then several training images or videos are selected to train the parameters in the two convolutional neural networks; according to actual needs, a convolutional neural network with a more appropriate computational complexity is selected, and the image or video to be enhanced is input into the selected network; finally, the network outputs the image or video with enhanced quality. This processing scheme can effectively enhance the video quality; users can specify the use of a convolutional neural network with a more appropriate computational complexity to enhance the quality of images or videos according to the computing power or remaining power of the device. However, in this processing scheme, since two convolutional neural networks of different complexity are designed and the user selects the network according to the situation of the device, the difference between the two networks is only the depth of the convolutional neural network. It is not feasible to improve the quality enhancement effect simply by deepening the network depth. Moreover, the network is not designed according to the characteristics of image video, that is, the network fails to utilize the temporal correlation between video frames. Therefore, the quality enhancement effect of this method is limited.

[0004] In the Chinese patent application with publication number CN108900848A and application name “A video quality enhancement method based on adaptive separable convolution”, adaptive separable convolution is applied as the first module in the network model, and each two-dimensional convolution is converted into a pair of one-dimensional convolution kernels in the horizontal and vertical directions. The parameter quantity is changed from n 2 Becomes n+n. And the adaptively changing convolution kernel learned by the network for different inputs is used to realize the estimation of motion vectors. By selecting two consecutive frames as network input, a pair of separable two-dimensional convolution kernels can be obtained for every two consecutive inputs, and then the two-dimensional convolution kernel is expanded into four one-dimensional convolution kernels. The obtained one-dimensional convolution kernel changes with the change of input, thereby improving the network adaptability. The invention replaces the two-dimensional convolution kernel with a one-dimensional convolution kernel, so that the parameters of the network training model are reduced and the execution efficiency is high. The processing scheme uses five encoding modules and four decoding modules, a separation convolution module and an image prediction module. Its structure is based on the traditional symmetrical encoding and decoding module network, and the last decoding module is replaced by a separation convolution module. Although it effectively reduces the parameters of the model, the quality enhancement effect needs to be further improved.

[0005] In a Chinese patent with publication number CN108307193A and application name “A method and device for multi-frame quality enhancement of lossy compressed video”, the multi-frame quality enhancement method disclosed therein includes: for the i-th frame of the decompressed video stream, the i-th frame is quality enhanced using the m frames associated with the i-th frame to play the i-th frame with enhanced quality; the m frames belong to the frames in the video stream, and each of the m frames has the same or corresponding number of pixels as the i-th frame, respectively, greater than a preset threshold; m is a natural number greater than 1. In specific applications, peak quality frames can be used to enhance non-peak quality frames between two peak quality frames. Method 3 reduces the quality fluctuations between multiple frames during video stream playback, and at the same time enhances the quality of each frame in the video after lossy compression. Although the multi-frame quality enhancement method takes into account the temporal information between adjacent frames, the designed multi-frame convolutional neural network (MF-CNN) is divided into a motion compensation subnet (MC-subnet) and a quality enhancement subnet (QE-subnet), where the motion compensation subnet relies heavily on optical flow estimation to compensate for the motion between non-peak quality frames and peak quality frames to achieve alignment. Any error in the optical flow calculation will introduce artifacts around the image structure in the aligned adjacent frames. However, accurate optical flow estimation itself is challenging and time-consuming, so the quality enhancement effect of the multi-frame quality enhancement method is still limited.

[0006] And in the U.S. patent application with publication number US20200404340A1 and titled “Multi-frame quality enhancement method and device for lossy compressed video”, a multi-frame quality enhancement method for lossy compressed video is disclosed, which is: identifying peak quality frames and non-peak quality frames in compressed video frames; using peak quality frames to enhance non-peak quality frames located between two peak quality frames. Although the scheme takes into account the temporal information between adjacent frames, the designed multi-frame convolutional neural network (MF-CNN) is divided into a motion compensation subnet (MC-subnet) and a quality enhancement subnet (QE-subenet), in which the motion compensation subnet relies heavily on optical flow estimation to compensate for the motion between non-peak quality frames and peak quality frames to achieve alignment. Any error in the optical flow calculation will introduce artifacts around the image structure in the aligned adjacent frames. However, accurate optical flow estimation itself is challenging and time-consuming, so the quality enhancement effect of this scheme is still limited.

[0007] As the number of videos increases exponentially and videos are becoming more high-definition, a series of problems will inevitably arise in the storage and transmission of videos. Compressing videos is one of the essential means in the future, but compressing videos will inevitably cause the compressed videos to lose a lot of details and cause serious artifacts, which seriously affects the video quality. Therefore, it is necessary to propose a compressed video quality enhancement scheme to further reduce the artifacts caused by compressed videos and restore the original structure of the video. Summary of the invention

[0008] The purpose of the present invention is to use an omniscient network to learn the spatiotemporal information and full-frequency domain information of the entire video, and effectively integrate the learned information to improve the quality of compressed frames, reduce the quality fluctuations caused by compressed video, and improve the user experience quality.

[0009] The present invention provides a method for enhancing the quality of compressed video based on an omniscient network, the method comprising the following steps:

[0010] Step 1, set up and train the omniscient network;

[0011] The omniscient network includes a spatiotemporal feature fusion (STFF) module and an overall frequency adaptive enhancement (OFAE) block, wherein the spatiotemporal feature fusion module is used to aggregate the spatiotemporal information of multiple frames, and the overall frequency adaptive enhancement block is used to learn the overall frequency information within the frame.

[0012] The omniscient network includes multiple branches, which are used to perform quality enhancement processing on a compressed video sequence of a specified length. Each branch sequentially inputs a video frame, and each branch sequentially includes: a spatiotemporal feature fusion module, a multi-layer feature extraction network and a layer of quality enhancement network. The input video frame of each branch is added to the enhanced video frame output by the quality enhancement network to obtain a final enhanced video frame;

[0013] The first empty feature fusion module of each branch is used to learn the temporal and spatial information of the R adjacent frames before and after the input video frame of the current branch and its own frame and generate the initial features of the current branch, where R is a positive integer;

[0014] The feature extraction network includes a spatiotemporal feature fusion module and a full-frequency adaptive enhancement block. The output of the spatiotemporal feature fusion module is spliced ​​with the input of the feature extraction network along the channel dimension and then input into the full-frequency adaptive enhancement block. The output of the full-frequency adaptive enhancement block is spliced ​​with its input along the channel dimension to obtain the output of the feature extraction network.

[0015] All feature extraction networks of all branches form a full feature enhancement network, and the initial features of each branch are propagated in the full feature enhancement network, first back-propagated and then forward-propagated. After multiple rounds of iterations, the feature map of each branch is obtained and input into the quality enhancement network;

[0016] The quality enhancement network consists of two convolutional layers and multiple full-frequency adaptive enhancement blocks located between the two convolutional layers;

[0017] The omniscient network is trained end-to-end using the collected training data to obtain a trained omniscient network;

[0018] Step 2: Input the compressed video sequence to be enhanced into the trained omniscient network, and obtain the corresponding enhanced video frame sequence based on its output.

[0019] In one possible implementation, each branch includes a 4-layer feature extraction network.

[0020] In one possible implementation, the quality enhancement network of each branch includes five full-band adaptive enhancement blocks.

[0021] In one possible implementation, the spatiotemporal feature fusion module includes: convolution layer 1, U-type network, convolution layer 2 and deformable convolution layer in sequence, and the input of convolution layer 1 is also input to the deformable convolution layer.

[0022] In one possible implementation, the U-shaped network of the spatiotemporal feature fusion module includes: three downsampling layers and two SKFF modules (selective kernel feature fusion modules), and the output of the first downsampling layer is also input to the last SKFF module, and the output of the second downsampling layer is also input to the first SKFF module. The SKFF module is used to adaptively model features of different scales, that is, to learn the weights of features of each scale based on the attention mechanism, and to fuse features of different scales in a weighted sum manner.

[0023] In one possible implementation, the full-frequency adaptive enhancement block includes three branches:

[0024] The first branch includes the downsampling layer and the convolution layer in sequence;

[0025] The second branch includes a downsampling layer, an SKFF module and a convolution layer in sequence, and the output of the downsampling layer of the second branch is subtracted from the upsampling result of the output of the downsampling layer of the first branch and then input into the SKFF module of the second branch, and the output of the convolution layer of the first branch is also input into the SKFF module of the second branch;

[0026] The third branch includes convolution layer 1, SKFF module, convolution layer 2 and channel attention network in sequence, and the output of convolution layer 1 of the third branch is subtracted from the upsampling result of the output of the downsampling layer of the second branch and then input into the SKFF module of the first branch, and the output of the convolution layer of the second branch is also input into the SKFF module of the third branch;

[0027] The input of the full-frequency adaptive enhancement block is added to the output of the third branch through a convolutional layer as the output of the full-frequency adaptive enhancement block.

[0028] In one possible implementation, the loss function used for end-to-end training of the omniscient network is:

[0029]

[0030] in, represents the enhanced video frame output by the omniscient network, represents the original video frame sample, ∈ represents a preset constant whose value is (0, 1), i represents the video frame number of the video sequence sample, and n represents the number of frames of the video sequence sample.

[0031] The technical solution provided by the present invention brings at least the following beneficial effects:

[0032] The present invention proposes a method for compressing multi-frame quality enhancement of video by fully utilizing the spatiotemporal information and full-frequency domain information of the entire video. First, the spatiotemporal information of adjacent frames is initially extracted through the spatiotemporal feature fusion module, and then a feature extraction network with a grid-like structure is used to realize bidirectional propagation of features to maximize the use of the spatiotemporal information and full-frequency information of the entire video. Finally, the full-frequency domain information in the frame is fully learned through the quality enhancement network, so that the subjective and objective quality of the video is significantly improved, while the bit rate is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0034] Figure 1 It is a visualized schematic diagram of the loss of frequency domain information caused by video compression in an embodiment of the present invention, wherein HEVC (High Efficiency Video Coding) refers to a video compressed by HEVC software (HM16.5).

[0035] Figure 2 It is a schematic diagram of a processing framework of a compressed video quality enhancement method based on an omniscient network provided by an embodiment of the present invention;

[0036] Figure 3 is a schematic diagram of the network structure of a spatiotemporal feature fusion (STFF) module in an embodiment of the present invention;

[0037] Figure 4 is a schematic diagram of a network structure of an OFAE module in an embodiment of the present invention;

[0038] Figure 5 It is a schematic diagram of the enhancement effect on the video sequences BQSquare and PartyScene when QP=37 in an embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0040] Today's network, image processing, and digital video technologies have developed rapidly. With the continuous advancement of technology, people have higher requirements for video quality, and the demand for video is also increasing exponentially. However, storing original videos on devices with limited storage space will inevitably take up huge amounts of space, and transmitting these videos will inevitably bring a series of problems. Therefore, it has become an inevitable trend to compress the original video through video encoding technology to save encoding bit rate. However, the most common video encoding technologies currently are often lossy, such as H.264 / AVC and H.265 / HEVC, which means that compressed video will inevitably cause the loss of original video details, thus affecting the quality of user experience. That is, compressed video will inevitably lead to the loss of low-frequency and high-frequency information in the video. Such as Figure 1 As shown in the figure, there is obvious frequency domain information loss in both videos. In view of this, the embodiment of the present invention proposes a compressed video quality enhancement method based on an omniscient network to improve the quality of compressed video. The OFAE module proposed in the present invention can effectively improve the different frequency domain information of compressed video, such as Figure 2 As shown in Figure 2, the video enhanced by the OFAE module restores more accurate frequency domain components.

[0041] As a possible implementation, the network structure of the omniscient network (OVQE) used in the compressed video quality enhancement method based on the omniscient network in the embodiment of the present invention is as follows: Figure 2 As shown, it includes two modules: a spatiotemporal feature fusion (STFF) module and an all-frequency adaptive enhancement block (OFAE) block. The former is intended to capture the spatiotemporal information in adjacent frames, while the latter is intended to adaptively restore different frequency domains of the compressed video. Information is designed to be propagated bidirectionally in a grid manner so that the results of all-round enhancement can be applied. Based on these two modules, the present invention can reduce the artifacts caused by compressed video and improve the quality of compressed video. Given a compressed video consisting of T frames in Represents the compressed frame at time t. The goal of this application is to better reduce compression artifacts and restore lost details of the video. In order to make full use of the spatiotemporal information of the entire video, these T frames are used as input to enhance T frames. This process can be expressed as:

[0042]

[0043] Among them, OVQE represents the omniscient network proposed in this application, that is, OVQE{} represents the output of the omniscient network, and θ is a learnable parameter, that is, the network parameter of the omniscient network.

[0044] The task of compressed video quality enhancement involves a complex process because it requires aggregating information not only from the spatiotemporal dimensions, but also from the frequency domain dimensions. The processing flow of the omniscient network of the present application is as follows. First, the features of the current frame are initialized by aggregating the spatiotemporal information of 2R+1 video frames using the spatiotemporal feature fusion module. Then, a network with a grid-like structure is used to realize the bidirectional propagation of features and maximize the use of the spatiotemporal information of the entire video. The features are iteratively refined in the grid-like structure network, and four different hidden features are generated. Finally, all hidden features are fused and placed in the quality enhancement network to further learn the full frequency domain information in the frame to generate an enhanced residual. The final enhanced frame is composed of the compressed frame. Plus enhanced residual composition.

[0045] like Figure 2 As shown, the omniscient network of the present invention includes two novel network modules: the STFF module, which is used to aggregate the spatiotemporal information of multiple frames, and the OFAE block (grid line square), which is used to learn the full-frequency information within the frame. Figure 2 The diagonal square in the figure consists of a STFF module and an OFAE block. The framework of the omniscient network can more effectively utilize the full-time, spatial and frequency information in the video. Initially, the STFF module is used to learn the spatiotemporal information of 2R+1 adjacent frames and generate the fused features of the current frame, which can be expressed as:

[0046]

[0047] in, represents continuous input frames, in this embodiment, R=3, F represents STFF module, θ 1 is a learnable parameter of the STFF module. Through the initialization of the STFF module, the initial features can utilize the spatiotemporal information of adjacent frames. It should be noted that for the frames at both ends of the video clip, 2R+1 frames (the R frames before and after the current frame and the current frame) are collected by directly copying the end frame. Then, the initial features will be propagated in the full feature enhancement network.

[0048] This is followed by a bidirectional propagation of features, where the features are first back-propagated and then forward-propagated. This iterates two rounds in total and generates four hidden states for each frame. After four propagations, the final connected hidden state is used to reconstruct the frame. Figure 2 Medium diagonal square.

[0049] First, for the frame In the first back propagation, another STFF module uses x t and the past estimated hidden states As input, it generates aligned features Mathematically, it can be expressed as follows:

[0050]

[0051] Among them, C represents the concatenation of features along the channel dimension, θ 2 is the learnable parameter of STFF during the first back propagation. In this embodiment, the number of input channels of F is 128 and the number of output channels is 64. Then a full frequency adaptive enhancement (OFAE) block is used to fuse and x t To perform full frequency domain enhancement and finally generate a new hidden state

[0052]

[0053] Where O represents the full-frequency adaptive enhancement block, θ 3 is the corresponding learnable parameter. In this embodiment, the number of input channels of O is 128 and the number of output channels is 64.

[0054] After that, the feature is forward propagated. Unlike the first backpropagation, in the forward propagation, the full spatiotemporal information (past, present and future) is used to further enhance the current frame. Mathematically, it can be expressed as follows:

[0055]

[0056] Among them, θ 4 are the corresponding learnable parameters. is the past hidden state generated by the first forward propagation. Represents the current and future hidden states generated by the first back propagation. In this embodiment, the number of input channels of F is 192 and the number of output channels is 64. Similarly, the full-frequency adaptive enhancement block is used to enhance the feature Perform full frequency domain enhancement:

[0057]

[0058] Similarly, for the second round of propagation, we can get and Hidden state, the final fused feature map is generated as:

[0059]

[0060] Finally, the quality enhancement network is crucial to restore video quality. It needs to decode the fused features into enhanced residuals and use the residuals to reconstruct high-quality frames. Considering that frequency domain information is an important part of the frame, this paper designs a lightweight quality enhancement network to improve video quality. Figure 2 As shown in the straight square, it consists of only 5 OFAE blocks and 2 convolutional layers to maintain a similar computational complexity to the quality enhancement module in STDF (for details, please refer to the document "Deng J, Wang L, Pu S, et al. Spatio-temporal deformable convolution for compressed video quality enhancement [C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2020, 34 (07): 10696-10703."). The first convolutional layer reduces the number of channels of the fused feature ht to 64; for the convenience of cascading, the number of input channels of the OFAE module of the quality enhancement network is modified to 64; the last convolutional layer outputs the enhanced residual. The final enhancement result It can be obtained by adding the residual map back to the compressed target frame:

[0061]

[0062] Among them, F QE represents the quality enhancement network, θ QE is the corresponding parameter.

[0063] As a possible implementation, the spatiotemporal feature fusion (STFF) module of the present invention is set as follows:

[0064] The STFF module is mainly composed of SKU-Net and deformable convolution. Considering the existence of violent object motion and the great success of U-Net in image restoration, this paper proposes an improved U-Net (SKU-Net) to implicitly align multiple adjacent frames to reduce the impact of object motion and predict additional learnable deformation offsets; secondly, the learned offsets are used to guide the deformable convolution to align multiple compressed frames or features to further collect information in adjacent frames.

[0065] like Figure 3As shown, SKU-Net performs multiple downsampling by using convolutions with a stride of 2 to extract information at different scales. When fusing multi-scale information, the SKFF module (for details, please refer to the literature "SW Zamir, A. Arora, SHKhan, H. Munawar, FSKhan, M.-H. Yang, and L. Shao, "Learning enriched features for fast image restoration and enhancement," IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–1, 2022") is used to adaptively model features of different scales and learn two attention activations S1 and S2. Use them to adaptively recalibrate the multi-scale feature maps L1 and L2. This process can be expressed by the following formula:

[0066] U=S1×L1+S2×L2

[0067] Where U represents the output of the SKFF module.

[0068] Compared with U-Net, SKU-Net can adaptively fuse features of different resolutions while retaining their unique characteristics.

[0069] It should be noted that Figure 3 The structure of the spatiotemporal feature fusion (STFF) module is shown. Unless otherwise specified, the stride of the convolution kernel is 1.

[0070] As a possible implementation, the OFAE block of the present invention is set as follows:

[0071] Since the frequency domain is also an important part of the video, it is difficult for existing methods to adaptively learn the full-frequency information of the compressed video, resulting in poor enhancement effect. Traditional methods use Fourier transform to decompose the high-frequency and low-frequency components of the image. In the embodiment of the present invention, in order to achieve better results, the feature is decomposed into three frequency domains. The present invention adaptively enhances the frequency domain of the frame by combining the SKFF module and channel attention. The cascade structure ensures information exchange between multiple frequency domains. The architecture of the OFAE module is as follows: Figure 4 It consists of two key lightweight components: full-frequency feature extraction and full-frequency feature enhancement. Figure 4 In the structure of the OFAE module shown, if not otherwise specified, the stride of the convolution kernel is 1.

[0072] For full-frequency feature extraction, first, a convolution with a step size of 4 is used to downsample the feature u (the input of the full-frequency adaptive enhancement block) to obtain the corresponding low-frequency component f l Then, the feature u is downsampled using a convolution with a stride of 2 to obtain the mid-frequency and low-frequency components f ml , remove the low frequency component f l Get the intermediate frequency component f m Similarly, the high-frequency component f is obtained by direct convolution hml , remove the low and medium frequencies to obtain the high frequency component f h The whole process can be expressed as follows:

[0073]

[0074]

[0075]

[0076] f m =f ml -Up(f l )

[0077] f h =f hml -Up(f ml )

[0078] Among them, Up represents the bilinear upsampling operation, that is, Up() represents the output of the bilinear upsampling operation, Conv↓ 2 and Conv↓ 4 They represent convolutions with stride 2 and stride 4 respectively. Conv represents ordinary convolution, and the symbol Indicates nested functions.

[0079] For full-frequency feature enhancement, the main task of full-frequency feature enhancement is to fuse and enhance different frequency features. A cascade enhancement procedure is adopted. First, the low-frequency component is enhanced, then the mid-frequency, and finally the high frequency. Mathematically, it can be expressed as:

[0080]

[0081]

[0082]

[0083]

[0084]

[0085]

[0086]

[0087] Among them, f l enc Indicates the enhanced low-frequency characteristics, Represents the fused low-frequency and mid-frequency features. Represents enhanced low- and mid-range capabilities, represents the full-frequency characteristics of the fusion, Indicates enhanced full-frequency characteristics, represents the full-frequency feature learned through channel attention, F skff represents the SKFF module, CA stands for channel attention, and when the input feature channel dimension of the OFAE block is different from that of the output, a 1×1 convolution is used to change the input channel dimension.

[0088] Since the omniscient network is fully convolutional, in the embodiment of the present invention, an end-to-end joint training is adopted, a video clip is used as input for joint training, and the Charbonnier Loss is used to optimize the network model:

[0089]

[0090] in, represents the enhanced video frame, Represents the original video frame, ∈ is a preset constant, which can be set according to specific needs, usually set at 10 -6 It is more appropriate to set it to 10 in this embodiment. -6 , n is the length of the video clip.

[0091] In order to demonstrate the effectiveness and superiority of the present invention, in this embodiment, qualitative and quantitative evaluations are performed on the HEVC standard test sequence.

[0092] Quantitative evaluation: The improved PSNR and SSIM (ΔPSNR and ΔSSIM) of the present invention are compared with the most advanced methods in recent years, DCAD, DS-CNN, MFQE, MFQE 2.0, STDF, RFDA, BasicVSR++ and Xu et al. to evaluate the performance of the present invention. PSNR stands for Peak Signal-to-Noise Ratio, which is an objective criterion for evaluating image quality, and SSIM stands for Structural Similarity, which is a full-reference image quality evaluation criterion. Table 1 shows the average results of ΔPSNR and ΔSSIM of the present invention on all frames of each test video. It can be seen from the data in the table that the present invention is consistently superior to the most advanced methods. Specifically, when QP=37, the present invention method outperforms the iterative method STDF by 0.34dB in terms of the average ΔPSNR of 18 standard test videos. Compared with the recursive method RFDA, OVQE achieves an improvement of 0.26dB. Compared with the multi-recursive method BasicVSR++, OVQE achieves a PSNR improvement of 0.18dB. Compared with the state-of-the-art method Xu et al., our method achieves an improvement of 0.14 dB.

[0093] Table 1 ΔPSNR (dB) and ΔSSIM ((×10 -2 )

[0094]

[0095] *Video resolution: Class A(2560×1600), Class B(1920×1080), Class C(832×480), Class D(480×240), Class E(1280×720).

[0096] In addition, in this embodiment, the reduction of BD-rate (an important indicator for evaluating the performance of video coding technology) is also used to evaluate the performance of the method of the present invention. By fine-tuning the number of channels of the STDF and RFDA models, the STDF-128 and RFDA-128 models are retrained. As shown in Table 2, the BD-rate of the method of the present invention is reduced by an average of 30.56%, which is better than the current advanced methods STDF (21.61%), RFDA-128 (24.37), STDF-128 (24.65) and BasicVSR++ (26.31).

[0097] Table 2 compares the BD-rate (%) reduction results on the HEVC standard test dataset.

[0098]

[0099] Among them, the method DCAD can refer to the literature "Wang T, Chen M, Chao HA Novel Deep Learning-Based Method of Improving Coding Efficiency from the Decoder-end for HEVC[C] / / Data Compression Conference(DCC)2017.IEEE,2007.";

[0100] Method DS-CNN can refer to the literature "Ren Y, Mai X, Tie L, et al. Enhancing Quality for HEVC Compressed Videos[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2017, PP: 1-1.";

[0101] Method MFQE can be found in the literature "Yang R, Xu M, Wang Z, et al. Multi-frame Quality Enhancement for Compressed Video[C] / / 2018IEEE / CVF Conference on Computer Vision and Pattern Recognition(CVPR).IEEE, 2018.";

[0102] For method MFQE 2.0, please refer to the literature "Guan Z, Xing Q, Xu M, et al. MFQE 2.0: A new approach for multi-frame quality enhancement on compressed video[J]. IEEE transactions on pattern analysis and machine intelligence, 2019.";

[0103] Method STDF possible references《Deng J,Wang L,Pu S,et al.Spatio-temporaldeformable convolution for compressed video quality enhancement[C] / / Proceedings of the AAAI Conference on Artificial Intelligence.2020,34(07):10696-10703.》;

[0104] Method RFDA can be referenced as follows: M. Zhao, Y. Xu, and S. Zhou, “Recursive Fusion and Deformable Spatiotemporal Attention for Video Compression Artifact Reduction,” in Proceedings of the 29th ACM International Conference on Multimedia, 2021.

[0105] Method BasicVSR++ Available references《KCChan, S. Zhou, X.

[0106] For the method, Xu et al can refer to the literature "Y.Xu, M.Zhao, J.Liu,

[0107] Qualitative assessment: Figure 5 The subjective quality of the video sequences BQSquare and PartyScene at QP = 37 is shown. It can be seen from the figure that compared with the current state-of-the-art methods, our method can better recover the structural details of the video.

[0108] The embodiment of the present invention provides a compressed video quality enhancement method based on an omniscient network, which utilizes an omniscient network to learn the spatiotemporal information and full-frequency domain information of the entire video, and can significantly improve the subjective and objective quality of the compressed video.

[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

[0110] The above are only some embodiments of the present invention. For those skilled in the art, several modifications and improvements can be made without departing from the creative concept of the present invention, which all belong to the protection scope of the present invention.

Claims

1. A method for enhancing compressed video quality based on an omniscient network, characterized in that: The following steps are involved: Step 1, set up and train the omniscient network; The omniscient network includes a spatiotemporal feature fusion module and a full-frequency adaptive enhancement block, wherein the spatiotemporal feature fusion module is used to aggregate the spatiotemporal information of multiple frames, and the full-frequency adaptive enhancement block is used to learn the full-frequency information in the frame. The omniscient network includes multiple branches, which are used to perform quality enhancement processing on a compressed video sequence of a specified length. Each branch sequentially inputs a video frame, and each branch sequentially includes: a spatiotemporal feature fusion module, a multi-layer feature extraction network and a layer of quality enhancement network. The input video frame of each branch is added to the enhanced video frame output by the quality enhancement network to obtain a final enhanced video frame; The first spatiotemporal feature fusion module of each branch is used to learn the spatiotemporal information of the R adjacent frames before and after the input video frame of the current branch and its own frame and generate the initial features of the current branch, where R is a positive integer; The feature extraction network includes a spatiotemporal feature fusion module and a full-frequency adaptive enhancement block. The output of the spatiotemporal feature fusion module is spliced ​​with the input of the feature extraction network along the channel dimension and then input into the full-frequency adaptive enhancement block. The output of the full-frequency adaptive enhancement block is spliced ​​with its input along the channel dimension to obtain the output of the feature extraction network. All feature extraction networks of all branches form a full feature enhancement network, and the initial features of each branch are propagated in the full feature enhancement network, first back-propagated and then forward-propagated. After multiple rounds of iterations, the feature map of each branch is obtained and input into the quality enhancement network; The quality enhancement network consists of two convolutional layers and multiple full-frequency adaptive enhancement blocks located between the two convolutional layers; The omniscient network is trained end-to-end using the collected training data to obtain a trained omniscient network; Step 2: Input the compressed video sequence to be enhanced into the trained omniscient network, and obtain the corresponding enhanced video frame sequence based on its output.

2. The method according to claim 1, characterized in that Each branch consists of a 4-layer feature extraction network.

3. The method according to claim 1, characterized in that The quality enhancement network of each branch includes 5 full-band adaptive enhancement blocks.

4. The method according to claim 1, characterized in that The spatiotemporal feature fusion module includes: convolution layer 1, U-shaped network, convolution layer 2 and deformable convolution layer in sequence, and the input of convolution layer 1 is also input to the deformable convolution layer.

5. The method according to claim 1, characterized in that The U-shaped network of the spatiotemporal feature fusion module includes three downsampling layers and two selective kernel feature fusion SKFF modules in sequence. The output of the first downsampling layer is also input to the last SKFF module, and the output of the second downsampling layer is also input to the first SKFF module. The SKFF module learns the weights of features at each scale based on the attention mechanism, and fuses features of different scales in a weighted sum manner.

6. The method according to claim 1, characterized in that The full-frequency adaptive enhancement block includes three branches: The first branch includes the downsampling layer and the convolution layer in sequence; The second branch includes a downsampling layer, a selective kernel feature fusion SKFF module and a convolution layer in sequence, and the output of the downsampling layer of the second branch is subtracted from the upsampling result of the output of the downsampling layer of the first branch and then input into the SKFF module of the second branch, and the output of the convolution layer of the first branch is also input into the SKFF module of the second branch; the SKFF module learns the weights of features of each scale based on the attention mechanism, and fuses features of different scales in a weighted sum manner; The third branch includes convolution layer 1, SKFF module, convolution layer 2 and channel attention network in sequence, and the output of convolution layer 1 of the third branch is subtracted from the upsampling result of the output of the downsampling layer of the second branch and then input into the SKFF module of the first branch, and the output of the convolution layer of the second branch is also input into the SKFF module of the third branch; The input of the full-frequency adaptive enhancement block is added to the output of the third branch through a convolutional layer as the output of the full-frequency adaptive enhancement block.

7. The method according to claim 1, characterized in that The loss function used for end-to-end training of the omniscient network is: in, represents the enhanced video frame output by the omniscient network, represents the original video frame sample, ∈ represents a preset constant, The value is (0, 1), i represents the video frame number of the video sequence sample, and n represents the number of frames of the video sequence sample.

Citation Information

Patent Citations

  • Image or video quality enhancement method based on convolution neural networks

    CN107481209A

  • Multi-frame quality enhancement method and device for lossy compressed video

    CN108307193A

  • Video quality enhancement method based on adaptive separable convolution

    CN108900848A

  • Multi-frame quality enhancement method and device for lossy compressed video

    US20200404340A1