A semi-blind method for removing VVC compression artifacts based on deep learning
By introducing quantitative parameter prediction and compression artifact removal models in the VVC compression artifact removal semi-blind method, the problem that the existing technology cannot effectively remove VVC compression artifacts in blind scenarios is solved, and efficient video quality improvement in blind scenarios is achieved.
Patent Information
- Application Number
- CN202111257660.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-10-27
AI Technical Summary
The existing VVC compression artifact removal methods based on deep learning are mostly non-blind methods and cannot work effectively in blind scenarios where quantization parameters cannot be obtained.
A semi-blind method for VVC compression artifact removal based on deep learning is proposed, including a quantitative parameter prediction method and a compression artifact removal method. By predicting the quantization parameters of VVC decoded video, the blind scene is converted into a non-blind scene, and the corresponding model is selected to remove compression artifacts.
Effectively remove the compression artifacts of VVC decoded videos, improve video quality, and perform well in blind scenarios where quantization parameters cannot be obtained.
Smart Images

Figure CN116074540B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to deep learning and VVC decoded video decompression technology, and specifically designs a semi-blind method for removing VVC compression artifacts based on deep learning, belonging to the post-processing technology of video coding and decoding. Background Art
[0002] With the rapid development of communication technology, people have more and more communication methods. In addition to voice communication methods such as walkie-talkies and mobile phones, image communication methods such as video phones and video conferences are more and more widely used. Compared with voice signals, image signals containing more information are more favored by people for their vividness and vividness. And video, which can represent continuous images of moving scenes, is also applied in various scenarios, such as surveillance videos for protecting cities, medical videos for assisting medical diagnoses, and live videos for leisure and entertainment. Thus, video has occupied an indispensable position in life.
[0003] Therefore, video technology has developed rapidly, and videos have gradually had a larger resolution and contained more information. This has also caused an undeniable characteristic of videos - large data volume and large required bandwidth, making the transmission and storage of videos a major difficulty. Versatile Video Coding (VVC) is a new generation of video coding and decoding standard. Compared with the High Efficiency Video Coding (HEVC) standard, VVC has more powerful compression performance and can more effectively remove a large amount of spatial redundancy and temporal redundancy information existing in video signals, thereby encoding videos into bitstream files with a large compression ratio. At the decoding end, the bitstream files can be decoded to reconstruct the video. However, due to operations such as quantization and transformation during the encoding process, some information of the reconstructed decoded video is lost, resulting in compression artifacts, reducing the video quality and the user's viewing experience. Therefore, it is very meaningful to remove the compression artifacts of the decoded video to enhance the video quality. Summary of the Invention
[0004] Many deep learning-based methods train models under the assumption that the Quantization Parameter (QP) is known, which are non-blind methods. In practical applications, the quantization parameter may not be obtainable, such as devices like televisions that only obtain the decoded video. Therefore, these methods are not very suitable for such blind scenarios where the quantization parameter cannot be obtained. As is well known, compared with blind methods, non-blind methods can train more effective models because they have known information. Therefore, combining the advantages of both, the present invention proposes a semi-blind method for removing VVC compression artifacts based on deep learning for blind scenarios, which mainly includes the following steps:
[0005] (1)A method for predicting video quantization parameters in VVC decoding is proposed.
[0006] (2)A method for removing video compression artifacts in VVC decoding is proposed and models with different quantization parameters are trained.
[0007] (3)According to the quantization parameter prediction method in (1), the quantization parameters of VVC decoded video are predicted, the blind scene without known quantization parameters is converted into a non-blind scene, and the corresponding trained model in (2) is selected according to the result to remove the compression artifacts of VVC decoded video. Description of the Drawings
[0008] Figure 1 It is the overall process of the semi-blind method for removing VVC compression artifacts.
[0009] Figure 2 It is the quantization parameter prediction network and its process.
[0010] Figure 3 The framework of the compression artifact removal network.
[0011] Figure 4 The framework of the motion compensation sub-network.
[0012] Figure 5 The framework of the decompression artifact removal sub-network. Detailed Implementation Manner
[0013] The present invention will be further described below with reference to the accompanying drawings.
[0014] (1) Figure 1 It shows the overall process of the semi-blind method for removing VVC compression artifacts. According to different quantization parameters, the compression artifacts of VVC decoded video are different. As the quantization parameter increases, the quality of the decoded video becomes lower and the compression artifacts become more serious. For the regions with rich texture in the decoded video frames, the distortion is more obvious, while there is no significant manifestation in the regions with rich texture. According to this characteristic, the Kirsh operator is used to extract the regions with rich texture in the video frames as the feature blocks representing the current video quality. Then, a prediction network is proposed to predict the quantization parameters of the feature blocks as the quantization parameters of the current video. The network adopts the idea of a classification network. First, the convolutional layer is used to initially extract the image block features, and then the convolutional layer with the activation function "Leaky ReLU" in the dense connection block and the dense connection method of the BN (Batch Normalization) layer are used to continuously strengthen the features. Then, the average pooling layer is used for downsampling and the transition dense connection block, and finally, regression classification is performed through the fully connected and softmax layers to output the preliminary prediction results. As Figure 2 shown, finally, a voting mechanism is used to select the prediction result with the most occurrences in the feature block results as the quantization parameter of the decoded video. The network is trained using the cross-entropy loss function.
[0015]
[0016] Among them, M represents the category, N represents the total number of samples, and y ic represents the sign function (0 or 1). If the true category of sample i is equal to c, then y ic takes 1, otherwise takes 0, and p ic represents the predicted probability that sample i belongs to category c;
[0017] (2) After predicting the quantization parameters of the VVC decoded video, a compression artifact removal network is proposed. This network mainly uses the spatio-temporal information of the decoded video to remove artifacts. The adjacent frames contain temporal information, and the current frame contains spatial information. However, directly using the temporal information of adjacent frames may lead to errors because there are motion errors between adjacent frames and the current frame. Therefore, a motion compensation sub-network is used to align adjacent frames. The structure of this sub-network is shown in Figure 4 , and this sub-network is trained under the supervision of the mean square error and the current frame,
[0018] F t ' -1 = N mc (F t , F t-1 ) (2)
[0019]
[0020] Among them, N mc (·) represents the motion compensation sub-network, F t represents the current frame, F t-1 represents the adjacent previous frame, F t+1 represents the adjacent next frame, F t ′ -1 represents the compensated previous frame, F t ′ +1 represents the compensated next frame. Formula (2) represents the motion compensation sub-network, and formula (3) represents the training function of the motion compensation sub-network;
[0021] (3) After obtaining the compensated adjacent frames, a de-compression artifact sub-network is proposed to remove the compression artifacts of the VVC decoded video. This network extracts spatio-temporal features through multi-scale convolutional layers and "Concat" operations, and then strengthens, maps features, and reconstructs video frames through convolutional layers and residual learning to achieve the purpose of de-compression. The structure of this sub-network is shown in Figure 5 , and this sub-network is trained through the L1 regularization function and under the supervision of the original frame,
[0022]
[0023] Among them, N ar (·) represents the pseudo-shadow removal network, representing the frame after removing compression artifacts, F R representing the original frame, Equation (4) represents the pseudo-shadow removal network, and Equation (5) represents the training function of the pseudo-shadow removal network;
[0024] (4) According to the quantization parameter prediction method in (1), the quantization parameters of the VVC decoded video are predicted, converting the blind scene without knowing the quantization parameters into a non-blind scene. According to the VVC compression artifact removal network in (2) and (3), models with quantization parameters of 32, 37, and 42 are pre-trained. This is because the larger the quantization parameter, the lower the quality of the decoded video, and the compression artifact removal work for low-quality decoded videos is more meaningful. According to the predicted quantization parameter results, the corresponding model is selected for compression artifact removal.
[0025] To better illustrate the effectiveness of the present invention, Table 1 shows the comparison of ΔPSNR (peak signal-to-noise ratio) between the present invention and other methods when the quantization parameter is 37, indicating that the present invention can effectively remove the compression artifacts of the VVC decoded video, and the experimental results are universal for other test sequences.
[0026] The comparison methods are as follows:
[0027] Method 1: The method proposed by Y. Dai et al., reference “A convolutional neural network approach for post-processing in hevc intra coding” in: International Conference on Multimedia Modeling, Springer, 2017, pp. 28–39.
[0028] Method 2: The method proposed by Yang et al., reference “Decoder-side hevc quality enhancement with scalable convolutional neural network” in: International Conference on Multimedia and Expo, IEEE, 2017, pp. 817–822.
[0029] Method 3: The method proposed by Yang et al., reference "Multi-frame quality enhancement for compressed video" in: Conference on Computer Vision and Pattern Recognition, 2018, pp. 6664 - 6673.
[0030] Table 1 Comparison of ΔPSNR between the present invention and other methods when the quantization parameter is 37
[0031]
Claims
1. A semi-blind method for removing VVC compression artifacts based on deep learning, characterized in that: (1) Use the Kirsh operator to extract texture-rich regions in the video frame as feature blocks representing the current video quality, and propose a quantization parameter prediction network to predict the quantization parameter of the VVC decoded video, converting the blind scenario without knowing the quantization parameter into a non-blind scenario: VVC decoded videos encoded with different quantization parameters have different compression artifacts. As the quantization parameter increases, the compression artifacts become more obvious and the video quality decreases. According to this characteristic, several feature blocks representing the decoded video are extracted, and a quantization parameter prediction network is proposed to predict the quantization parameter for each feature block. Then, using a voting mechanism, the quantization parameter with the most results in the feature blocks is selected as the quantization parameter of the decoded video to achieve the prediction purpose. The quantization parameter prediction network is trained with the cross-entropy loss function as follows, Among them, M represents the category, N represents the total number of samples, and y ic represents the sign function (0 or 1). If the true category of sample i is equal to c, then y ic takes 1, otherwise takes 0, and p ic represents the predicted probability that sample i belongs to category c; (2) Propose a VVC decoded video compression artifact removal network based on the spatio-temporal information of the decoded video: The network combines the spatio-temporal information of the decoded video and uses the spatial information of the current frame of the decoded video and the temporal information of its adjacent frames for compression. Since directly using adjacent frames may introduce negative information, first align the adjacent frames by using a motion compensation sub-network, and then propose a de-artifact sub-network. It extracts spatio-temporal features by using multi-scale convolutional layers and "Concat" operations, and then continuously reconstructs the video frame through convolutional layers to strengthen, map features, and residual learning to achieve the purpose of de-compression. The entire VVC decoded video compression artifact removal network is trained with the joint loss function as follows, L ar = ||F R -F t e || 1 (3) L 去伪 = a * L ar + b * L mc (4) Among them, F t represents the current frame, F′ t-1 represents the previous frame after compensation, F′ t+1 represents the next frame after compensation, represents the frame after removing compression artifacts, F R represents the original frame, L mc represents the motion compensation sub-network, L ar represents the de-artifacting sub-network; (3) Use the network proposed in (2) to train the compression artifact removal network with different quantization parameters, use the network proposed in (1) to predict the quantization parameter of the VVC decoded video, and then use the network trained with the corresponding quantization parameter to remove the compression artifacts of the VVC decoded video: The method uses the quantization parameter prediction network to predict the quantization parameter of the VVC decoded video, converts the scenario without knowing the quantization parameter into a non-blind scenario, and then pre-trains several models with the VVC decoded video compression artifact removal network. Since removing the compression artifacts of low-quality decoded videos is more obvious and more meaningful, and the larger the quantization parameter, the lower the quality of the decoded video, so train models with quantization parameters of 32, 37, and 42. According to the predicted quantization parameter, use the corresponding model to remove the compression artifacts of the VVC decoded video.