Fast Coding Unit Partitioning Method and System under VVC Standard

Through a method based on convolutional neural network, the network is trained using grayscale pixel values, quantization parameters and optimal intra prediction mode to output the partition label of CU, solving the problem of difficult division of texture complex coding units in the prior art, and achieving the effect of reducing calculation complexity and improving coding efficiency.

CN116033153BActive Publication Date: 2025-06-24NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211669820.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-25
Publication Date
2025-06-24
Estimated Expiration
2042-12-25

AI Technical Summary

Technical Problem

When existing video encoding technology deals with coding units with complex textures, it is difficult to effectively divide, resulting in high computational complexity and low encoding efficiency.

Method used

Using a convolutional neural network-based method, the network uses grayscale pixel values, quantization parameters and optimal intra prediction mode as inputs to output the partition label of the CU to achieve rapid partitioning of the CU.

Benefits of technology

This reduces the computational complexity of encoding, improves encoding efficiency, and significantly reduces encoding time while ensuring video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116033153B_ABST
    Figure CN116033153B_ABST
Patent Text Reader

Abstract

The present invention discloses a fast coding unit partitioning method and system under the VVC standard. The method includes: Step 1, collecting data during the coding of different video sequences: the pixel values of a CU with a size of 32×32, quantization parameters, the best intra prediction mode, and the final partitioning label; Step 2, using the data saved in Step 1 to train a constructed convolutional neural network to obtain a model; Step 3, in actual coding, for each CU with a size of 32×32, calling the trained model to input the grayscale values, quantization parameters, and the best intra prediction mode within the CU into the trained network, and the network will output the final partitioning label, and each label corresponds to a different partitioning method. The fast CU partitioning method proposed by the present invention can reduce the coding complexity while ensuring the video quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of video encoding and decoding processing, and particularly relates to a method and system for fast partitioning of coding units under the VVC standard. Background Art

[0002] The rapid development of multimedia devices, the continuous enhancement of people's demand for high-quality video applications, and the high-speed growth of video data volume pose challenges to existing video compression technologies.

[0003] The partitioning methods specified by the VVC standard include quadtree partitioning, binary tree partitioning, and ternary tree partitioning. Among them, binary tree partitioning and ternary tree partitioning can be further divided into horizontal partitioning and vertical partitioning. In addition, the present invention regards non-partitioning as a new partitioning method. Therefore, for most CUs, there are 6 selectable partitioning methods. To determine the final optimal partitioning method, it is necessary to traverse all partitioning methods through the rate distortion optimization (RDO) process. The finally determined optimal partitioning method will have the minimum RD cost. The selection of the CU partitioning method is calculated from top to bottom first and then selected from bottom to top. This method can bring good partitioning effects and improve the encoding efficiency, but it has a high computational complexity.

[0004] Generally, the traditional feature-based CU fast partitioning methods pay more attention to the texture information of the coding block itself. Such methods usually use mathematical tools such as gradients and variances to find the relationship between a certain feature of the CU itself and the final partitioning result. Although there is a strong correlation between the partitioning of the CU and the extracted features, there are still problems. For CUs with complex textures, this method cannot decide whether further partitioning is needed. In addition, operations such as variance and gradient on the CU will introduce unnecessary calculations.

[0005] In recent years, deep learning has been widely used in image classification problems. Therefore, it has become a general trend to use the method based on convolutional neural networks to handle the CU fast partitioning problem. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and system for fast partitioning of coding units under the VVC standard by using a convolutional neural network, so as to achieve the purpose of reducing the computational complexity.

[0007] The technical solution for achieving the purpose of the present invention is as follows: In the first aspect, the present invention provides a method for fast partitioning of coding units under the VVC standard, including:

[0008] Step 1: Select 12 standard test sequences provided by VVC. For each video sequence, select 15 frames of data for data collection. During the data collection process, save the data with horizontal division or vertical division traces. Specifically, the division trace means that if the coding result contains a CU of 32×H and H < 32, it is marked as horizontal division, and if the coding result contains a CU of W×32 and W < 32, it is marked as vertical division. The specific data to be saved includes: the grayscale pixel values of the 32×32-sized CU, the quantization parameter when encoding the current CU, the best intra-prediction mode selected for the current CU, and the optimal label value corresponding to the current CU in the original encoder. The specific labels are divided into three categories: horizontal division, vertical division, and other division cases;

[0009] Step 2: Input the collected data into the network group by group. The 32×32 CU first forms a 64×1×1 global feature after four layers of convolution. Then, the QP with a dimension of 1×1 and the best intra-prediction mode with a dimension of 1×1 are transformed through a fully connected layer to 32×1×1. Subsequently, it is superimposed with the 64×1×1 global feature after convolution, and the dimension becomes 96×1×1 after superposition. Finally, the 96×1×1 passes through two 1×1 convolutional layers to finally output the prediction probability of the current CU division, corresponding to the collected data label. Thus, the network training is completed;

[0010] Step 3: During the actual coding process, when the encoder encodes a 32×32-sized CU, load the convolutional neural network model, input the grayscale pixel values of the 32×32-sized CU, and the network model will give the final division label, and realize the early decision of the division method according to the final label.

[0011] In a second aspect, the present invention provides a fast coding unit division system under the VVC standard, including:

[0012] A data collection module, which is used to collect data for training the convolutional neural network. Select 12 standard test sequences provided by VVC. For each video sequence, select 15 frames of data for data collection. During the data collection process, save the data with horizontal division or vertical division traces. Specifically, the division trace means that if the coding result contains a CU of 32×H and H < 32, it is marked as horizontal division, and if the coding result contains a CU of W×32 and W < 32, it is marked as vertical division. The specific data to be saved includes: the grayscale pixel values of the 32×32-sized CU, the quantization parameter when encoding the current CU, the best intra-prediction mode selected for the current CU, and the label value corresponding to the current CU. The specific labels are divided into three categories: horizontal division, vertical division, and other division cases;

[0013] The network training module is used to input the collected data into the network in groups. The 32×32 CU first forms a 64×1×1 global feature after four convolutional layers. Then, the QP with a dimension of 1×1 and the best intra prediction mode with a dimension of 1×1 are transformed through a fully connected layer into 32×1×1. Subsequently, it is superimposed with the 64×1×1 global feature after convolution, and the dimension becomes 96×1×1 after superposition. Finally, the 96×1×1 is passed through two 1×1 convolutional layers to finally output the prediction probability of the current CU partition, corresponding to the label of the collected data. Thus, the network training is completed;

[0014] The model loading module is used to input the pixel values of the 32×32 CU, quantization parameters, and the best intra prediction mode into the trained network during the actual encoding process. The network outputs the final partition label of the CU for making an early decision on the CU partition method.

[0015] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method described in the first aspect.

[0016] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in the first aspect.

[0017] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the method described in the first aspect.

[0018] Compared with the existing algorithms, the present invention has the following advantages:

[0019] (1) Compared with the traditional feature-based CU fast partitioning method, the convolutional neural network-based CU fast partitioning method does not require manual screening of features related to the partitioning result, and only needs to input the grayscale pixel values of the CU into the network;

[0020] (2) The convolutional neural network-based CU fast partitioning method does not introduce a large amount of additional calculations, and the loading and prediction of the convolutional neural network account for a very small proportion of the time in the entire encoding process. Description of the Drawings

[0021] Figure 1 It is a flowchart of the CU fast partitioning method based on a convolutional neural network.

[0022] Figure 2 It is a multi-type tree partitioning structure under the VVC standard.

[0023] Figure 3It is a convolutional neural network structure diagram under the CU fast partitioning algorithm based on convolutional neural network. Detailed implementation manner

[0024] Combined with Figures 1 - 3 , the present invention provides a fast partitioning method for coding units (CUs) under the video VVC standard. For CUs with a size of 32×32, the specific processing steps for making a fast decision on the partitioning method are as follows:

[0025] Step 1, for different video sequences, N frames of data are selected for each video sequence for data acquisition. During the data acquisition process, the data with obvious horizontal partitioning or vertical partitioning traces is saved. The specific data saved includes: the grayscale pixel values of the CUs with a size of 32×32, the quantization parameters when encoding the current CU, the best intra prediction mode selected for the current CU, and the label value corresponding to the current CU. The specific labels are divided into three categories: horizontal partitioning, vertical partitioning, and other partitioning situations;

[0026] Step 2, the collected data is input into the network in groups. The 32×32 CU first forms a 64×1×1 global feature after four convolutions. Then, the QP with a dimension of 1×1 and the best intra prediction mode with a dimension of 1×1 are transformed through a fully connected layer into 32×1×1. Subsequently, it is superimposed with the 64×1×1 global feature after convolution, and the dimension becomes 96×1×1 after superimposition. Finally, the 96×1×1 passes through two 1×1 convolutional layers, and finally outputs the prediction probability of the current CU partitioning, corresponding to the collected data label. Thus, the network training is completed;

[0027] Step 3, during the actual encoding process, when the encoder encodes a CU with a size of 32×32, the convolutional neural network model is loaded, and the grayscale pixel values of the CU with a size of 32×32 are input. The network model will give the final partitioning label, and the early decision of the partitioning method is realized according to the final label.

[0028] For the data acquisition in Step 1, if there are obvious horizontal partitioning traces in the final partitioning result, that is, the size of this block is 32×W, W×16, it indicates that this block uses horizontal binary tree partitioning or horizontal ternary tree partitioning. Therefore, the label of this block is set to 0 and then collected; if there are obvious vertical partitioning traces in the final partitioning result, that is, the size of this block is H×32, H×16, it indicates that this block uses vertical binary tree partitioning or vertical ternary tree partitioning. Therefore, the label of this block is set to 1 and then collected; in other cases, they are all collected as the case with label 2.

[0029] For the four convolution operations in Step 2, four convolution kernels and the same stride are set.

[0030] Connect the features obtained after the convolutional layer in step 2 with the quantization parameter QP and the best intra prediction mode PredMode through a fully connected layer.

[0031] The present invention will be described in detail below with reference to embodiments.

[0032] Embodiment

[0033] The embodiment shown herein is a fast partitioning method for coding units under the VVC standard, and its process is as Figure 1 shown, and its steps include:

[0034] Step 1: Data collection. Select 12 video sequences, namely BQTerrace, basketballdrill, basketballDrive, fourpeople, bubble, johnny, chinaspeed, BQsquare, basketballpass, Cactus, Racehorses, BQMall. Select 1 frame every 8 frames, and a total of 15 frames are selected for each video. Collect data in the manner of step 1;

[0035] Step 2: Input the data set into the network for training, and after training, obtain a network model for fast partitioning prediction of VVC coding blocks;

[0036] Step 3: Model deployment: During the actual VVC coding process, for each CU with a size of 32×32, input the luminance value, QP value, and best intra prediction mode value of the CU into the above-trained network to obtain a prediction of the optimal partitioning of the current CU, and implement an early decision on the partitioning method according to the output value of the network. Specifically:

[0037] If the size of the current CU is 32×32, input it into the network for partitioning prediction. If it is judged to be horizontally partitioned, only perform horizontal partitioning coding; if it is judged to be vertically partitioned, only perform vertical partitioning coding; if it does not meet the above two cases, classify it as other cases, and select the partitioning method according to the original VTM process. If the size of the current CU does not belong to 32×32, encode it according to the original VTM partitioning method.

[0038] Compare the above method with the original VTM10.0 model in terms of performance. Table 1 gives the specific experimental results, and the experimental results use three measurement indicators: PSNR, BDBR, and time saving rate. Among them, BDBR represents the saving of the bit rate compared with the original VTM 10.0 method under the condition of the same objective video quality. The comparison of the encoding time is the saving of the encoding time compared with the original VTM10.0 method.

[0039] Table 1 Comparison of Coding Results between the Method of the Present Invention and the Method of VTM6.0

[0040]

[0041] The fast CU partitioning method proposed by the present invention can reduce the coding complexity while ensuring the video quality.

[0042] The above are only the preferred embodiments of the present invention. It should be noted that for those skilled in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A fast partitioning method for coding units under the VVC standard, characterized in that, Including: Step 1: Select 12 standard test sequences provided by VVC. For each video sequence, select 15 frames of data for data collection. During the data collection process, save the data with horizontal or vertical division traces. Specifically, the division trace means that if the coding result contains a CU of 32×H and H < 32, it is marked as horizontal division, and if the coding result contains a CU of W×32 and W < 32, it is marked as vertical division; The specific data to be saved includes: the grayscale pixel values of 32×32-sized CUs, the quantization parameters when encoding the current CU, the best intra-prediction mode selected for the current CU, and the optimal label value corresponding to the current CU in the original encoder. The specific labels are divided into three categories: horizontal division, vertical division, and other division cases; Step 2: Input the collected data into the network group by group. The 32×32 CU first forms a 64×1×1 global feature after four convolutional layers. Then, the QP with a dimension of 1×1 and the best intra-prediction mode with a dimension of 1×1 are transformed through a fully connected layer into 32×1×1. Subsequently, it is superimposed with the 64×1×1 global feature after convolution, and the dimension becomes 96×1×1 after superposition. Finally, the 96×1×1 passes through two 1×1 convolutional layers to finally output the prediction probability of the current CU division, corresponding to the collected data label. Thus, the network training is completed; Step 3: During the actual coding process, when the encoder encodes a 32×32-sized CU, load the convolutional neural network model, input the grayscale pixel values of the 32×32-sized CU, and the network model will give the final division label, and realize the early decision of the division method according to the final label.

2. The fast partitioning method of coding units under the VVC standard according to claim 1, wherein For the data collection in Step 1, if there is a horizontal division trace in the final division result, that is, the size of the coding unit is 32×W or W×16, it means that the coding unit uses horizontal binary tree division or horizontal ternary tree division. Therefore, set the coding unit label to 0 and then collect it; if there is a vertical division trace in the final division result, that is, the size of the coding unit is H×32 or H×16, it means that the coding unit uses vertical binary tree division or vertical ternary tree division. Therefore, set the coding unit label to 1 and then collect it; in other cases, regard it as the case with label 2 for collection.

3. The fast partitioning method of coding units under the VVC standard according to claim 1, wherein, For the four convolutional operations in Step 2, set the four convolutional kernels and the strides to be the same.

4. A fast coding unit partitioning system under the VVC standard, characterized in that, Including: A data collection module for collecting data for training the convolutional neural network. Select 12 standard test sequences provided by VVC. For each video sequence, select 15 frames of data for data collection. During the data collection process, save the data with horizontal or vertical division traces. Specifically, the division trace means that if the coding result contains a CU of 32×H and H < 32, it is marked as horizontal division, and if the coding result contains a CU of W×32 and W < 32, it is marked as vertical division; The specific data saved include: the grayscale pixel value of the CU of size 32×32, the quantization parameter when encoding the current CU, the best intra prediction mode selected by the current CU, and the label value corresponding to the current CU. The specific labels are divided into three categories: horizontal division, vertical division and other division situations; The network training module is used to input the collected data into the network in groups. The 32×32 CU first passes through four layers of convolution to form a 64×1×1 global feature. Then the QP with a dimension of 1×1 and the best intra-frame prediction mode with a dimension of 1×1 are transformed through a fully connected layer to 32×1×1. Then, it is superimposed with the 64×1×1 global feature after convolution. After superposition, the dimension becomes 96×1×1. Finally, the 96×1×1 passes through two 1×1 convolution layers, and finally outputs the predicted probability of the current CU division, which corresponds to the collected data label. At this point, the network training is completed; The model loading module is used to input the 32×32 CU pixel value, quantization parameter and optimal intra-frame prediction mode into the trained network during the actual encoding process. The network outputs the final partition label of the CU to decide the CU partition method in advance.

5. The fast partitioning system for coding units under the VVC standard according to claim 4, wherein For the data acquisition module, if there are traces of horizontal division in the final division result, that is, the size of the coding unit is 32×W, W×16, it means that the coding unit adopts horizontal binary tree division or horizontal ternary tree division, so the coding unit label is set to 0 for collection; if there are traces of vertical division in the final division result, that is, the size of the coding unit is H×32, H×16, it means that the coding unit adopts vertical binary tree division or vertical ternary tree division, so the coding unit label is set to 1 for collection; in other cases, they are all regarded as label 2 for collection.

6. The fast partitioning system for coding units under the VVC standard according to claim 4, wherein For the four convolution operations, four convolution kernels and the same step size are set.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method according to any one of claims 1 to 3 is implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.

9. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Early termination method for reduction and block division of VVC intra-frame coding unit candidate prediction modes

    CN109688414A

  • 3D video depth map intra-frame fast coding method based on deep neural network

    CN112770120A