VVC block division rapid prediction method based on convolutional neural network
Through the fast prediction method based on convolutional neural network, the problem of large amount of VVC block division is solved, and the encoding time and efficiency are reduced while ensuring quality.
Patent Information
- Application Number
- CN202510478899.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
AI Technical Summary
The existing video encoding standard VVC has a significant increase in the amount of calculation in block division mode, resulting in limited application in scenarios with high real-time requirements, and it is difficult for traditional methods to achieve a good balance between accuracy and computing efficiency.
Using a fast prediction method based on convolutional neural network, the CNN-SVM model is collected, trained and deployed, and the encoding complexity is reduced.
On the premise of ensuring coding quality, it significantly reduces coding time, improves coding efficiency, and reduces calculation complexity.
Smart Images

Figure CN120378619A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video coding and decoding, and particularly relates to a fast prediction method for VVC block partitioning based on a convolutional neural network. Background Art
[0002] With the rapid development of video applications, higher requirements are put forward for video coding efficiency. VVC adopts a block-based hybrid coding framework, actively innovates while inheriting most of the coding tools of HEVC, introduces new coding tools in each key module, and deeply optimizes the original coding tools. Different from previous coding standards, VVC has made a major breakthrough in the block partitioning method. In addition to the traditional quadtree partitioning, it also introduces multi-type tree partitioning, greatly improving the coding flexibility and adaptability.
[0003] However, this QTMT partitioning structure will lead to a significant increase in coding computation, limiting its application in scenarios with high real-time requirements. Traditional block partitioning algorithms are difficult to achieve a good balance between accuracy and computational efficiency. Therefore, a new algorithm is needed to quickly and accurately perform VVC block partitioning and improve coding efficiency.
[0004] Such as Figure 1 , the coding block partitioning adopts a recursive design. Facing coding blocks with obvious horizontal or vertical features, as well as smooth coding blocks, complex multiple partitioning methods cause a lot of ineffective redundant coding. Therefore, to reduce the algorithm complexity, quickly judging the partitioning type of coding blocks is an effective way. The literature "Zhang, D.; Li, Q. An Efficient CU Partition Algorithm for VVC Intra Coding. J. Phys. Conf. Ser. 2021, 1815, 012006." uses the texture information and residual coefficient distribution of CUs to skip low-probability CU partition types in advance. In addition, horizontal or vertical splitting modes are skipped according to the direction of the edges in the block. The literature "Yang, H.; Shen, L.; Dong, X. Low-complexity CTU partition structure decision and fast intra mode decision for versatile video coding. IEEE Trans. Circuits Syst. Video Technol. 2019, 30, 1668–1682." uses a decision tree to transform the QTMT partitioning process into multiple binary classification problems at each decision layer, reducing the coding time.
[0005] In recent years, it has become increasingly difficult to continue improving the coding performance using traditional methods within the video coding framework, and video coding based on neural networks has gradually become possible. Summary of the Invention
[0006] The present invention proposes a fast prediction method for VVC block partitioning based on a convolutional neural network, aiming to reduce the computational complexity of multi-type partitioning of coding units and save coding time on the premise of ensuring coding quality.
[0007] The technical solution to achieve the object of the present invention is: a fast prediction method for VVC block partitioning based on a convolutional neural network, comprising the following steps:
[0008] Step 1: Data collection
[0009] Select several different types of videos, select 1 frame every M frames, and a total of N frames are selected for each video. Encode them with the original VVC encoder under different quantization parameters QP, and collect the following data as the training set: the coding unit (CU) with a size of 32×32, the vertical direction feature, the horizontal direction feature, and the quantization parameter QP as input information, and the category to which the optimal partitioning method of this CU belongs during the coding process: horizontal partitioning, vertical partitioning, non-partitioning, or quadtree partitioning as the corresponding label. Label the CU and its corresponding label according to the following rules: Select according to the coding result, label the CU block containing 32×H (H<32) in the coding result as horizontal partitioning, label the CU block containing W×32 (W<32) in the coding result as vertical partitioning, label the CU that is not further partitioned in the coding result as non-partitioning, and label the CU that continues to be quadtree partitioned in the coding result as a complex form.
[0010] Step 2: Data training
[0011] For the CU in the training set, first de-mean the luminance value of the CU and preprocess it with an improved Canny operator as the input of the convolutional neural network (CNN). After four convolutional calculations, perform max pooling, and input the feature vector extracted by the fully connected layer and the prior knowledge QP parameter into the SVM model together to complete block partitioning classification, and this result corresponds to the label. After training, a CNN-SVM model for fast prediction of VVC block partitioning is obtained.
[0012] Step 3: Model deployment
[0013] During the actual VVC coding process, for each CU with a size of 32×32, input the CU luminance value, vertical direction feature, horizontal direction feature, and QP value into the above-trained CNN-SVM model to obtain the prediction of the optimal partitioning of the current CU, and perform subsequent coding according to this prediction.
[0014] Further, for the labeling of the CUs in step 1 and their corresponding tags, it is selected according to the coding result. The CU blocks containing 32×H (H < 32) in the coding result are labeled as horizontal division, the CU blocks containing W×32 (W < 32) are labeled as vertical division, the CUs that are not further divided in the coding result are labeled as non - division, and the CUs that continue with quadtree division in the coding result are labeled as complex form.
[0015] Further, for the four - time convolution calculation in step 2, the convolution kernel has a size of 3×3, and the convolution kernels of each convolution do not overlap.
[0016] Further, in step 2, the features obtained from the 32×32 - sized CU are concatenated with the features of QP.
[0017] Further, the judgment method of the convolutional neural network built in step 2 is as follows:
[0018] The features after four - layer convolution of the 32×32 - sized CU are concatenated with the features of QP to obtain multiple features, which are then input into a support vector machine (SVM) to judge the division type of the current coding unit.
[0019] Compared with the prior art, the significant advantages of the present invention are: (1) The neural network has strong learning ability and feature extraction ability, and the SVM has high classification accuracy; (2) The network model is small, and the time occupied in the overall coding can be ignored. Brief Description of the Drawings
[0020] Figure 1 is the division structure of the coding unit in VVC.
[0021] Figure 2 is the model training diagram of the fast prediction method for VVC block division based on convolutional neural network of the present invention.
[0022] Figure 3 is the coding flow chart of the fast prediction method for VVC block division based on convolutional neural network of the present invention. Detailed Embodiments
[0023] The present invention uses a convolutional neural network to judge the division method of 32×32 coding blocks in VVC. Combining Figure 2 , the specific steps are as follows:
[0024] Step 1: Data collection. Select several different types of videos. Select 1 frame every M frames, and a total of N frames are selected for each video. Encode with the original VVC encoder under different quantization parameters QP, and collect the following data as the training set: the 32×32-sized coding unit (CU) therein, vertical direction features, horizontal direction features, and quantization parameter QP as input information, and the category to which the optimal partitioning method of this CU belongs during the encoding process: horizontal partitioning, vertical partitioning, non-partitioning, or quadtree partitioning as the corresponding label. Label the CU and its corresponding label according to the following rules: Select according to the encoding result. Label the CU block containing 32×H (H < 32) in the encoding result as horizontal partitioning, label the CU block containing W×32 (W < 32) in the encoding result as vertical partitioning, label the CU that is not further partitioned in the encoding result as non-partitioning, and label the CU that continues to be quadtree partitioned in the encoding result as a complex form.
[0025] Step 2: Data training. For the CU in the training set, first de-mean the luminance value of the CU and preprocess it with an improved Canny operator as the input of the convolutional neural network (CNN). After four convolutional calculations, perform max pooling, and input the feature vector extracted by the fully connected layer and the prior knowledge QP parameter into the SVM model together to complete the block partitioning classification, and this result corresponds to the label. After training, a CNN-SVM model for fast prediction of VVC block partitioning is obtained.
[0026] Step 3: Model deployment. During the actual VVC encoding process, for each 32×32-sized CU, input the CU luminance value, vertical direction features, horizontal direction features, and QP value into the above-trained CNN-SVM model to obtain the prediction of the optimal partitioning of the current CU, and perform subsequent encoding according to this prediction.
[0027] Furthermore, for the labeling of the CU and its corresponding label in Step 1, select according to the encoding result. Label the CU block containing 32×H (H < 32) in the encoding result as horizontal partitioning, label the CU block containing W×32 (W < 32) in the encoding result as vertical partitioning, label the CU that is not further partitioned in the encoding result as non-partitioning, and label the CU that continues to be quadtree partitioned in the encoding result as a complex form.
[0028] Furthermore, for the four convolutional calculations in Step 2, the convolutional kernel has a size of 3×3, and the convolutional kernels of each convolution do not overlap.
[0029] Furthermore, in Step 2, the features obtained from the 32×32-sized CU and the features of QP are concatenated together.
[0030] Furthermore, the judgment method of the convolutional neural network built in Step 2 is as follows:
[0031] The features of the 32×32-sized CU after four layers of convolution are connected with the features of QP to obtain multiple features, which are then input into a support vector machine (SVM) to determine the partitioning type of the current coding unit;
[0032] The following is a further specific description of the technical solution of the present invention through embodiments.
[0033] Embodiment
[0034] The embodiment shown is a fast prediction method for VVC block partitioning based on a convolutional neural network, and its process is as Figure 2 shown, and its steps include:
[0035] Step 1: Data collection. Select video sequences FoodMarket2, PeopleOnStreet, BQTerrace, BQMall, FourPeople, BasketballPass, select 1 frame every 8 frames, and a total of 10 frames are selected for each video. Use the above method to collect the data set;
[0036] Step 2: Input the data set into the network for training, and after training, obtain a network model for fast prediction of VVC block partitioning;
[0037] Step 3: Model deployment: In the actual VVC coding process, for each 32×32-sized CU, input the CU pixel values, CU vertical vector, CU horizontal vector, and QP value into the above-trained network to obtain a prediction of the optimal partitioning of the current CU, and perform subsequent coding according to this prediction. Specifically:
[0038] If the current CU is 32×32 in size, input it into the network for partitioning prediction. If it is judged to be horizontally partitioned, only perform horizontal partitioning coding; if it is judged to be vertically partitioned, only perform vertical partitioning coding; if it is judged not to be partitioned, do not perform subsequent partitioning coding; if it is judged to be in a complex form, code according to the original VTM partitioning method.
[0039] If the current CU is not 32×32 in size, code according to the original VTM partitioning method.
[0040] Compare the performance of the above method with the original VTM6.0 model of VTM. Table 1 gives the coding performance comparison in terms of coding time and BDBR. Among them, BDBR represents the loss of bit rate compared with the original VTM 6.0 method under the condition of the same objective video quality. The coding time compares the savings in coding time compared with the original VTM6.0 method.
[0041] Table 1 Comparison of coding results between the method of the present invention and VTM6.0
[0042]
[0043] Under the condition of ensuring the same coding platform and the number of coding frames, the algorithm proposed in this paper reduces the coding time by an average of 34.24% compared with the standard coding algorithm, and the BDBR increases by 1.02%. While improving the coding efficiency, the coding quality is ensured.
[0044] The present invention is not limited to the content involved in the claims and the above embodiments. Any invention created according to the concept of the present invention should fall within the protection scope of the present invention.
Claims
1. A fast prediction method for VVC block partitioning based on convolutional neural network, characterized in that, It includes the following steps: Step 1: Data collection Select several different types of videos. Select 1 frame every M frames, and a total of N frames are selected for each video. Encode with the original VVC encoder under different quantization parameters QP, and collect the following data as the training set: the coding unit (CU) with a size of 32×32, the vertical direction feature, the horizontal direction feature, and the quantization parameter QP as input information, and the category to which the optimal partitioning method of this CU belongs during the encoding process: horizontal partitioning, vertical partitioning, non-partitioning, or quadtree partitioning as the corresponding label. Label the CU and its corresponding label according to the following rules: Select according to the encoding result. Label the CU block containing 32×H (H < 32) in the encoding result as horizontal partitioning, label the CU block containing W×32 (W < 32) in the encoding result as vertical partitioning, label the CU that is not further partitioned in the encoding result as non-partitioning, and label the CU that continues to be quadtree partitioned in the encoding result as a complex form. Step 2: Data training For the CU in the training set, first de-mean the luminance value of the CU, and preprocess it with an improved Canny operator as the input of the convolutional neural network (CNN). After four convolutional calculations, perform max pooling, and input the feature vector extracted by the fully connected layer and the prior knowledge QP parameter into the SVM model together to complete the block partitioning classification, and this result corresponds to the label. After training, a CNN-SVM model for fast prediction of VVC block partitioning is obtained. Step 3: Model deployment During the actual VVC encoding process, for each CU with a size of 32×32, input the CU luminance value, vertical direction feature, horizontal direction feature, and QP value into the above-trained CNN-SVM model to obtain the prediction of the optimal partitioning of the current CU, and perform subsequent encoding according to this prediction.
2. The fast prediction method for VVC block partitioning based on a convolutional neural network according to claim 1, characterized in that, For the annotation of the CU in Step 1 and its corresponding label, select according to the encoding result. Label the CU block containing 32×H (H < 32) in the encoding result as horizontal partitioning, label the CU block containing W×32 (W < 32) in the encoding result as vertical partitioning, label the CU that is not further partitioned in the encoding result as non-partitioning, and label the CU that continues to be quadtree partitioned in the encoding result as a complex form.
3. The fast prediction method for VVC block partitioning based on a convolutional neural network according to claim 1, characterized in that, For the four convolutional calculations in Step 2, the convolutional kernel has a size of 3×3, and the convolutional kernels of each convolution do not overlap.
4. The fast prediction method for VVC block partitioning based on a convolutional neural network according to claim 1, wherein Step 2 connects the features obtained from the 32×32-sized CU with the features of QP.