CNN-based VCC coding mode rapid selection method

Through the fast selection method of VCC encoding mode based on CNN, the encoding mode probability is predicted and the encoding path is optimized, which solves the problem of high complexity in VVC encoding technology, and achieves efficient coding efficiency and mass balance.

CN120238655APending Publication Date: 2025-07-01CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510597920.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

While improving coding efficiency, the existing VVC encoding technology has led to a significant increase in algorithm complexity, making it difficult to meet the processing needs of real-time transmission and low-power devices.

Method used

The CNN-based VCC encoding mode fast selection method is adopted, and the encoding mode probability of different sizes of CUs is predicted through the trained encoding mode selection model, and the encoding path is optimized by combining the decision tree model.

Benefits of technology

It significantly reduces the encoding complexity, improves encoding efficiency, and maintains encoding quality, providing a feasible solution for real-time encoding and mobile deployment of ultra-high-definition videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238655A_ABST
    Figure CN120238655A_ABST
Patent Text Reader

Abstract

The invention provides a CNN-based VCC coding mode rapid selection method, and the method comprises the steps: inputting a brightness component of a CTU into a trained VCC coding mode selection model, and obtaining the prediction probability of each coding mode of CU of different sizes; the encoding modes are sorted in a descending order, and an encoder preferentially checks the encoding mode with the highest probability; according to a preset coding mode threshold value, a decision tree model is adopted to judge whether to terminate the check of the coding mode in advance; judging whether the coding mode selected by the CU is an Intra mode or not; if the mode is the Intra mode, inputting the brightness component into a transformation mode selection model, and outputting the prediction probability of each transformation mode of the CUs with different sizes; the transformation modes are sorted in a descending order, and the encoder preferentially checks the transformation mode with the highest probability; according to a preset transformation mode threshold value, judging whether to terminate the inspection of the transformation mode in advance or not by adopting a decision tree model; and selecting a transformation mode to complete CU coding. The method can ensure the coding quality and significantly improve the coding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of video coding, and particularly relates to a method for quickly selecting a VCC coding mode based on CNN. Background Art

[0002] With the rapid development of applications such as ultra-high-definition video, real-time communication, and screen content sharing, higher requirements are put forward for the compression efficiency of video coding standards. As a new generation of international video coding standard, H.266 / Versatile Video Coding (VVC) can reduce the bit rate by 30%-50% compared with HEVC under the same visual quality by introducing innovative technologies such as multi-type tree partitioning (QTMT) and screen content coding (SCC). However, the improvement of coding efficiency has led to a significant increase in algorithm complexity. For example, the QTMT structure supports the flexible combination of quadtrees, binary trees, and ternary trees, expanding the partitioning method of each coding unit (CU) from a limited number of modes in HEVC to hundreds of possibilities, resulting in an exponential growth of the partitioning search space. At the same time, the SCC technology for screen content (such as text and graphics) further introduces modes such as intra-block copy (IBC) and palette (PLT), which need to be independently evaluated under various CU segmentation methods, exacerbating the computational and memory access overhead. Such a trade-off problem between complexity and efficiency restricts the wide application of VVC.

[0003] In the field of intra-frame coding, VVC has significantly improved the compression performance of texture and edge regions by introducing 65 directional prediction modes, multi-transform selection (MTS), low-frequency non-separable transform (LFNST), etc. However, the superimposed effect of technical complexity has led to an exponential growth in the amount of calculation: in the rough mode decision (RMD) stage, the SATD values of hundreds of modes need to be calculated, and in the transform selection stage, various transform kernel combinations need to be traversed and intensive matrix operations need to be performed. In the VTM reference model, intra-frame prediction accounts for more than 60% of the overall coding time, and the calculation of the transform selection module accounts for more than 40%. In the 4K / 8K ultra-high-definition video scenario, the single-frame coding time can reach dozens of seconds, which is difficult to meet the processing requirements of real-time transmission and low-power devices. Although some studies have proposed mode screening methods based on heuristic rules, their generalization ability in complex textures and dynamic screen content is limited, which is prone to performance loss.

[0004] As an important application scenario of VVC, the particularity of SCC further exacerbates the complexity. Under the QTMT partitioning structure, SCC needs to synchronously optimize the CU segmentation method and coding modes (such as IBC and PLT) for high-frequency content such as text and charts. Since IBC and PLT need to be independently evaluated under multiple CU sizes and have a strong coupling with the intra prediction mode, the search space of their joint decision-making shows a trend of combinatorial explosion. For example, a 64×64 CU may generate more than 200 partitioning and mode combinations under the QTMT structure, and traditional exhaustive search will result in a huge amount of rate-distortion cost (RD Cost) calculation and memory access overhead. Although existing technologies attempt to reduce some of the computational amount, it will seriously affect the coding quality. Therefore, there is an urgent need for a method that can significantly improve the coding efficiency while ensuring the coding quality. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the present invention provides a fast VCC coding mode selection method based on CNN, which includes:

[0006] S1: Obtain the luminance component of the CTU and preprocess it, and input the preprocessed CTU into the trained VCC coding mode selection model to obtain the prediction probabilities of each coding mode of different-sized CUs; where the coding modes include PLT mode, IBC mode, and Intra mode;

[0007] S2: Sort the coding modes in descending order according to the prediction probabilities of the coding modes, and the encoder preferentially checks the coding mode with the highest probability;

[0008] S3: According to the preset coding mode threshold, use a decision tree model to judge whether to terminate the inspection of the coding mode in advance. If the inspection is terminated, output the optimal coding mode; otherwise, replace the coding mode in the sorting order and continue the inspection until the selection of the coding mode is completed;

[0009] S4: Judge whether the coding mode selected by the CU is the Intra mode. If so, execute step S5; otherwise, complete the CU coding according to the selected coding mode;

[0010] S5: Input the luminance component of the CTU into the transform mode selection model for processing, and output the prediction probabilities of each transform mode of different-sized CUs; where the transform modes include BDPCM, DCT-II, and other higher-order transforms;

[0011] S6: Sort the transform modes in descending order according to the prediction probabilities of the transform modes, and the encoder preferentially checks the transform mode with the highest probability;

[0012] S7: According to the preset transformation mode threshold, use the decision tree model to determine whether to terminate the inspection of the transformation mode in advance. If the inspection is terminated, output the optimal transformation mode; otherwise, replace the transformation mode in the sorting order and continue the inspection until the selection of the transformation mode is completed;

[0013] S8: The selected Intra-mode CUs complete the CU coding according to the Intra mode and the selected transformation mode.

[0014] Preferably, the VCC coding mode selection model includes a main network and three sub-networks, namely the first sub-network, the second sub-network and the third sub-network; the main network is used to extract basic features from the luminance component of the CTU, and the sub-networks process CUs of different sizes respectively.

[0015] Preferably, the process of the main network processing the luminance component of the CTU includes:

[0016] Extract basic features from the luminance component of the CTU through the convolutional layer conv1 to obtain basic features;

[0017] Use the first sub-network to process the basic features and output the coding mode probabilities of CUs with sizes of 64×64, 32×32, 16×16, 8×8 and 4×4;

[0018] Use the second sub-network to process the basic features and output the coding mode probabilities of CUs with sizes of 32×16, 32×8, 32×4, 16×8, 16×4 and 8×4;

[0019] Use the third sub-network to process the basic features and output the coding mode probabilities of CUs with sizes of 16×32, 8×32, 4×32, 8×16, 4×16 and 4×8.

[0020] Preferably, the process of the first sub-network processing the basic features includes:

[0021] Use consecutive convolutional layers conv2~conv5 to downsample the basic features to obtain the first, second, third and fourth local features; among them, the number of channels of the convolutional layer doubles step by step;

[0022] Use consecutive deconvolutional layers deconv1~deconv4 to upsample the fourth local feature layer by layer to obtain the first, second, third and fourth global features;

[0023] After processing the fourth local feature through the convolutional layer conv6 and the Softmax function, output the coding mode probability of the CU with a size of 64×64;

[0024] Concatenate the third local feature and the first global feature. After processing the concatenated feature through the convolutional layer conv7 and the Softmax function, output the coding mode probability of the CU with a size of 32×32;

[0025] Concatenate the second local feature and the second global feature. After processing the concatenated feature through the convolutional layer conv8 and the Softmax function, output the coding mode probability of the CU with a size of 16×16;

[0026] Concatenate the first local feature and the third global feature. After processing the concatenated feature through the convolutional layer conv9 and the Softmax function, output the coding mode probability of the CU with a size of 8×8;

[0027] Concatenate the basic feature and the fourth global feature. After processing the concatenated feature through the convolutional layer conv10 and the Softmax function, output the coding mode probability of the CU with a size of 4×4.

[0028] Preferably, the processing process of the second subnet for the basic feature includes:

[0029] Use consecutive convolutional layers conv11 - conv15 to downsample the basic feature to obtain the fifth, sixth, seventh, eighth, and ninth local features; among them, the number of channels in the convolutional layer doubles step by step;

[0030] Use the convolutional layer conv31 to downsample the sixth local feature to obtain the tenth local feature;

[0031] After processing the tenth local feature through the convolutional layer conv14 and the Softmax function, output the coding mode probability of the CU with a size of 16×8;

[0032] Use consecutive transposed convolutional layers deconv5~deconv8 to upsample the ninth local feature layer by layer to obtain the fifth, sixth, seventh, and eighth global features;

[0033] After processing the ninth local feature through the convolutional layer conv16 and the Softmax function, output the coding mode probability of the CU with a size of 32×16;

[0034] Concatenate the eighth local feature and the fifth global feature. After processing the concatenated feature through the convolutional layer conv17 and the Softmax function, output the coding mode probability of the CU with a size of 32×8;

[0035] Concatenate the seventh local feature and the sixth global feature. After processing the concatenated feature through the convolutional layer conv18 and the Softmax function, output the coding mode probability of the CU with a size of 32×4;

[0036] Concatenate the sixth local feature and the seventh global feature. After processing the concatenated feature through the convolutional layer conv19 and the Softmax function, output the coding mode probability of the CU with a size of 16×4 at the back;

[0037] Concatenate the fifth local feature and the eighth global feature. After processing the concatenated feature through the convolutional layer conv20 and the Softmax function, output the coding mode probability of the CU with a size of 8×4 at the back.

[0038] Preferably, the processing process of the third subnet for the basic feature includes:

[0039] Use consecutive convolutional layers conv21 - conv25 to downsample the basic feature to obtain the tenth, eleventh, twelfth, thirteenth, and fourteenth local features; among them, the number of channels in the convolutional layer doubles step by step;

[0040] Use the convolutional layer conv32 to downsample the eleventh local feature to obtain the fifteenth local feature;

[0041] After processing the fifteenth local feature through the convolutional layer conv12 and the Softmax function, output the coding mode probability of the CU with a size of 8×16 at the back;

[0042] Upsample the fourteenth local feature layer by layer through consecutive deconvolutional layers deconv9~deconv12 to obtain the ninth, tenth, eleventh, and twelfth global features;

[0043] After processing the fourteenth local feature through the convolutional layer conv26 and the Softmax function, output the coding mode probability of the CU with a size of 16×32 at the back;

[0044] Concatenate the thirteenth local feature and the ninth global feature. After processing the concatenated feature through the convolutional layer conv27 and the Softmax function, output the coding mode probability of the CU with a size of 8×32 at the back;

[0045] Concatenate the twelfth local feature and the tenth global feature. After processing the concatenated feature through the convolutional layer conv28 and the Softmax function, output the coding mode probability of the CU with a size of 4×32 at the back;

[0046] Concatenate the eleventh local feature and the eleventh global feature. After processing the concatenated feature through the convolutional layer conv29 and the Softmax function, output the coding mode probability of the CU with a size of 4×16 at the back;

[0047] Concatenate the tenth local feature and the twelfth global feature. After processing the concatenated feature through the convolutional layer conv30 and the Softmax function, output the coding mode probability of the CU with a size of 4×8 at the back.

[0048] Preferably, the preset coding mode threshold includes the IBC mode threshold T Intra of 0.8, the PLT mode threshold T PLT of 0.6, the IBC Serch mode threshold T IBC-Serch of 0.7 and the IBC Merge mode threshold T IBC-Merg of 0.7.

[0049] Preferably, the network structures of the VCC coding mode selection model and the transformation mode selection model are the same.

[0050] Preferably, the preset transformation mode threshold includes the BDPCM mode threshold T BDPCM of 0.7, the DCT-II mode threshold T DCT-I of 0.6 and the other high-order transformation mode threshold T Other of 0.6.

[0051] The beneficial effects of the present invention are as follows:

[0052] The present invention provides a method for quickly selecting a VCC coding mode based on CNN. By using the VCC coding mode selection model based on CNN to predict the coding mode probability, it solves the efficiency bottleneck caused by the traditional method relying on manual rules. And by combining the probability sorting encoder to preferentially check the coding mode with the highest probability, according to the preset coding mode threshold, using the decision tree model to judge whether to terminate the check of the coding mode in advance and skip the redundant calculation, the coding complexity is significantly reduced. At the same time, for the transformation mode in the Intra mode, it is selected to further optimize the calculation path, achieving an efficient balance between coding efficiency and coding quality, and providing a feasible solution for real-time coding of ultra-high-definition videos and mobile deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following introduces the related technical solution drawings of the embodiments of the present invention. It should be understood that the drawings introduced below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0054] Figure 1 is a schematic flow chart of a method for quickly selecting a VCC coding mode based on CNN provided in an embodiment of the present invention;

[0055] Figure 2 is a schematic diagram of the main network architecture of CNN in an embodiment of the present invention;

[0056] Figure 3 is a schematic diagram of the first subnet structure of CNN in an embodiment of the present invention;

[0057] Figure 4 It is a schematic diagram of the second subnet structure of the CNN in the embodiment of the present invention;

[0058] Figure 5 It is a schematic diagram of the third subnet structure of the CNN in the embodiment of the present invention;

[0059] Figure 6 It is a schematic diagram showing the loss of coding efficiency of the coding mode at different T values in the embodiment of the present invention;

[0060] Figure 7 It is a schematic diagram showing the loss of coding efficiency of the transform mode at different T values in the embodiment of the present invention. Detailed implementation manners

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0062] The embodiment of the present invention provides a method for quickly selecting a VCC coding mode based on CNN, and the method steps are as Figure 1 shown.

[0063] In the intra-frame coding framework of VVC SCC in the embodiment of the present invention, the best coding mode is selected for CUs of different sizes. There are differences in the processing characteristics of different coding modes, and the adaptability of CUs in different content scenarios is also different. Specifically, the IBC mode relies on the intra-frame search mechanism and reuses highly similar blocks found in the current frame, so it performs excellently in the SCC scenario. Since small-sized CUs can provide more accurate block matching and reduce the matching error, the application of the IBC mode on these small-sized CUs has more advantages. In contrast, although larger-sized CUs usually contain more complex texture information, due to the relatively concentrated color distribution of screen content, the palette-based PLT mode can usually encode more effectively and reduce the bitrate overhead.

[0064] In the embodiments of the present invention, in order to effectively predict candidate coding modes and exclude low-probability options, a systematic analysis of the coding mode distributions of different types of CUs is carried out. Using the VTM-17.0 encoder, coding tests are performed on 13 representative test sequences under the CTC standard. These sequences cover various screen content types such as images, texts, and mixed content, and can comprehensively reflect the coding characteristics in different scenarios. Under the AI configuration, the test sequences are coded using QP22, 27, 32, and 37 respectively, and the coding mode distributions of CUs of different sizes are counted. The coding mode distribution is shown in Table 1. It can be seen that the selection of the coding mode has an obvious correlation with the size of the CU. Larger CUs tend to use the PLT mode, while smaller CUs are more suitable for the IBC mode.

[0065] Table 1 Coding mode distribution of coding units of different sizes

[0066]

[0067] In the implementation of the present invention, the PLT mode depends on the color distribution inside the CU and does not need to consider spatial correlation. Therefore, for larger CUs with relatively uniform colors and simple textures, the PLT mode can provide efficient color index coding, thereby reducing bitrate consumption. At the same time, the PLT mode avoids the cumulative error caused by pixel-level prediction and can significantly improve the coding efficiency in screen content coding. The IBC mode searches for the optimal matching block in the encoded area through the intra-frame search mechanism and uses the block vector to indicate the best matching position of the current CU. Therefore, it is particularly efficient in scenarios with repeated patterns, text edges, or screen GUI interface elements. Since smaller CUs can provide a finer-grained matching area, the IBC mode can more accurately utilize the intra-frame redundant information, thereby improving the coding compression efficiency.

[0068] In the implementation of the present invention, the selection of the Intra mode in the CU presents specific applicability rules. Different from the PLT and IBC modes, the Intra mode mainly performs prediction based on the spatial correlation of adjacent pixels, so it performs better in regions with strong spatial correlation. In screen content coding, for example, in regions with gradient backgrounds or low contrast, it is difficult to effectively compress through the IBC or PLT mode. At this time, the Intra mode can provide good reconstruction quality at a lower bit rate. Therefore, in regions where the texture is relatively uniform or the edge information is not obvious, the Intra mode often becomes the preferred encoding scheme. It should be noted that the PLT mode and the IBC mode will generate additional overhead during the encoding process. For example, the PLT mode needs to store palette and index information, while the IBC mode needs to additionally record block vector information, and these overheads will affect the encoding efficiency to a certain extent. Therefore, in practical applications, the encoder needs to comprehensively consider the content characteristics of the CU during the mode selection stage to achieve the best balance between the bit rate and the encoding complexity. For regions with relatively simple textures, although all three modes of PLT, IBC, and Intra can achieve effective encoding, the Intra mode can usually complete the encoding task at a lower code rate, so it has more advantages in low-complexity regions.

[0069] In the embodiments of the present invention, there are extremely significant differences in the complexities presented by different coding modes. At the same time, there is an obvious internal correlation between the size of the CU and the coding mode it adapts to. In VVC, the default size of the CTU is set to 128×128. This standard has high flexibility and allows the CTU to be finely divided, and the minimum can be divided into blocks as small as 4×4. Based on the unique QTMT coding structure of VVC, a CTU can be divided into up to 17 different size specifications of CUs. Predicting each CU individually will bring considerable complexity to the prediction process. Therefore, the embodiments of the present invention combine the characteristics of coding modes, CU sizes, etc., and design dedicated mode predictors for 17 different size CUs.

[0070] The embodiments of the present invention select the luminance component of a CTU as the initial input data and first perform a preprocessing operation of mean removal on it. The key role of this step is to eliminate the DC component in the data, so that the subsequent feature extraction process can focus more on the effective features of the data and improve the accuracy and efficiency of feature extraction. After the preprocessing is completed, the data enters the convolutional layer for feature extraction. Through multiple levels of alternating convolutional and deconvolutional operations, the feature information in the data is deeply mined, and finally the mode probabilities corresponding to 17 different size CUs are generated.

[0071] The VCC coding mode selection model constructed in the embodiments of the present invention includes a main network and three sub-networks, namely the first sub-network, the second sub-network, and the third sub-network, which cooperate to complete the mode prediction tasks for CUs of different sizes. The main network undertakes the core responsibility of extracting basic features from the luminance component of the CU, while the three sub-networks focus on processing CUs of different sizes in square, horizontal, and vertical directions respectively according to the shape characteristics of the CU.

[0072] In the embodiments of the present invention, the main network receives a luminance component of size 64×64 as input, and the structure of the main network is as Figure 2 shown. At the initial stage of the main network, a 4×4 convolutional layer is used to perform preliminary feature extraction on the input data. This convolutional layer can capture the basic features in the input data and lay a foundation for subsequent in-depth feature mining. The first sub-network, the second sub-network, and the third sub-network are used to process the basic features respectively, and gradually increase the number of channels of the feature map through multiple downsampling operations. The downsampling operation can not only effectively reduce the data volume and computational complexity, but also aggregate and abstract the features at different scales, enabling the network to learn more critical and representative feature information. After completing multiple convolutions and downsamplings, the feature maps generated by the convolutional layer and the transposed convolutional layer are concatenated. The function of the transposed convolutional layer is opposite to that of the convolutional layer, and its main role is to enlarge the size of the feature map and restore some of the spatial information lost during the downsampling process. After concatenation, these feature maps are merged, and finally a last group of feature maps with a kernel size of 1×1 and a stride of 1 is generated. This group of feature maps highly condenses the global feature information of the input data, and based on this, an accurate prediction of the CU mode is completed.

[0073] In the embodiments of the present invention, the first sub-network processes the basic features, and the output results include coding mode probabilities of sizes 64×64, 32×32, 16×16, 8×8, and 4×4. The structure of the first sub-network is as Figure 3As shown. The kernel size of convolutional layers conv2 to conv5 is set to 2×2. According to the characteristics of the non-overlapping CU partitioning structure, the stride of conv2 - conv5 is precisely set to the width of the non-overlapping convolutional kernel size to downsample the basic features, obtaining the first, second, third, and fourth local features. By setting it this way, in the feature map, the receptive field of each node can always maintain the same size as a CU. In this way, the feature maps generated by conv2 to conv5 can accurately reflect the local features of CUs in the range from 4×4 to 64×64. In each downsampling step, the number of channels of the feature map is doubled. This operation enables the network to learn richer feature information at different scales and enhances the network's ability to express the local features of CUs. Then, the fourth local feature is upsampled layer by layer through consecutive deconvolutional layers deconv1 to deconv4 to obtain the first, second, third, and fourth global features; after the fourth local feature is processed by convolutional layer conv6 and the Softmax function, the encoding mode probability of the CU with a size of 64×64 is output; the third local feature and the first global feature are concatenated, and after the concatenated feature is processed by convolutional layer conv7 and the Softmax function, the encoding mode probability of the CU with a size of 32×32 is output; the second local feature and the second global feature are concatenated, and after the concatenated feature is processed by convolutional layer conv8 and the Softmax function, the encoding mode probability of the CU with a size of 16×16 is output; the first local feature and the third global feature are concatenated, and after the concatenated feature is processed by convolutional layer conv9 and the Softmax function, the encoding mode probability of the CU with a size of 8×8 is output; the basic feature and the fourth global feature are concatenated, and after the concatenated feature is processed by convolutional layer conv10 and the Softmax function, the encoding mode probability of the CU with a size of 4×4 is output. The Softmax function can convert the feature map into a probability distribution form, thus intuitively presenting the possibilities of different modes, facilitating the selection of the optimal CU mode, and this method can fully integrate the feature information at different scales, further improving the accuracy and reliability of the prediction.

[0074] In the embodiment of the present invention, the processing of the basic features by the second subnet outputs mode probabilities with sizes such as 32×16, 32×8, 32×4, 16×8, 16×4, 8×4, etc. The structure of the second subnet is as Figure 4As shown. The processing process includes: using consecutive convolutional layers conv11 - conv15 to downsample the basic features to obtain the fifth, sixth, seventh, eighth, and ninth local features; among them, the number of channels in the convolutional layers doubles step by step; using convolutional layer conv31 to downsample the sixth local feature to obtain the tenth local feature; after processing the tenth local feature through convolutional layer conv14 and the Softmax function, output the coding mode probability of the CU with a size of 16×8 at the back; through consecutive transposed convolutional layers deconv5 - deconv8 to upsample the ninth local feature layer by layer to obtain the fifth, sixth, seventh, and eighth global features; after processing the ninth local feature through convolutional layer conv16 and the Softmax function, output the coding mode probability of the CU with a size of 32×16 at the back; concatenate the eighth local feature and the fifth global feature, and after processing the concatenated feature through convolutional layer conv17 and the Softmax function, output the coding mode probability of the CU with a size of 32×8 at the back; concatenate the seventh local feature and the sixth global feature, and after processing the concatenated feature through convolutional layer conv18 and the Softmax function, output the coding mode probability of the CU with a size of 32×4 at the back; concatenate the sixth local feature and the seventh global feature, and after processing the concatenated feature through convolutional layer conv19 and the Softmax function, output the coding mode probability of the CU with a size of 16×4 at the back; concatenate the fifth local feature and the eighth global feature, and after processing the concatenated feature through convolutional layer conv20 and the Softmax function, output the coding mode probability of the CU with a size of 8×4 at the back.

[0075] In the embodiment of the present invention, the third subnet processes the basic features and outputs the mode probabilities with sizes such as 16×32, 8×32, 4×32, 8×16, 4×16, 4×8, etc. The structure of the third subnet is as Figure 5As shown in the figure. The processing process includes: using consecutive convolutional layers conv21 - conv25 to downsample the basic features to obtain the tenth, eleventh, twelfth, thirteenth, and fourteenth local features; among them, the number of channels in the convolutional layers doubles step by step; using convolutional layer conv32 to downsample the eleventh local feature to obtain the fifteenth local feature; after processing the fifteenth local feature through convolutional layer conv12 and the Softmax function, output the coding mode probability of the CU with a size of 8×16; upsample the fourteenth local feature layer by layer through consecutive transposed convolutional layers deconv9 - deconv12 to obtain the ninth, tenth, eleventh, and twelfth global features; after processing the fourteenth local feature through convolutional layer conv26 and the Softmax function, output the coding mode probability of the CU with a size of 16×32; concatenate the thirteenth local feature and the ninth global feature, and after processing the concatenated feature through convolutional layer conv27 and the Softmax function, output the coding mode probability of the CU with a size of 8×32; concatenate the twelfth local feature and the tenth global feature, and after processing the concatenated feature through convolutional layer conv28 and the Softmax function, output the coding mode probability of the CU with a size of 4×32; concatenate the eleventh local feature and the eleventh global feature, and after processing the concatenated feature through convolutional layer conv29 and the Softmax function, output the coding mode probability of the CU with a size of 4×16; concatenate the tenth local feature and the twelfth global feature, and after processing the concatenated feature through convolutional layer conv30 and the Softmax function, output the coding mode probability of the CU with a size of 4×8.

[0076] In the embodiments of the present invention, the working principles of the second subnet and the third subnet are similar to that of the first subnet. They both process the input data through multiple convolution and downsampling operations, and obtain the final prediction result by means of feature concatenation and combination. At the end of each subnet, the final output result is processed through the Softmax layer. The Softmax layer converts the output data into a probability distribution, thereby realizing the probability prediction and selection of the optimal CU mode, and ensuring that the entire network architecture can efficiently and accurately determine the most suitable coding mode for CUs of different sizes and directions.

[0077] In the embodiments of the present invention, to ensure that there is no data overlap between the training set and the test set, a data set is constructed by statistically analyzing the same video sequences. After screening and processing, a data set containing 300,000 samples is finally generated, and the size of each sample is 16×16. After the data set is built, it enters the data set division link. The data set is divided into two parts, namely the training set and the validation set. By means of the random sampling method, two-thirds of the total number of all samples are extracted to form the training set; the remaining one-third is used as the validation set. After the division, the model performance can be evaluated in real time during training and targeted optimization can be carried out. The model training is carried out based on the PyTorch deep learning framework. To accelerate the training speed, a GeForce GTX3090 GPU is used to let the model perform 100 rounds of iterative training, and the Adam optimizer is selected. Before training, we set the initial learning rate to 0.001 and the batch size to 1024. In this way, the training efficiency is significantly improved and the model can converge faster.

[0078] In the embodiments of the present invention, in order to train the network to accurately predict the modes of 17 different-sized CUs, a standard loss function is adopted as the optimization objective. In the VVC SCC coding mode prediction, using the standard loss function can ensure the stable training of the model, improve the prediction performance, and enhance the generalization ability. The standard loss function is applicable to multi-class classification problems, can effectively measure the gap between the predicted probability and the true distribution, and avoid the model overfitting to a single mode. In addition, the standard loss function has been widely verified and has stable gradient update characteristics, making the optimization process more efficient and avoiding the problems of gradient disappearance or explosion. In contrast, a custom loss function may bring problems such as unstable gradients or too high computational complexity, increasing the difficulty of model tuning. Due to the certain uncertainty in the selection of coding modes, adopting the standard loss function can better adapt to probability prediction and improve the robustness of the model. At the same time, the standard loss function has been highly optimized in mainstream deep learning frameworks, enabling efficient calculation and accelerating training. Therefore, using the standard loss function is a key strategy to ensure the stability and optimization effect of the coding mode prediction network. The loss function in the training process of the VCC coding mode selection model is expressed as:

[0079]

[0080] Among them, L i is the loss value of the i-th sample, y i is the label of the i-th sample, p i is the probability that the i-th sample is a positive class, and N is the total number of samples.

[0081] In the embodiments of the present invention, through the training process, a VCC coding mode selection model is obtained. Using the VCC coding mode selection model, the coding mode probability can be determined, including P Intra 、PPLT , P IBC-Serch and P IBC-Merge . In an ideal situation, the network model can accurately predict all patterns, thereby avoiding a large amount of redundant search during the RDO process, and thus reducing the encoding complexity. However, due to the complexity of the encoder and the accuracy of the model, the performance of the encoder will decline. In order to reduce the impact of incorrect prediction on the encoder efficiency, the embodiments of the present invention select the encoding mode through probability values to achieve a balance between encoding efficiency and complexity.

[0082] Before the encoder in the embodiments of the present invention performs pattern checking, the patterns can be re-ordered to change the order of checking. In this way, the encoder can preferentially check those patterns that are most likely to meet the requirements, improving the encoding efficiency. After completing the check of the current candidate pattern, the encoder will obtain the encoding-related information of this pattern. Based on this information, the encoder can determine whether it is necessary to stop checking the remaining patterns according to the actual performance of the pattern. During this process, a decision tree is used to predict whether the pattern selection is terminated in advance. Decision trees are widely regarded as effective classifiers, known for their powerful classification performance and low complexity. When constructing a decision tree model, it is crucial to select features. Good features can improve the model accuracy. By screening features that are strongly correlated with the target variable, capturing key information, filtering noise, enabling the decision tree to better learn the data pattern, and also preventing overfitting, reducing the data dimension, reducing the interference caused by irrelevant or redundant features, and enhancing the generalization ability and stability of the model. In terms of feature selection, the RDcost of the currently checked pattern, the number of bits occupied by the encoding block, the variance of the absolute value of the residual, the width of the CU and the height of the CU, the quantization parameter QP, and the pattern information of its parent CU are selected.

[0083] In the embodiment of the present invention, a decision tree model is adopted to determine whether to terminate the inspection of the coding mode in advance. The performance of the decision tree model is closely related to the diversity and correlation of the training data set. Therefore, video sequences consistent with the statistical experiment are selected as the extraction source of the training samples. The specific operation process is as follows: use the VTM encoder deployed with the CNN classifier to encode these video sequences, then calculate the intermediate coding data and texture feature information of the CU, and obtain the true labels according to the coding results generated by the standard encoder. After data balancing processing, 200,000 samples are generated for the termination models of different modes. Normally, the larger the depth setting of the decision tree, the higher the prediction accuracy of the model usually is. However, because an overly deep decision tree is extremely likely to capture the noise and outliers in the data, which in turn causes the generalization performance of the model to decrease and leads to overfitting. Therefore, when training the decision tree, the maximum depth of the leaf nodes is limited to 6. In addition, to ensure that each leaf node has a certain number of samples to improve the generalization ability and stability of the model, the minimum number of samples in the leaf node is set to one-thousandth of the total number of samples. For the decision tree used for early termination of mode selection, to ensure the feature quality, for the decision tree models constructed for different modes, the top 6 features with the highest contribution are selected. After completing the model construction and derivation, pruning operations need to be performed on each redundant node in the decision tree model to achieve a more ideal balance between the coding complexity and the coding efficiency.

[0084] In the embodiment of the present invention, to achieve the goal of improving the coding speed and maintaining the coding efficiency, it is necessary to screen out those coding modes with lower probabilities. The sorted coding mode probabilities can be obtained through the VCC coding mode selection model. If the coding mode probability is greater than the threshold, the process is terminated in advance; if it is less than the threshold, it is skipped. Therefore, in the process of pursuing the balance between the coding efficiency and the complexity, accurately determining the optimal threshold is of great significance. In view of this, a series of commonly used T values in related research are selected for testing, including 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, and 0.9. By deeply analyzing the effects generated by these different threshold values, the coding efficiency and loss under different T values are obtained as Figure 6 shown. As the parameter T gradually increases, the BDBR shows a continuous decreasing trend; conversely, when T gradually decreases, the BDBR will continuously increase. It can be obtained that a lower T value will cause more coding modes to be wrongly screened out. Because a large number of modes are misjudged and excluded, this not only causes the BDBR to increase significantly, but also makes the coding speed achieve a more remarkable improvement. Therefore, there is an obvious trade-off relationship between the BDBR and the improvement of the coding speed. When the IBC mode threshold T Intra is 0.8, the PLT mode threshold T PLT is 0.6, the IBC Serch mode threshold TIBC-serch is 0.7 and the IBC Merge mode threshold T IBC-Merg When it is 0.7, the decline rate of BDBR tends to be stable and gradually slows down.

[0085] In the embodiments of the present invention, the Intra mode encoding is very complex. The encoder needs to check various prediction modes and techniques respectively to find the mode with the minimum RD cost as the final prediction mode. Although VVC has improved the intra-frame encoding efficiency by introducing technological innovations, its complexity has increased exponentially. Therefore, not only the mode needs to be optimized, but also the transformation in the Intra mode needs to be optimized.

[0086] The embodiments of the present invention introduce the MTS mechanism in the Intra mode. This mechanism significantly enhances the encoding flexibility and adaptability, enabling the encoder to more effectively process the complex structural information in screen content. The traditional DCT-II, as the basic transformation in VVC, is still retained. However, to further improve the encoding performance, VVC additionally introduces two new transformation matrices, DCT-VIII and DST-VII. The introduction of these transformations is mainly to better adapt to the complex texture structures and directional features in screen content, such as different types of visual content like text, graphic edges, and gradient backgrounds, making the transformed residual signal sparser and thus improving the encoding efficiency. In addition, to more efficiently process the unique flat areas and simple textures in screen content, VVC SCC introduces the BDPCM technology. BDPCM directly performs differential encoding on pixel values without transformation, thus avoiding the computational overhead in the transformation process. This technology is particularly suitable for areas with obvious horizontal or vertical direction structures, such as regular content like tables, text, and user interface elements, enabling it to effectively reduce the bit rate while ensuring the encoding accuracy.

[0087] In order to optimize the selection of transform candidates and reduce unnecessary computational overhead, the embodiments of the present invention conduct a systematic statistical analysis on the distribution differences of different transform matrices, and count the transform types selected by different-sized CUs in the optimal coding mode. The transform mode distribution is shown in Table 2. It can be seen that the optimal transform selection for different CU sizes shows obvious regularity, which is highly consistent with the selection trend of coding modes. Specifically, most CUs only need to select between BDPCM and DCT-II, and as the CU size decreases, the usage ratio of BDPCM increases significantly. The main reason for this phenomenon is that BDPCM only makes predictions in the horizontal and vertical directions, and its coding efficiency depends on the local correlation of pixels within the CU. Since small-sized CUs usually have relatively simple texture structures or regular directional characteristics, BDPCM can provide better compression effects with lower computational complexity, significantly enhancing its applicability to small-sized CUs. On the other hand, larger-sized CUs often contain more complex texture details, and the energy distribution of their residual signals in the transform domain is relatively dispersed. Therefore, DCT-II still needs to be used for transformation to fully remove redundant information and improve the concentration of spectral energy. In addition, in some special scenarios, the application ratios of DCT-VIII and DST-VII transforms also increase, indicating that high-order transform modes can still provide additional coding gains under certain specific contents.

[0088] Table 2 Transform Distribution of Coding Units of Different Sizes

[0089]

[0090]

[0091] Through the analysis of the transform distribution of coding units of different sizes, the embodiments of the present invention conclude that the optimal transform selection for different CU sizes has significant regularity. In particular, the applicability of BDPCM to small-sized CUs is significantly enhanced, while DCT-II remains the mainstream choice in most cases. This finding provides an important theoretical basis for optimizing the transform decision-making strategy in the VVC SCC coding framework, enabling the encoder to adaptively adjust the transform mode according to the size, content characteristics, and directional information of the CU in practical applications, thereby further improving the coding efficiency and reducing the computational overhead.

[0092] In the embodiment of the present invention, the same network as the VCC coding mode selection model is used to obtain the transformed mode probabilities. In the coding mode of VVC, there are extremely significant differences in the coding complexity between different steps. At the same time, there are also huge differences in the transformation distributions of coding units of different sizes. Given the QTMT partition structure adopted by VVC, this complex structure will give rise to more than a dozen CU size types. If a separate classification model is trained for each CU size, undoubtedly many difficult problems will be faced. On the one hand, the training process will be extremely cumbersome, requiring a large amount of time and computing resources to adjust the model parameters and optimize the model structure for different-sized CUs one by one. On the other hand, from the perspective of the operation of the encoder, multiple separate classification models are not conducive to the encoder for parallel prediction. Parallel prediction requires a high degree of consistency and coordination in the model structure and data processing flow, while numerous independent classification models will disrupt this coordination, resulting in low efficiency of the encoder in parallel processing data and making it difficult to fully utilize the parallel computing power of the hardware. However, in order to meet the actual needs of VVC coding and simplify the calculation process, the network outputs the probabilities of 3 modes in the transformation mode. In this way, it is possible to provide support for the calculation of the transformation mode probabilities in VVC coding in a relatively simple and effective way, and to balance the relationship between coding complexity and computing efficiency to a certain extent.

[0093] Through the above process, all training models can be obtained in the embodiment of the present invention. Using these models, the transformed mode probabilities: P BDPCM , P DCT and P Other can be determined. Before the encoder starts the mode check, the modes can be sorted to change the order of the check process. By this means, the encoder can give priority to checking those most potential modes, thus effectively improving the overall efficiency of coding. When the encoder completes the check of the current candidate mode, it will obtain the detailed coding information of this mode. Based on this information, the encoder can flexibly judge whether it is necessary to stop the check of the remaining modes according to the performance of the mode in actual applications. In this process, the decision tree is still used as an important means to predict the early termination of mode selection. As a classification tool widely used and recognized as effective in many fields, the decision tree performs excellently in various data analysis and processing tasks with its powerful classification efficiency and low complexity. When building the decision tree model, the selection of features is particularly crucial. In terms of feature selection, the size of the transformed CU, the coding bits, the RD cost, the variance of the residual coefficients, the quantization parameter QP, and the transformation type of its parent CU are selected.

[0094] In the embodiments of the present invention, a series of commonly used T values in related research are selected for testing, covering 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, and 0.9. Through a rigorous experimental process, the effects produced by these different threshold values are analyzed, and the loss of coding efficiency corresponding to different T values is obtained. The loss situation is as Figure 7 shown. As the T value increases, the corresponding BDBR gradually decreases. A smaller T value excludes more coding modes, resulting in a greater increase in BDBR but a faster coding speed. There is a delicate balance between BDBR and speed improvement. When the BDBPCM mode threshold T BDPCM is 0.7, the DCT-II mode threshold T DCT-II is 0.6, and the other high-order transform mode threshold T Other is 0.6, the decrease in BDBR is stable and slows down. To achieve the goal of improving coding speed and maintaining coding efficiency, it is necessary to screen out those coding modes with lower probabilities. The sorted coding mode probabilities can be obtained through the network model. If the coding mode probability is greater than the threshold, it is terminated in advance; if it is less than the threshold, it is skipped in advance.

[0095] The embodiments of the present invention also conduct performance verification, integrating this method into VTM-17.0 for testing. In addition, a comparative analysis is carried out for another three existing methods - MLFM, OLFM, and FastSCC. The performance comparison is shown in Table 3. Considering the practicality of the method, the test results cover the inference time of the used model. It can be seen that the improvement ratios of this method compared with MLFM, OLFM, and FastSCC in terms of average coding speed are 49.33%, 22.32%, 16.51%, and 26.98% respectively. At the same time, the average BDBR values of this method compared with MLFM, OLFM, and FastSCC are 1.81%, 1.61%, 0.72%, and 3.22% in turn. Compared with MLFM and FastSCC, from the comprehensive dimension of coding speed improvement and coding efficiency loss, this method has excellent performance. Compared with the OLFM method, the coding speed of this method is increased by nearly 32.82%, and at the same time, the BDBR increases by 1.09%. Although the coding efficiency loss is slightly higher than that of OLFM, due to the significant improvement in coding speed, this slight increase in coding efficiency loss can be ignored. The superiority of this method over other methods in performance is mainly attributed to two key factors. First, there are significant differences in the mode distributions of CUs of different sizes. By designing convolutional layers for each CU according to the size to predict candidate modes, higher prediction accuracy can be achieved. Second, with the help of GPU parallel processing, the trained model based on all CUs is used to infer candidate modes, and the corresponding inference time is almost negligible.

[0096] Table 3 Performance comparison of different methods

[0097]

[0098]

[0099] In summary, the present invention first statistically analyzes the distribution of the VVC SCC intra-coding modes, and then proposes a fast VCC coding mode selection method based on CNN for the statistical results, including network structure design, model training, early termination of mode selection, decision tree implementation, and threshold selection. Finally, the performance of this method is evaluated. This method can save 49.33% of the coding time while increasing the BDBR by 1.81%, indicating that the performance loss is negligible while significantly reducing the coding time.

[0100] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to the embodiments of the present invention without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A VCC coding mode fast selection method based on CNN, characterized in that: The following steps are involved: S1: Obtain the brightness component of CTU and preprocess it, input the preprocessed CTU into the trained VCC coding mode selection model, and obtain the coding mode probabilities of CUs of different sizes; The coding modes include PLT mode, IBC mode and Intra mode; S2: Sort the coding modes in descending order according to their probability, and the encoder checks the coding mode with the highest probability first; S3: according to the preset coding mode threshold, a decision tree model is used to determine whether to terminate the coding mode check in advance. If the check is terminated, the optimal coding mode is output; otherwise, the coding mode is changed in the sorting order and the check is continued until the selection of the coding mode is completed; S4: Determine whether the coding mode selected by the CU is the Intra mode. If so, execute step S5. Otherwise, complete the CU coding according to the selected coding mode. S5: Input the brightness component of the CTU into the transform mode selection model for processing, and output the transform mode probabilities of CUs of different sizes; the transform modes include BDPCM, DCT-II and other high-order transforms; S6: sorting the transform modes in descending order according to the transform mode probabilities, and the encoder first checks the transform mode with the highest probability; S7: according to the preset transformation mode threshold, a decision tree model is used to determine whether to terminate the transformation mode check in advance. If the check is terminated, the optimal transformation mode is output; otherwise, the transformation mode is changed in the sorting order and the check is continued until the transformation mode selection is completed; S8: Select the CU in Intra mode and complete CU encoding according to the Intra mode and the selected transform mode.

2. The method for quickly selecting a VCC coding mode based on CNN according to claim 1, characterized in that: The VCC coding mode selection model includes a main network and three sub-networks, wherein the three sub-networks are respectively a first sub-network, a second sub-network and a third sub-network; the main network is used to extract basic features from the brightness component of the CTU, and the sub-network outputs coding mode probabilities of CUs of different sizes.

3. The method for quickly selecting a VCC coding mode based on CNN according to claim 2, characterized in that: The main network processes the brightness component of CTU in the following ways: The basic features of the brightness component of CTU are extracted through the convolution layer conv1 to obtain the basic features; The first subnet is used to process the basic features and output the coding mode probabilities of CUs of sizes 64×64, 32×32, 16×16, 8×8, and 4×4; The second subnet is used to process the basic features and output the coding mode probabilities of CUs of size 32×16, 32×8, 32×4, 16×8, 16×4, and 8×4. The third subnet is used to process the basic features and output the coding mode probabilities of CUs of sizes 16×32, 8×32, 4×32, 8×16, 4×16, and 4×8.

4. The method for quickly selecting a VCC coding mode based on CNN according to claim 2, characterized in that: The first subnet processes the basic features as follows: The basic features are downsampled using continuous convolutional layers conv2 to conv5, and the four convolutional layers output the first, second, third, and fourth local features respectively; the number of channels in the convolutional layer doubles step by step; The fourth local feature is upsampled layer by layer through consecutive deconvolution layers deconv1~deconv4, and the four deconvolution layers output the first, second, third and fourth global features respectively; After the fourth local feature is processed by the convolution layer conv6 and the Softmax function, the coding mode probability of the CU with a size of 64×64 is output; Concatenate the third local feature and the first global feature, process the concatenated feature through the convolution layer conv7 and the Softmax function, and output the coding mode probability of the CU of size 32×32; Concatenate the second local feature and the second global feature, process the concatenated feature through the convolution layer conv8 and the Softmax function, and output the coding mode probability of the CU of size 16×16; Concatenate the first local feature and the third global feature, process the concatenated feature through the convolution layer conv9 and the Softmax function, and output the coding mode probability of the CU of size 8×8; The basic features and the fourth global features are concatenated, and the concatenated features are processed by the convolution layer conv10 and the Softmax function to output the coding mode probability of the CU of size 4×4.

5. The method for quickly selecting a VCC coding mode based on CNN according to claim 2, characterized in that: The second subnet processes the basic features as follows: The basic features are downsampled using continuous convolutional layers conv11-conv15, and the four convolutional layers output the fifth, sixth, seventh, eighth and ninth local features respectively; the number of channels in the convolutional layer doubles step by step; The convolution layer conv31 is used to downsample the sixth local feature to obtain the tenth local feature; After the tenth local feature is processed by the convolution layer conv14 and the Softmax function, the coding mode probability of the CU with a size of 16×8 is output; The ninth local feature is upsampled layer by layer through consecutive deconvolution layers deconv5~deconv8, and the four deconvolution layers output the fifth, sixth, seventh and eighth global features respectively; After the ninth local feature is processed by the convolution layer conv16 and the Softmax function, the coding mode probability of the CU with a size of 32×16 is output; Concatenate the eighth local feature and the fifth global feature, process the concatenated feature through the convolution layer conv17 and the Softmax function, and output the coding mode probability of the CU of size 32×8; Concatenate the seventh local feature and the sixth global feature, process the concatenated feature through the convolution layer conv18 and the Softmax function, and output the coding mode probability of the CU of size 32×4; Concatenate the sixth local feature and the seventh global feature, process the concatenated feature through the convolution layer conv19 and the Softmax function, and output the coding mode probability of the CU of size 16×4; The fifth local feature and the eighth global feature are concatenated, and the concatenated features are processed by the convolution layer conv20 and the Softmax function to output the coding mode probability of the CU of size 8×4.

6. The method for quickly selecting a VCC coding mode based on CNN according to claim 2, characterized in that: The third subnet processes the basic features as follows: The basic features are downsampled using continuous convolutional layers conv21-conv25, and the four convolutional layers output the tenth, eleventh, twelfth, thirteenth and fourteenth local features respectively; the number of channels in the convolutional layer doubles step by step; The convolution layer conv32 is used to downsample the eleventh local feature to obtain the fifteenth local feature; After the fifteenth local feature is processed by the convolution layer conv12 and the Softmax function, the coding mode probability of the CU with a size of 8×16 is output; The fourteenth local feature is upsampled layer by layer through consecutive deconvolution layers deconv9~deconv12, and the four deconvolution layers output the ninth, tenth, eleventh and twelfth global features respectively; After the fourteenth local feature is processed by the convolution layer conv26 and the Softmax function, the coding mode probability of the CU with a size of 16×32 is output; Concatenate the thirteenth local feature and the ninth global feature, process the concatenated feature through the convolution layer conv27 and the Softmax function, and output the coding mode probability of the CU of size 8×32; Concatenate the twelfth local feature and the tenth global feature, process the concatenated feature through the convolution layer conv28 and the Softmax function, and output the coding mode probability of the CU of size 4×32; Concatenate the eleventh local feature and the eleventh global feature, process the concatenated feature through the convolution layer conv29 and the Softmax function, and output the coding mode probability of the CU of size 4×16; The tenth local feature and the twelfth global feature are concatenated, and the concatenated features are processed by the convolution layer conv30 and the Softmax function to output the coding mode probability of the CU of size 4×8.

7. The method for quickly selecting a VCC coding mode based on CNN according to claim 1, characterized in that: The preset coding mode thresholds include the IBC mode threshold T Intra is 0.8, PLT mode threshold T PLT is 0.6, IBC Serch mode threshold T IBC-Serch is 0.7 and the IBC Merge mode threshold T IBC-Merge is 0.

7.

8. The method for quickly selecting a VCC coding mode based on CNN according to claim 1, characterized in that: The network structure of the VCC coding mode selection model is the same as that of the transformation mode selection model.

9. The method for quickly selecting a VCC coding mode based on CNN according to claim 1, characterized in that: The preset conversion mode threshold includes the BDPCM mode threshold T BDPCM is 0.7, DCT-II mode threshold T DCT- The threshold T for other high-order transformation modes is 0.6 Other is 0.6.