A fast video encoding method based on deep neural network
By adopting a fast video encoding method based on deep neural network in HEVC video encoding, using convolution and self-attention mechanisms to divide CU blocks and select PU modes, the problems of complex and high encoding time of HEVC video encoding are solved, and efficient video encoding is achieved.
Patent Information
- Application Number
- CN202111599851.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-12-24
AI Technical Summary
The existing HEVC video encoding technology is complex and has a long encoding time, making it difficult to meet the requirements of real-time applications.
A fast video encoding method based on deep neural network is adopted, and a deep neural network with convolution and self-attention mechanism is used to perform CU block division and PU mode selection, reducing coding complexity and time.
While ensuring encoding performance, the video encoding time is significantly reduced, the average time efficiency can reach 70%, and the code rate increases by only about 2%.
Smart Images

Figure CN114286093B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of high efficiency video coding (HEVC), and in particular relates to a low-complexity fast HEVC intra-frame video coding method. Background Art
[0002] High Efficiency Video Coding (HEVC), the latest video coding standard developed by the Joint Development Team for Video Coding (JCT-VC) in 2012, significantly improves coding performance. Compared with the previous video coding standard H.264 / AVC, HEVC uses several carefully designed methods to save about 50% of the bit rate at the same video compression quality. Specifically, for intra-frame coding, the coding unit (CU) based on the quadtree structure is recursively divided into 64×64 to 8×8 blocks; in addition, for the prediction unit (PU), up to 35 intra-frame prediction modes are allowed, including DC, Planar, and 32 angular prediction modes. Both technologies are conducive to improving coding performance, but at the cost of greatly increasing the coding complexity, which is difficult to meet the requirements of real-time applications. Therefore, it is necessary to study fast video coding methods.
[0003] So far, many HEVC fast intra-frame coding algorithms have been proposed, which can be roughly divided into two categories: fast partitioning of CU blocks and fast selection of intra-frame modes. Since the partitioning of CU is a top-down recursive partitioning through rate-distortion optimization, and its block size is flexible, its coding complexity is very high. Many fast coding methods try to predict the CU partitioning mode in advance to avoid lengthy recursive RDO search. In terms of intra-frame mode selection, the current HEVC encoder uses a three-step algorithm to accelerate the intra-frame mode determination process. The optimization of intra-frame mode selection is currently mainly focused on simplifying the RMD process or RDO calculation. The optimization algorithm uses traditional heuristic methods to manually extract the texture features of CU blocks or use the relevant features between adjacent CUs. Methods in the field of machine learning such as decision trees, support vector machines, and Bayesian decision-making have also been applied to CU depth decision-making, all of which have achieved certain optimization effects; and convolutional neural networks, due to their excellent local feature extraction capabilities, have outstanding image classification performance in the field of computer vision. In recent years, many neural network structures have been designed to automatically extract texture and object features in coding blocks. However, due to the prediction accuracy of the neural network, a high accuracy means fewer cases of incorrectly predicting CUs, which will not lead to excessive unnecessary bit rate increases. However, the network model will be too complex, which will in turn bring additional time overhead. Therefore, the design of the network structure needs to strike a balance between complexity and prediction accuracy.
[0004] In 2017, the Google team proposed a Transformer neural network model based on the self-attention mechanism. This model uses the self-attention mechanism instead of the sequential structure of RNN, which allows the model to be trained in parallel and has global information. This is effective when processing sequences, and it quickly became the model of choice in the field of natural language processing (NLP), and then gradually expanded to the field of computer vision. In 2021, the Google team published an article titled "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale" at the International Conference on Learning Representations (ICLR). The Vision Transformer (ViT) network model described in the article introduced the Transformer structure into the classification task of computer vision for the first time, and achieved better performance than the CNN model. The ViT model directly migrates the encoder module in the Transformer and adds a classification vector to output the classification results. Compared with CNN (convolutional neural network), the ViT network can integrate the information of the entire image, even the lowest level information, and then bring better classification results, which is beyond the reach of CNN. The huge potential shown by ViT has led to an explosive growth in research on Transformer and attention mechanisms. Summary of the invention
[0005] The purpose of the present invention is to address the shortcomings of existing HEVC video encoding, such as high complexity and long overall encoding time, and propose a fast video encoding method based on deep neural network, which utilizes the powerful feature extraction and learning capabilities of deep neural network based on convolution and self-attention mechanism to achieve fast encoding while ensuring encoding performance.
[0006] The present invention proposes a fast video encoding method based on a deep neural network, and its specific implementation includes a CU division module based on a deep neural network and a PU mode selection module based on neighborhood correlation.
[0007] The CU partitioning module based on deep neural network first uses the neural network to predict the partitioning results of each CU block from top to bottom, and then uses the threshold to optimize the prediction results to reduce the occurrence of erroneous predictions. Finally, the encoder determines whether to terminate the partitioning of the current CU block in advance according to the partitioning results during encoding.
[0008] The PU mode selection module based on neighborhood correlation first uses a neural network to predict the position of the best mode in the candidate mode list after RMD rough selection, and avoids some redundant modes from entering the RDO calculation by discarding the subsequent modes. At the same time, it optimizes the most likely mode (MPM) and reduces the number of MPMs in the candidate list, thereby reducing the amount of mode calculation and reducing the complexity of PU mode selection.
[0009] The present invention utilizes the correlation between the CU depth and the PU prediction mode in the neighborhood in the video image, reduces the complexity of CU recursive partitioning, simplifies the selection process of the intra-frame prediction mode, and effectively improves the coding efficiency of HEVC.
[0010] The technical solution adopted by the present invention to solve the technical problem is as follows
[0011] (I) CU division module based on deep neural network
[0012] Step (I), construct a dataset for network model training in HEVC intra-frame mode: The dataset comes from YUV video sequences of various resolutions including: CIF (352×288), 480p (832×480), 720p (1280×720), 1080p (1920×1080), WQXGA (2560×1600). The samples of the dataset consist of the luminance component of the CU block and the corresponding sample label, where the sample label is obtained by encoding the luminance component using the HEVC reference software HM16.9. The dataset includes a training set, a validation set, and a test set, and each dataset is divided into four subsets according to four QPs (22, 27, 32, 37).
[0013] Step (II) constructs a deep neural network with three CU blocks of 64×64, 32×32, and 16×16 to form a hierarchical convolutional network (HCT) structure. The hierarchical convolutional network HCT is a combination of ViT and CNN. The HCT is trained through the corresponding training set, and the HCT model is determined and saved by the validation set. Finally, the test set determines the generalization ability of the HCT model. The objective function of HCT model training is the cross entropy loss function (CrossEntropyLoss):
[0014]
[0015] Where output is the output vector of the HCT model, target is the sample label value, and L is the length of the output vector.
[0016] Step (III), the hierarchical convolutional network HCT consists of a convolutional module, a Transformer encoder module, a sequence pooling layer, and a fully connected layer. First, the brightness component of the CU block is sent to the hierarchical convolutional network HCT, and the convolution module outputs a feature map with local feature information. The convolution module contains a convolution layer and a maximum pooling layer (Maxpool). Each layer is activated by a linear rectifier function (ReLU) to improve the nonlinearity of the model. Then the feature map is flattened into one dimension and flipped. Assume that the input image x∈R C×H×W , where C represents the number of input images, H is the height of the image, and W is the width of the image. The output after the convolution module is as follows:
[0017] x 0 =Transpose(Flatten(MaxPool(Conv2d(x)))) (2)
[0018] Next, the feature data x 0 The sum of the position vector and the position vector is fed into the Transformer encoder module to extract global information. The encoder module has 7 layers, each of which consists of a multi-headed self-attention layer (MSL) and a feed-forward convolution layer (FCL). Both sub-layers are preceded by a layer normalization (LN) operation to improve the robustness and generalization ability of the model. 0 After passing through the multi-head self-attention layer, the output data is the same as x 0 Add to get new feature data x 1 , x 1 After the feedforward convolution layer, its output value is the same as x 1 Add to get the characteristic data x 2 , the formula is as follows:
[0019] x 1 =x 0 +MSL(LN(x 0 )) (3)
[0020] x 2 =x 1 +FCL(LN(x 1 )) (4)
[0021] Finally, the classification vector is obtained through the sequence pooling layer. The sequence pooling uses the mapping transformation T: R b×n×d→R b×d , b represents the batch size (batch_size), n represents the number of feature data, and d represents the size of each feature data. This operation converts the entire Transformer encoder module output feature data x 2 It is directly transformed into a classification vector, which contains relevant information about each part of the input image, and is used to replace the additional classification vector added in ViT. Finally, the classification vector is output as a binary classification result after the full connection layer and Softmax, and the final prediction value is the subscript of the maximum output value (0 or 1).
[0022] Step (IV): The HCT model is trained using the stochastic gradient descent (SGD) method, and a total of 12 HCT models with the highest accuracy for the three types of CU blocks under four QPs are saved. The trained HCT model uses an early termination mechanism to predict the partitioning results of 64×64, 32×32, and 16×16 blocks from top to bottom. The prediction results of the model are divided into two categories: 0 represents no partitioning and 1 represents partitioning. When the prediction result of a certain type of block is 0, the quadtree partitioning will not be continued during encoding. In this way, some redundant block partitioning operations can be avoided by terminating the partitioning early, thereby reducing the encoding time complexity.
[0023] In order to reduce the additional coding performance loss caused by the model's incorrect prediction, this module uses threshold optimization to improve coding performance. The present invention uses the similarity (SD) between the two-category vectors to compare with the threshold λ. When SD is less than the threshold λ, we can use the original coding method to check the CU block, which reduces the misjudgment of the CU division result and improves the coding performance, achieving a trade-off between coding performance and complexity. The calculation formula of the similarity SD is as follows:
[0024]
[0025] output i is the binary classification output vector of the i*i size block. Here we divide the threshold λ into three categories according to the block size, and the size ratio is 4:2:1.
[0026] (II) PU mode selection module based on neighborhood correlation:
[0027] Step (1), obtain the sample label value label∈[0, 1, 2] of each PU block during intra-frame mode selection through HM coding, and the acquisition rules are as follows: for PU blocks of sizes 64×64, 32×32, and 16×16, the original length of the candidate list after RMD rough selection is 3. If the best mode of the PU block during mode selection is the first in the candidate list after RMD rough selection, label=0, and the length of the corresponding candidate list after RMD rough selection becomes 1; if the best mode of the PU block during mode selection is the second in the candidate list after RMD rough selection, label=1, and the length of the corresponding candidate list after RMD rough selection becomes 2; in other cases, label=2, and the length of the corresponding candidate list after RMD rough selection is 3. For 8×8 and 4×4 PU blocks, since the original length of their candidate lists is 8, we also divide it into three intervals to correspond to label=0, 1, and 2, respectively: if the best mode of the PU block after mode selection is located in the first or second position in the candidate list after RMD rough selection, label=0, and the corresponding length of the candidate list after RMD rough selection becomes 2; if the best mode of the PU block during mode selection is located in the third or fourth position in the candidate list after RMD rough selection, label=1, and the corresponding length of the candidate list after RMD rough selection becomes 4; in other cases, label=2, and the corresponding length of the candidate list after RMD rough selection is 8.
[0028] Step (2), the data set of the PU mode selection module also comes from the video sequence mentioned in the CU partition module based on the deep neural network, and adds 8×8 and 4×4 PU block data on the basis of the block partition data set. The model of PU blocks of 64×64, 32×32, and 16×16 is similar to the model of the block partition module, but the number of layers of the Transformer encoder module becomes 1. The model corresponding to the 8×8 and 4×4 PU blocks is simplified on the basis of the number of layers of the encoder module becoming 1, that is, no dimensionality reduction operation is performed and the maximum pooling layer is removed to reduce the complexity of the model. In addition, the model training uses the mean square error loss function (MSELoss) for regression training:
[0029]
[0030] Where output is the model output value vector with a length of 3, value is the true value vector obtained by comparing output with label, and N is the number of input images in each training.
[0031] The rules for obtaining the true value vector are as follows: Assuming output = [x, y, z], when label = 0, if the maximum value in the output appears at the subscript 0, then value = output; if the maximum value appears at the subscript 1, then value = [y, x, z]; if the maximum value appears at the subscript 2, then value = [z, y, x]. Similarly, when label = 1, if the maximum value in the output appears at the subscript 0, then value = [y, x, z]; if the maximum value appears at the subscript 1, then value = output; if the maximum value appears at the subscript 2, then value = [x, z, y]. When label = 2, if the maximum value in the output appears at the subscript 0, then value = [z, y, x]; if the maximum value appears at the subscript 1, then value = [x, z, y]; if the maximum value appears at the subscript 2, then value = output.
[0032] The beneficial effects of the present invention are as follows:
[0033] (1) The present invention adopts a batch prediction method for video sequences, and the prediction results of CU blocks or PU blocks can be obtained by running only once, which can significantly reduce the prediction time of the neural network. The present invention adopts a method of terminating CU block division in advance, and uses the prediction results to reduce the number of modes of PU blocks entering RDO calculation after RMD rough selection, and also optimizes the MPM mode to reduce the number of modes in the candidate list, thereby achieving fast encoding.
[0034] (2) Compared with the CNN structure in the prior art, the HCT model can not only automatically extract relevant local features of image blocks, but also has the ability to extract global information, which improves the prediction accuracy and generalization ability of the model. The parallel computing characteristics reduce the computational complexity of the model and greatly reduce the memory resources consumed. The data set used in the present invention is much smaller than the data set required by the CNN network in the prior art, which greatly reduces the training time of the model. While ensuring that the time complexity is greatly reduced, the bit rate does not increase much, and the practicality is stronger.
[0035] (3) The present invention simulates five video sequences with different resolutions, namely A (2560×1600), B (1920×1080), C (832×480), D (416×240), and E (1280×720). The experimental results show that the average time efficiency can reach about 70%, and the bit rate only increases by about 2%. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 HEVC original intra-frame coding flow chart
[0037] Figure 2 It is the overall algorithm diagram of the present invention;
[0038] Figure 3 This is a flow chart of the CU division module method of the present invention;
[0039] Figure 4 It is a schematic diagram of the data set of the present invention;
[0040] Figure 5 It is a schematic diagram of the HCT model of the present invention;
[0041] Figure 6 This is a flow chart of the PU mode selection module method of the present invention;
[0042] Figure 7 A schematic diagram of the correspondence between the data set label of the PU mode selection module of the present invention and the optimal mode position after RMD rough selection;
[0043] Figure 8 A schematic diagram of the corresponding relationship between the data set label and the true value vector during the PU mode selection module model training of the present invention; DETAILED DESCRIPTION
[0044] The original HEVC intra-frame coding process is as follows Figure 1 As shown in the figure, first, the 64×64 CU block must first be selected by the PU mode to calculate the rate-distortion cost, and then divided into 4 32×32 CU blocks through the quadtree. These CU blocks are selected by the PU mode to calculate the rate-distortion cost, and then divided down to 8×8 size. Finally, the rate-distortion cost of the CU block and its 4 sub-CU blocks is compared from bottom to top to finally obtain the division result. Most of the deep learning methods in the prior art use convolutional neural networks to automatically extract local features in CU or PU blocks for training to optimize the downward division operation of CU blocks or the rate-distortion calculation of PU mode selection, and the network model often requires a complex network structure or a large-scale data set to train in order to achieve good prediction accuracy and generalization ability.
[0045] The core improvements of the present invention may include: 1. Proposing a hierarchical convolutional network, namely HCT, which only needs to be trained with a small data set to achieve prediction results similar to or even better than the CNN model under a large data set, and quickly output the CU block division results through batch prediction to reduce the complexity of intra-frame modes. 2. The HCT structure is lightweight and used for PU mode selection, and the initial 3 (corresponding to 64×64, 32×32, 16×16 PU blocks) or 8 (corresponding to 8×8, 4×4 PU blocks) modes are divided into [1,2,3] or [2,4,8] for prediction, and the MPM is optimized to reduce the time complexity of intra-frame coding by reducing the number of modes entering the RDO calculation.
[0046] The present invention will be further described below in conjunction with the accompanying drawings and implementation examples.
[0047] A fast video encoding method based on deep neural network, the overall algorithm framework is as follows Figure 2 As shown, it is composed of a CU partitioning module based on a deep neural network and a PU mode selection module based on neighborhood correlation. When the CU block is intra-coded, the PU mode selection will first be used to calculate the rate-distortion cost. At this time, the PU mode selection module is first used for optimization, and the number of candidate modes calculated by the RDO is reduced by the prediction results of the lightweight HCT model; after the PU mode selection is completed, the encoder will make a CU depth decision to determine whether the CU block is divided. At this time, the CU partitioning module based on the neural network is optimized, and the prediction results are obtained from the HCT model to determine whether to terminate the division in advance, otherwise the division continues to the PU mode selection and CU partitioning decision of the sub-CU block. The simulation of the present invention adopts the HEVC official reference software HM16.9 for compression encoding. The test conditions refer to the general test conditions of JCT-VC (JCTVC-R1015), and the full intra-frame coding configuration file encoder_intra_main.cfg of the HM16.9 model is used. The flowchart of the CU block partitioning module method based on a deep neural network in the present invention is shown as follows. Figure 3 As shown, the specific steps are as follows:
[0048] Step (I): construct the database required for HCT model training. The relevant data sources are as follows: Figure 4 First, from Figure 4 In the video sequence shown, the brightness component of one frame of the image is extracted every 50 frames, and the binary labels corresponding to the three CU blocks of 64×64, 32×32, and 16×16 in the image are obtained by encoding. The label value is 0 (representing no division) or 1 (representing division). The final dataset sample size is 9,865,968, and it is divided into 12 sub-datasets according to its QP value and CU block size for network model training.
[0049] Step (II), constructing the HCT model for training the above database and CU partition prediction, such as Figure 5 As shown in the figure, the input of the HCT model is the brightness component of the CU block, and each sub-dataset corresponds to a network model. Each CU block sent to the network will pass through the convolution module, Transformer encoder module, sequence pooling layer and fully connected layer, and finally output a binary classification vector. When predicting, the index subscript corresponding to the maximum value in the vector will be selected as the block partition prediction result. The HCT model is a combination of CNN and Vision Transformer networks. The configuration and function description of each part are as follows:
[0050] 1) Convolution module. This module is used to extract local features of CU blocks and improve the prediction accuracy of the model. Its formula is as follows:
[0051] x 0 =Transpose(Flatten(MaxPool(Conv2d(x)))) (1)
[0052] Since the input image is a batch operation, a number of CU brightness components x of size [1, W, W] (W∈[64, 32, 16]) are selected for convolution operation, the convolution kernel size is 3×3, the step size is 2, and the padding is 1; 128 feature maps of size W / 2×W / 2 are obtained and then the maximum pooling is performed, and its dimension becomes [T, 128, W / 4, W / 4]. Both the convolution and maximum pooling layers are activated by the linear rectification function (ReLU) to improve the nonlinearity of the model, and then the feature map is flattened into one dimension, and the dimension becomes x after flipping. 0 =[T, W / 4×W / 4, 128]; Finally, the learnable position information (position vector) is added. This information comes from the ViT network structure, which allows the model to learn the correlation between different positions, thereby improving the prediction accuracy.
[0053] 2) Transformer encoder module. The encoder module proposed in the present invention has a total of 7 layers, each of which is composed of a multi-headed self-attention layer (Multi-headed Self-attention Layer, MSL) and a feed-forward convolution layer (Feed-forward Convolution Layer, FCL). This module can parallelize the feature map data and obtain the global information of the image. The multi-headed self-attention layer comes from the ViT network structure, and the feed-forward layer in the original ViT encoder module is actually composed of two fully connected layers combined to perform a simple dimensionality increase and reduction operation. Too many training parameters will lead to excessive calculation. The present invention changes it to two convolution layers with a convolution kernel of 1x1 to optimize the model parameters. The input and output of the module remain unchanged, both of which are [T, W / 4×W / 4, 128]. The relevant calculation formulas are shown below;
[0054] x 1 =x 0 +MSL(LN(x 0 )) (2)
[0055] x 2 =x 1 +FCL(LN(x 1 )) (3)
[0056] The feature data x output by the convolution module 0 After passing through the multi-head self-attention layer, the output data is the same as x 0 Add to get new feature data x 1 , x 1 After the feedforward convolution layer, its output value is the same as x 1 Add to get the characteristic data x 2 Layer Normalization (LN) operations are added before the multi-head self-attention layer and the feed-forward convolution layer to stabilize the model and regularize it.
[0057] 3) Sequence pooling layer. The data output after the Transformer encoder module contains local information and global information of the input image. Different from the additional classification vectors added in ViT, the present invention maps these data into classification vectors through sequence pooling operations. First, we reduce the dimension of the Transformer encoder module output data x = [T, W / 4×W / 4, 128] to [T, W / 4×W / 4, 1], and then normalize the data through Softmax, and then flip it to y = [T, 1, W / 4×W / 4]; multiply y and x to obtain a vector of dimension [T, 1, 128], and finally obtain a binary classification vector after the dimensionality reduction through the fully connected layer: [T, 128]=>[T, 2]. Each CU block outputs two classification values, representing the probability of division and non-division. The present invention uses the subscript (0 or 1) where the maximum value is located as the representation of non-division and division.
[0058] 4) Other layers. The multi-head self-attention layer and feed-forward convolution layer in the Transformer encoder module are added with dropout operation, with a random dropout probability of 10%, in order to prevent overfitting and thus improve the generalization ability of the network.
[0059] Step (III), with the database and network model, we can train the model. According to the three CU block sizes and four QPs, the block partitioning module of the present invention requires a total of 12 HCT models. The network model is built and trained under the Pytorch deep learning library, and the required loss function is the cross entropy loss function:
[0060]
[0061] Where output is the network model output vector, target is the sample label value, and L is the length of the output vector. The model parameters obtained after training the training set are then verified by the validation set for model accuracy and whether to save the model parameters. This is one iteration. The number of iterations in the present invention is unified as 100 times, and the batch size is 64. Finally, the network model with the highest accuracy in the test set is obtained for prediction.
[0062] Step (IV), after the model is trained, it can be used for encoding. The HEVC reference software used in the present invention is HM16.9. When the encoder starts encoding, it will first use the HCT network model to predict the three CU blocks of the to-be-encoded frames of the video sequence, and then use the obtained prediction results for the quadtree division judgment of the CU block in the intra-frame mode. If the current CU block prediction result is 0, the division is terminated in advance, otherwise the quadtree division judgment is continued downward.
[0063] In order to reduce the additional coding performance loss caused by the model's misprediction, this module also adds threshold optimization to improve coding performance. Since the output of the model is a binary classification vector, we convert it into a soft classification value between [0, 1] through the Softmax normalization operation. The values at the subscripts 0 and 1 represent the probability of the model predicting no division and division. When the distance between the two is larger, that is, the probability value of one side is getting larger and larger, and the probability value of the other side is getting smaller and smaller, the model will have a clearer and clearer prediction ability for the two categories, and it is not easy to make wrong predictions; on the contrary, the smaller it is, the closer the probability values of the two are, and finally they are equal. At this time, the model's prediction ability for the two categories will become more and more vague and uncertain, and it is difficult to judge accurately, and then it is easy to make wrong predictions. Therefore, the present invention uses the similarity (SimilarDegrees, SD) between the soft classification values to compare with the threshold λ. When SD is less than the threshold λ, we can use the original coding method to check the CU block, so that the situation of misjudgment of the CU division result can be reduced, thereby improving the coding performance, and achieving a trade-off between coding performance and complexity. The calculation formula of the similarity SD is as follows:
[0064]
[0065] output i is the binary classification output vector of the i*i size block, and T is the number of input images. It should be noted that different CU blocks have different contents, and the complexity reduction achieved is also different, and the larger the block, the more likely it is to be divided, so the threshold range should be smaller. Here we divide the constant λ into three categories according to the block size, and λ 64×64 =2λ 32×32 =4λ 16×16 ,Through experimental verification, this fixed value combination can achieve better ,encoding performance.,After the above optimization operations, the time complexity of video encoding ,can be effectively reduced without much bit rate increase.
[0066] The PU mode selection module based on neighborhood correlation has a method flow chart as shown below: Figure 6 This module uses the lightweight and improved Light-HCT network model for prediction. The specific steps are as follows:
[0067] Step (1), first construct the data set required for this module. The length of the initial candidate mode list in the PU mode selection is 3 (corresponding to PU block sizes of 64×64, 32×32, 16×16) or 8 (corresponding to PU block sizes of 8×8 and 4×4). It is observed that the position of the best mode of the PU block in the candidate mode list after RMD rough selection is not fixed, and the mode after the best mode is redundant, which will increase the RDO calculation amount. Therefore, the present invention excludes these redundant modes by predicting the position of the best mode in the candidate list after RMD rough selection, so as to reduce the subsequent RDO calculation amount. The data set for PU mode selection covers PU blocks from 64×64 to 4×4. The label value of each PU block is label∈[0, 1, 2]. The predicted best mode position is also divided into three categories according to the label category and corresponds to the label value one by one. The corresponding relationship is as follows Figure 7 As shown, the corresponding relationship rules are as follows:
[0068] ① The length of the initial candidate list of PU blocks of sizes 64×64, 32×32, and 16×16 is 3. If the best mode of the current PU block is located in the first position in the mode list after RMD rough selection, label = 0; if the best mode is located in the second position in the mode list after RMD rough selection, label = 1; in other cases, label = 2.
[0069] ②For PU blocks of 8×8 and 4×4 sizes, since the length of the initial candidate list is 8, the position of the segmented interval is used to correspond to the label value. If the best mode of the current PU block is located at the 1st or 2nd position in the mode list after RMD rough selection, label = 0; if the best mode of the current PU block is located at the 3rd or 4th position in the mode list after RMD rough selection, label = 1; in other cases, label = 2.
[0070] Step (2), use the Pytorch deep learning library to build a lightweight HCT model, namely the Light-HCT model. Different from the HCT model, Light-HCT reduces the number of layers of the Transformer encoder module from the original 7 layers to 1 layer. Compared with the classification accuracy function of HCT, Light-HCT is more inclined to fit the position of the best mode in the PU block after RMD rough selection, and the mode selection of the PU block is jointly determined by texture features, boundary curvature, quantization parameters and boundary direction. The deep learning method cannot automatically extract so many features, so it is difficult to train using the classification method. In particular, since the 8×8 and 4×4 PU blocks are very small and do not require a very complex network structure, we have simplified the convolution module based on the Light-HCT model, changing the parameters of the convolution layer to a convolution kernel size of 3×3, with a step size and padding of 1, that is, no dimensionality reduction operation, and the subsequent maximum pooling layer is removed.
[0071] Step (3), linear regression prediction is used for model training, and the loss function uses the MSELoss function in Pytorch:
[0072]
[0073] Where output is the network output value vector, length is 3, and value is the true value vector, which is obtained by comparing output and label, such as Figure 8 As shown, the value vector acquisition rules are as follows:
[0074] ①label=0. If the maximum value in output=[x,y,z] appears at the subscript 0, then value=output; if the maximum value appears at the subscript 1, then value=[y,x,z]; if the maximum value appears at the subscript 2, then value=[z,y,x].
[0075] ②In the case of label = 1. If the maximum value in output = [x, y, z] appears at the subscript 0, then value = [y, x, z]; if the maximum value appears at the subscript 1, then value = output; if the maximum value appears at the subscript 2, then value = [x, z, y].
[0076] ③In the case of label = 2. If the maximum value in output = [x, y, z] appears at the subscript 0, then value = [z, y, x]; if the maximum value appears at the subscript 1, then value = [x, z, y]; if the maximum value appears at the subscript 2, then value = output.
[0077] The purpose of the above exchange mechanism is to move the maximum value position frequency predicted by the model closer to the position corresponding to the label value, so as to fit the corresponding category of each PU block and achieve the purpose of balancing coding performance and complexity.
[0078] Step (4), the present invention also optimizes the most likely mode MPM part after RMD rough selection. First, the mode with the minimum SAD and SATD cost value in the candidate mode list after RMD rough selection is obtained, and its SATD cost is recorded as J SATDmin, SAD (Sum of Absolute Difference) is the sum of absolute differences, which represents the size of the residual value between the original image block and the predicted image block, and SATD (Hadamard transformed SAD) refers to the size of the residual after transformation. The predicted residual is first transformed by Hadamard, and then the absolute value of the residual is summed. These two cost values reflect the RD-cost size of the PU block to a certain extent, and can be used for preliminary mode screening. The calculation formulas of these two cost functions are as follows:
[0079] TD(x,y)=|Orig(x,y)-Pred(x,y)| (7)
[0080] SAD=∑ x,y TD(x,y) (8)
[0081]
[0082] After reaching the MPM part, since there are three MPMs obtained from adjacent PU blocks, firstly, these MPMs are compared in order to see if they are the same as the mode with the smallest SAD and SATD costs. If so, the MPM process is terminated and this mode is sent to the RDO calculation as the best mode, and other modes are discarded; otherwise, the SATD cost value of this MPM is calculated and compared with the adaptive threshold AT. The adaptive threshold AT defined in the present invention is:
[0083] AT=ρ×J SATDmin (10)
[0084] The proportionality coefficient ρ = 1.3, which is obtained based on the statistical analysis of a large number of video sequence experiments. If the SATD cost of the MPM is greater than AT, the mode is not added to the RDO candidate mode list, otherwise it is added to the candidate list, provided that the mode does not exist in the candidate list, and then the next MPM is judged to replace the original encoder's MPM operation until all MPM mode judgments are completed.
Claims
1. A fast video encoding method based on deep neural network, Features The specific implementation includes a CU partitioning module based on a deep neural network and a PU mode selection module based on neighborhood correlation. When encoding a CU block intraframe, the PU mode selection module will first be used to calculate the rate-distortion cost. At this time, the PU mode selection module based on neighborhood correlation is first used for optimization, and the number of candidate modes calculated by RDO is reduced through the prediction results of the lightweight HCT model. After the PU mode selection is completed, the encoder will make a CU block depth judgment to determine whether the CU block is divided. At this time, the CU partitioning module based on a deep neural network is optimized, and the prediction results are obtained from the HCT model to determine whether to terminate the division in advance. Otherwise, the PU mode selection and CU block partitioning judgment of the sub-CU block will continue to be divided downward. The CU division module based on deep neural network is specifically implemented as follows: Step (I), construct a data set for network model training in HEVC intra-frame mode: the data set comes from YUV video sequences of various resolutions including: CIF (352×288), 480p (832×480), 720p (1280×720), 1080p (1920×1080), WQXGA (2560×1600); HEVC encoder HM16.9 is used to encode the images in the data set to obtain CU blocks and their positive and negative sample labels; the data set includes a training set, a validation set and a test set, and each data set is divided into four subsets according to four QPs (22, 27, 32, 37); Step (II) constructs a deep neural network with three CU blocks of 64×64, 32×32, and 16×16 to form a hierarchical convolutional network HCT structure. The hierarchical convolutional network HCT is composed of ViT and CNN. The HCT is trained through the corresponding training set, and the HCT model is determined and saved by the validation set. The generalization ability of the HCT model is judged by the test set; the objective function of the HCT model training is the cross entropy loss function: Where output is the output vector of the HCT model, target is the label value, and N is the length of the output vector; Step (III), the hierarchical convolutional network HCT consists of a convolutional module, an encoder module, a sequence pooling layer, and a fully connected layer; first, the brightness component of the CU block is sent to the hierarchical convolutional network HCT, and a feature map with local feature information is output through the convolutional module. The convolutional module contains a convolutional layer and a maximum pooling layer, and each layer is activated by a linear rectification function to improve the nonlinearity of the model; then the feature map is flattened into one dimension and exchanged with the number of feature maps, that is, the flattening and flipping operation; assuming that the input image x∈R C×H×W , where C represents the number of input images, H is the height of the image, W is the width of the image, and the output feature data x after the convolution module is 0 as follows: x 0 =Transpose(Flatten(MaxPool(Conv2d(x)))) (2) Next, the feature data x 0 The vector is added to the position vector and sent to the Encoder module for global information extraction. The Encoder module has 7 layers, each of which consists of a multi-head self-attention layer MSL and a feed-forward convolution layer FCL. Both sub-layers are preceded by layer normalization LN operations. 0 After passing through the multi-head self-attention layer, the output data is the same as x 0 Add to get new feature data x 1 , x 1 After the feedforward convolution layer, its output value is the same as x 1 Add to get the characteristic data x 2 , the formula is as follows: x 1 =x 0 +MSL(LN(x 0 )) (3) x 2 =x 1 +FCL(LN(x 1 )) (4) Finally, the classification vector is obtained through the sequence pooling layer. The sequence pooling adopts the mapping transformation T:R b×n×d →R b×d , b represents the batch size, n represents the number of feature data, and d represents the size of each feature data; this operation converts the entire Encoder output feature data x 2 Directly transformed into a classification vector, which contains relevant information of each part of the input image, to replace the additional classification vector added in ViT; finally, the classification vector is output as a binary classification result through a fully connected layer and softmax, and the final prediction value is the subscript of the maximum output value; Step (IV), the HCT model is trained using the stochastic gradient descent method, and a total of 12 HCT models with the highest accuracy for the three types of CU blocks under four QPs are saved. The trained HCT model uses an early termination mechanism to predict the partitioning results of 64×64, 32×32, and 16×16 blocks from top to bottom. The prediction results of the model are divided into two categories: 0 represents no partitioning, and 1 represents partitioning; when the prediction result of a certain type of block is 0, the quadtree partitioning will not be continued during encoding; The contrast value between the two-class vectors is used as the threshold Thr. When Thr is less than the fixed value λ, the CU block is checked in the original encoding method. The formula is as follows: Among them, output i is the binary classification output vector of the i*i size block, and the fixed value λ is divided into three categories according to the block size, and the size ratio is 4:2:1; The PU mode selection module based on neighborhood correlation is specifically implemented as follows: Step (1), obtain the sample label value label∈[0,1,2] of each PU block during intra-frame mode selection through HM coding, and the acquisition rule is as follows: for PU blocks of size 64×64, 32×32, and 16×16, the original length of the candidate list after RMD rough selection is 3. If the best mode of the PU block during mode selection is the first in the candidate list after RMD rough selection, label=0, and the length of the candidate list after RMD rough selection becomes 1; if the best mode of the PU block during mode selection is the second in the candidate list after RMD rough selection, label=1, and the length of the candidate list after RMD rough selection becomes 2; in other cases, label=2, and the length of the candidate list after RMD rough selection becomes 2. The length of the candidate list after RMD rough selection is 3; for 8×8 and 4×4 PU blocks, since the original length of the candidate list is 8, we also divide it into three intervals to correspond to label=0, 1, and 2, respectively: if the best mode of the PU block after mode selection is located in the first or second position in the candidate list after RMD rough selection, then label=0, and the length of the candidate list after RMD rough selection becomes 2; if the best mode of the PU block during mode selection is located in the third or fourth position in the candidate list after RMD rough selection, then label=1, and the length of the candidate list after RMD rough selection becomes 4; in other cases, label=2, and the length of the candidate list after RMD rough selection is 8; Step (2), the data set of the PU mode selection module also comes from the video sequence mentioned in the block partition module, and adds PU block data of 8×8 and 4×4 sizes on the basis of the block partition data set; The pytorch deep learning library is used to build a lightweight HCT model, namely the Light-HCT model. Light-HCT reduces the number of layers of the Encoder module from the original 7 layers to 1 layer; the model is trained using the mean square error loss function for regression training: Where output is the model output vector with a length of 3, value is the true value vector obtained by comparing output with label, and N is the number of input images in each training; the rules for obtaining the true value vector are as follows: assuming output = [x, y, z], label = 0, if the maximum value in the output appears at the subscript 0, then value = output; if the maximum value appears at the subscript 1, then value = [y, x, z]; if the maximum value appears at the subscript 2, then value = [z, y, x]; similarly, when label = 1, if the maximum value in the output appears at the subscript 0, then value = [y, x, z]; if the maximum value appears at the subscript 1, then value = output; if the maximum value appears at the subscript 2, then value = [x, z, y]; when label = 2, if the maximum value in the output appears at the subscript 0, then value = [z, y, x]; if the maximum value appears at the subscript 1, then value = [x, z, y]; if the maximum value appears at the subscript 2, then value = output.