VVC inter-frame coding acceleration method and device based on lightweight neural network

The probability of candidate division mode is calculated by lightweight neural network MFLCNN, which solves the problem of low feature correlation in VVC inter-frame encoding, realizes encoding acceleration and hardware cost reduction, and improves coding efficiency.

CN120263987AActive Publication Date: 2025-07-04HUAQIAO UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510664268.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-04
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

In the existing VVC inter-frame encoding technology, the correlation between features and the final division mode is low, resulting in less obvious encoding acceleration effect, and complex neural network models increase hardware cost and time consumption.

Method used

The lightweight neural network MFLCNN is adopted to obtain real-time feature data and residual information, and use the residual processing module, feature processing module and feature fusion module to calculate the probability of candidate division modes, select several modes with the highest probability for encoding, and skip the mode with lower probability to achieve coding acceleration.

Benefits of technology

A significant encoding acceleration effect was achieved, with an average time saving of 47.57%, and the bit rate of the code stream was only increased by 2.15%, which was better than the existing methods and reduced hardware cost and time consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263987A_ABST
    Figure CN120263987A_ABST
Patent Text Reader

Abstract

The invention discloses a VVC inter-frame coding acceleration method and device based on a lightweight neural network, and relates to the technical field of video coding, and the method comprises the following steps: obtaining real-time feature data and residual error information before the start of a division mode test of a CU; the feature data and the residual information are sent to an MFLCNN network, and the probability of each candidate division mode is calculated; a plurality of division modes with the highest probability are selected according to trade-off configuration preset in advance, and a plurality of modes with the lower probability are actively skipped in the encoding process, so that encoding acceleration is realized; and the MFLCNN network extracts residual features from residual information by using a residual processing module, extracts variable features from feature data by using a feature processing module, and finally integrates the residual features and the variable features by using a feature fusion module, and outputs a probability value of a possible division mode. According to the method, feature analysis and mode prediction are carried out by using the MFLCNN, VVC inter-frame coding acceleration is realized, the realization is simple, and the effect is remarkable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video coding technology, and in particular to a VVC inter-frame coding acceleration method and device based on a lightweight neural network. Background Art

[0002] Video coding acceleration technology is to reduce the computational complexity and time consumption in the video coding process by optimizing the algorithm, while maintaining the rate-distortion performance of video coding as much as possible, so as to improve the efficiency and real-time performance of video coding. This technology is widely used in video streaming, video conferencing, video surveillance, video storage and other fields, and is of great significance for reducing video processing costs and improving user experience.

[0003] At present, the research on VVC video coding acceleration technology mainly focuses on the use of three features: original pixels, residual pixels, and motion vectors, and reduces the encoding time by terminating the pattern test in the encoding process in advance. These methods can reduce the complexity of VVC inter-frame coding to a certain extent. However, an in-depth analysis of these existing methods will reveal that they have many obvious defects. First, the correlation between the features selected and used by these methods and the final division pattern is low. In the encoding process, due to this low correlation, although some time is saved on the surface, the rate-distortion performance of the encoder is greatly reduced. The reduction in rate-distortion performance directly leads to an increase in the amount of bitstream data, which is an unfavorable factor for video storage and transmission, and will increase storage costs and bandwidth requirements. Second, in order to obtain the motion vector data required by the decision maker, some methods have to spend a lot of extra time for motion search. Motion search itself is a process with a large amount of computation. This extra time consumption seriously weakens the effect of encoding acceleration, making it difficult to effectively achieve the originally expected acceleration goal. Third, in order to enable the decision maker to obtain better pattern prediction effects, some methods often require the neural network model to be extremely complex. Although complex neural network models may improve the accuracy of pattern prediction to a certain extent, they also bring huge time consumption. In some cases, it is even necessary to rely on GPU acceleration to achieve the purpose of encoding acceleration, which undoubtedly increases the hardware cost and system complexity, and limits the promotion of these methods in practical applications. Therefore, how to design a lightweight VVC inter-frame fast algorithm has become a difficult problem that needs to be broken through in the field of video coding, and more innovative and targeted technologies are needed to deal with this problem. Summary of the invention

[0004] The purpose of the present invention is to solve the problems in the prior art such as low correlation between the selected features and the final division mode and unclear acceleration effect.

[0005] The technical solution adopted by the present invention to solve its technical problems is to provide a VVC inter-frame coding acceleration method based on a lightweight neural network, including the following steps:

[0006] Before the start of the CU partition mode test, obtain the real-time feature data and residual information therein;

[0007] Send the feature data and residual information into the MFLCNN network to calculate the probabilities of each candidate partition mode;

[0008] Select several partition modes with the highest probabilities according to the pre-set trade-off configuration, and then the encoder tests a mode with the lowest rate-distortion cost from them;

[0009] The MFLCNN network uses a residual processing module to extract residual features from the residual information, uses a feature processing module to extract variable features from the feature data, and finally uses a feature fusion module to integrate the residual features and variable features to output the probability values of possible partition modes.

[0010] Preferably, before the start of the CU partition mode test, obtaining the real-time feature data and the motion search residual image therein specifically includes: in the stage after all non-partition mode tests in the encoder are completed and before the partition mode test starts, collecting the motion search residual image and auxiliary features;

[0011] Among them, the motion search residual image is the result of subtracting the predicted image obtained by the non-partition mode from the original image, and the auxiliary features include intermediate variables in the best non-partition mode selected by the encoder.

[0012] Preferably, the auxiliary features include: quantization parameter QP value and the QP value of the current frame the total distortion of the motion compensation image of the best non-partition mode the overall rate-distortion cost of the best non-partition mode the number of encoded bits conventional merge mode flag variable merge mode flag variable merge mode flag variable with motion vector difference CU no-residual flag whether to use the no-residual method for the merge mode with motion vector difference affine mode flag geometric partition mode the depth of the current CU partition the partition depth in the quadtree manner the partition depth in the binary tree manner and the partition depth in the multi-tree manner , provide information on the upper-level partitioning mode.

[0013] Preferably, the auxiliary features need to be preprocessed, expressed as:

[0014] ;

[0015] ;

[0016] ;

[0017] ;

[0018] .

[0019] Preferably, the residual processing module of the MFLCNN network includes several different processing structures for compressing residual images of different sizes into a unified data size; each processing structure includes several downsampling sub-modules and a flattening operation, and the downsampling sub-module includes a Conv2D layer, a LayerNorm layer, and a ReLU layer connected in sequence; the processing process of the residual processing module includes the following steps:

[0020] Receive the absolute value generated by the conversion of the residual image and input it into the corresponding processing unit according to the CU size;

[0021] The corresponding processing unit uses the downsampling sub-module to compress residual images of different sizes into a unified data size.

[0022] Preferably, the feature processing module of the MFLCNN network includes three fully connected layers, and each fully connected layer is followed by a LeakyReLU layer; the FPM receives an auxiliary feature group with an input dimension of 16 and outputs variable features of the same dimension.

[0023] Preferably, the processing of the downsampling sub-module of the MFLCNN network includes the following steps:

[0024] Select the structure for integrating features according to the CU size; for larger-sized CUs, use a normal structure to integrate features, and the normal structure includes an input layer with a dimension of 32, a hidden layer with a dimension of 16, a hidden layer with a dimension of 8, and an output layer with a dimension of 6; for smaller-sized CUs, use a simplified structure to integrate features, and the simplified structure includes an input layer with a dimension of 32, a hidden layer with a dimension of 16, and an output layer with a dimension of 4.

[0025] Generate the probability value of each partitioning mode through the Softmax layer, and the output dimension depends on the number of candidate partitioning modes for the given CU size.

[0026] The present invention also provides a VVC inter-frame coding acceleration device based on a lightweight neural network, including:

[0027] A data collection module, which acquires real-time feature data and residual information therein before the start of the CU partition mode test;

[0028] A mode prediction module, which sends the feature data and residual information into the MFLCNN network to calculate the probabilities of each candidate partition mode;

[0029] A mode selection module and an encoding module, which select several partition modes with the highest probabilities according to a pre-set trade-off configuration, and then the encoder tests a mode with the lowest rate-distortion cost from them;

[0030] The MFLCNN network extracts residual features from the residual information by using a residual processing module, extracts variable features from the feature data by using a feature processing module, and finally integrates the residual features and variable features by using a feature fusion module to output the probability values of possible partition modes.

[0031] The present invention has the following beneficial effects:

[0032] (1) The present invention designs an extremely lightweight neural network for feature analysis and mode prediction, realizing the acceleration of VVC inter-frame coding, with simple implementation and remarkable effects;

[0033] (2) The present invention proposes an auxiliary feature group with higher correlation with the final partition mode to make up for the information loss of commonly used residual features, original pixels, and motion vector features.

[0034] The following further elaborates on the present invention in detail with reference to the drawings and embodiments, but the present invention is not limited to the embodiments. Description of the Drawings

[0035] Figure 1 It is a method step diagram of a VVC inter-frame coding acceleration method based on a lightweight neural network according to an embodiment of the present invention;

[0036] Figure 2 It is a structure diagram of the MFLCNN network of a VVC inter-frame coding acceleration method based on a lightweight neural network according to an embodiment of the present invention;

[0037] Figure 3 It is a structural schematic diagram of the residual processing module of a VVC inter-frame coding acceleration method based on a lightweight neural network according to an embodiment of the present invention;

[0038] Figure 4 It is a structural schematic diagram of the downsampling sub-module of a VVC inter-frame coding acceleration method based on a lightweight neural network according to an embodiment of the present invention;

[0039] Figure 5 This is a schematic structural diagram of a VVC inter-frame coding acceleration device based on a lightweight neural network according to an embodiment of the present invention. Specific implementation manners

[0040] In order to effectively solve the problems of insufficient acceleration amplitude of the current VVC video inter-frame coding method and a large decrease in the rate-distortion performance of the encoder, the present invention provides a VVC inter-frame coding acceleration method based on a lightweight neural network. As shown in Figure 1 the following steps are included:

[0041] S101, before the start of the CU partition mode test, obtain the real-time feature data and residual information therein;

[0042] S102, send the feature data and residual information into the MFLCNN network to calculate the probabilities of each candidate partition mode;

[0043] S103, select several partition modes with the highest probabilities according to the pre-set trade-off configuration, and then the encoder tests a mode with the lowest rate-distortion cost from them;

[0044] The MFLCNN network extracts residual features from the residual information by using a residual processing module, extracts variable features from the feature data by using a feature processing module, and finally integrates the residual features and variable features by using a feature fusion module to output the probability values of possible partition modes; select several partition modes with the highest probabilities, and then the encoder tests a mode with the lowest rate-distortion cost from them. By narrowing the mode traversal range of the encoder during the coding process, coding acceleration is achieved.

[0045] Specifically, the MFLCNN network (lightweight neural network) extracts residual features from the residual information by using a residual processing module, extracts variable features from the feature data by using a feature processing module, and finally integrates the residual features and variable features by using a feature fusion module to output the probability values of possible partition modes.

[0046] Specifically, in S101, after the tests of non-partition modes such as the fusion mode, the motion search mode, and the affine mode are completed, before the start of the CU partition mode test, extract the test results of the best non-partition mode. These features include the motion search residual image and a series of auxiliary variables; the residual image is the result of subtracting the predicted image obtained by the non-partition mode from the original image; the auxiliary features include 16 variables, which are described as follows:

[0047] Quantization parameter QP value and the QP value of the current frame and the total distortion of the motion compensation image of the best non-partition mode and the overall rate-distortion cost of the best non-partition mode , the number of encoded bits , the conventional merge mode flag variable , the merge mode flag variable , the merge mode flag variable with motion vector difference , the CU no-residual flag , whether to use the no-residual method for the merge mode with motion vector difference , the affine mode flag , the geometric partition mode ; and four partition stage indicator variables, , , , , which respectively represent the depth of the current CU partition, the partition depth in the quadtree mode, the partition depth in the binary tree mode, and the partition depth in the multi-tree mode, to provide information on the upper-level partition mode.

[0048] represents the total distortion of the motion-compensated image of the best non-partition mode; represents the overall rate-distortion cost of this non-partition mode. represents the number of bits used for encoding. The higher the values of these three parameters, the lower the prediction accuracy of the non-partition mode is usually indicated, thus increasing the possibility of the CU being partitioned. The value ranges of these three variables in the encoder are very wide, so they cannot be directly input into the neural network for training. Therefore, these variables need to be preprocessed. The preprocessing methods are as follows:

[0049] ;

[0050] ;

[0051] ;

[0052] and have the same preprocessing methods as above:

[0053] ;

[0054] ;

[0055] Specifically, in the S102, the residual image and the preprocessed auxiliary features are sent into the MFLCNN network for feature analysis and calculation of the partition mode probability. See Figure 2As shown, MFLCNN consists of three main components: a Residual Processing Module (RPM) for extracting features from residual pixels, a Feature Processing Module (FPM) for processing the auxiliary feature set, and a Feature Fusion Module (FFM) for integrating the information from the first two modules and outputting the probability values of possible partitioning patterns. The details of each module are described as follows:

[0056] The RPM is used to analyze the matching degree of motion-compensated images in different regions of the CU. Before inputting to the RPM, the residual image is converted to its absolute value. See Figure 3 As shown, the RPM consists of 14 different structures, each structure being specifically designed for a CU of a specific size. Each structure includes several downsampling sub-modules (DSM) and a flattening operation to compress residual images of different sizes into a unified data size. The structure of the DSM is shown in Figure 4 As shown, it consists of three layers: a Conv2D layer, a LayerNorm layer, and a ReLU layer. Five parameters are specifically designed for the DSM and provided to the Conv2D layer: represents the number of input channels of the Conv2D layer, represents the number of output channels, represents the kernel size, represents the kernel stride, represents the padding. After the Conv2D layer, the LayerNorm layer normalizes the data, normalizing across all dimensions except the batch dimension. For example, if the feature map generated by the Conv2D layer has dimensions, LayerNorm will normalize the data within the dimensions. After normalization, the data is passed through the ReLU activation function to introduce non-linearity into the network.

[0057] The FPM module is used to analyze the auxiliary feature group. The FPM consists of three fully connected layers, each followed by a LeakyReLU layer. The FPM accepts an auxiliary feature group with an input dimension of 16 and outputs features of the same dimension for subsequent use. The FFM is used to integrate and analyze the output features of the first two modules and generate the probability values of possible partitioning patterns through a Softmax layer. The output dimension of the FFM depends on the number of candidate partitioning patterns for a given CU size. The FFM has two structures: a normal structure and a simplified structure. For some larger-sized CUs, the FFM module adopts the normal structure to more comprehensively analyze the information from the first two modules. For smaller-sized CUs, the FFM adopts the simplified structure to save more time and reduce the risk of network overfitting. Finally, the test of partitioning patterns with lower probabilities is skipped to achieve an improvement in coding speed.

[0058] Specifically, the experimental platform is Ubuntu 20.24.2, and the CPU is AMD EPYC 9654 96 Core processor. In the time comparison test, both the original encoder and the accelerated encoder were tested using this experimental platform. This solution achieved an average time saving of 47.57%, and the bit rate of the bitstream only increased by 2.15% under the same quality. The method proposed by Pan et al. in the literature "A cnn-based fast inter coding method for vvc" achieved a time saving of 30.63%, and the bit rate of the bitstream increased by 3.18% under the same quality. The method proposed by Tissier et al. in the literature "Machine learning based efficient qt-mtt partitioning for vvc inter coding" achieved a time saving of 43.4%, and the bit rate of the bitstream increased by 2.33% under the same quality. This solution is superior to the comparison solutions in terms of both time saving and suppression of bit rate increase.

[0059] See Figure 5 As shown, it is a schematic structural diagram of an apparatus for accelerating VVC inter-frame coding based on a lightweight neural network according to an embodiment of the present invention, including:

[0060] A data collection module 501, which acquires real-time feature data and residual information therein before the start of the CU partition mode test;

[0061] A mode prediction module 502, which sends the feature data and residual information into the MFLCNN network to calculate the probabilities of each candidate partition mode;

[0062] A model selection module 503, which selects several partition modes with the highest probabilities according to a pre-set trade-off configuration, and the encoder then tests a mode with the lowest rate-distortion cost from them;

[0063] The apparatus realizes coding acceleration by narrowing the mode traversal range of the encoder during the coding process; the implementation of each module in the apparatus is the same as that of each module in a method for accelerating VVC inter-frame coding based on a lightweight neural network, and will not be described repeatedly here.

[0064] It can be seen that a method and apparatus for accelerating VVC inter-frame coding based on a lightweight neural network proposed by the present invention acquire real-time feature data and residual information therein before the start of the CU partition mode test, input and design the MFLCNN network to calculate the probabilities of each candidate partition mode, then select several partition modes with the highest probabilities according to a pre-set trade-off configuration, and actively skip several partition modes with lower probabilities during the coding process, thereby realizing coding acceleration.

[0065] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for accelerating VVC inter-frame coding based on a lightweight neural network, characterized in that, It includes the following steps: Before the partition mode test of the CU starts, obtain the real-time feature data and residual information therein; Send the feature data and residual information into the MFLCNN network to calculate the probabilities of each candidate partition mode; Select several partition modes with the highest probabilities according to the pre-set trade-off configuration, and then the encoder tests a mode with the lowest rate-distortion cost from them; The MFLCNN network extracts residual features from the residual information by using the residual processing module, extracts variable features from the feature data by using the feature processing module, and finally integrates the residual features and variable features by using the feature fusion module to output the probability values of possible partition modes.

2. The VVC inter-frame coding acceleration method based on a lightweight neural network according to claim 1, wherein Before the partition mode test of the CU starts, obtain the real-time feature data and motion search residual image therein. Specifically, in the stage after all non-partition mode tests in the encoder are completed and before the partition mode test starts, collect the motion search residual image and auxiliary features; Among them, the motion search residual image is the result of subtracting the predicted image obtained by the non-partition mode from the original image, and the auxiliary features include intermediate variables in the best non-partition mode selected by the encoder.

3. The VVC inter-frame coding acceleration method based on a lightweight neural network according to claim 2, wherein The auxiliary features include: quantization parameter QP value , QP value of the current frame , total distortion of the motion compensated image in the best non-partition mode , overall rate-distortion cost in the best non-partition mode , number of encoded bits , conventional merge mode flag variable , merge mode flag variable , merge mode flag variable with motion vector difference , CU no-residual flag , whether to use the no-residual method for the merge mode with motion vector difference , affine mode flag , geometric partition mode , depth of the current CU partition , partition depth in the quadtree manner , partition depth in the binary tree manner and partition depth in the multi-tree manner , providing information on the upper-level partition mode.

4. The VVC inter-frame coding acceleration method based on a lightweight neural network according to claim 3, wherein The auxiliary features need to be preprocessed, which is expressed as: ; ; ; ; 。 5. The VVC inter-frame coding acceleration method based on a lightweight neural network according to claim 1, wherein The residual processing module of the MFLCNN network includes several different processing structures for compressing residual images of different sizes into a unified data size; each processing structure includes several downsampling sub-modules and a flattening operation, and the downsampling sub-module includes a Conv2D layer, a LayerNorm layer, and a ReLU layer connected in sequence; the processing process of the residual processing module includes the following steps: Receive the absolute value generated by converting the residual image and input it into the corresponding processing unit according to the CU size; The corresponding processing unit uses the downsampling sub-module to compress residual images of different sizes into a unified data size.

6. The VVC inter-frame coding acceleration method based on a lightweight neural network according to claim 1, wherein The feature processing module of the MFLCNN network includes three fully connected layers, and each fully connected layer is followed by a LeakyReLU layer; the FPM receives an auxiliary feature group with an input dimension of 16 and outputs variable features with the same dimension.

7. The VVC inter-frame coding acceleration method based on a lightweight neural network according to claim 1, wherein The processing of the downsampling sub-module of the MFLCNN network includes the following steps: Select the structure for integrating features according to the CU size; for larger-sized CUs, use a common structure to integrate features. The common structure includes an input layer with a dimension of 32, a hidden layer with a dimension of 16, a hidden layer with a dimension of 8, and an output layer with a dimension of 6; for smaller-sized CUs, use a simplified structure to integrate features. The simplified structure includes an input layer with a dimension of 32, a hidden layer with a dimension of 16, and an output layer with a dimension of 4; Generate the probability value of each partition mode through the Softmax layer, and the output dimension depends on the number of candidate partition modes for the given CU size.

8. An inter-frame coding acceleration device for VVC based on a lightweight neural network, characterized in that It includes: A data collection module, which obtains the real-time feature data and residual information therein before the partition mode test of the CU starts; A mode prediction module, which sends the feature data and residual information into the MFLCNN network to calculate the probabilities of each candidate partition mode; A mode selection module and an encoding module select several partitioning modes with the highest probabilities according to a pre-set trade-off configuration, and then the encoder tests a mode with the lowest rate-distortion cost from them; The MFLCNN network uses a residual processing module to extract residual features from residual information, uses a feature processing module to extract variable features from feature data, and finally uses a feature fusion module to integrate the residual features and variable features to output the probability values of possible partitioning modes.

Citation Information

Patent Citations

  • Novel reconstruction method based on distributed compressed video sensing system

    CN112637599A

  • VVC multi-level fast inter-frame coding system and method based on neural network

    CN117915104A

  • VVC-SCC intra-frame coding method and device based on multi-stage irregular coding unit division

    CN119299671A

  • Inter coding using deep learning in video compression

    WO2024006167A1