VVC-SCC Coding Method and Apparatus Based on Multi-Size and Neighborhood Combined Prediction Patterns

By combining prediction patterns with multiple sizes and neighborhoods, an intra-frame coding unit prediction pattern classification model is constructed, which solves the problem of high computational complexity of the VVC-SCC encoder and achieves savings in coding time and improvement in prediction accuracy.

CN120639979BActive Publication Date: 2025-10-28HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511124481.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-10-28
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing VVC-SCC encoders have high computational complexity when processing screen content video, and existing fast prediction algorithms ignore the relationship between patterns and pixels, resulting in low encoding efficiency.

Method used

A prediction mode combining multiple sizes and neighborhoods is adopted. By constructing a prediction mode classification model for multi-size intra-coding units, prediction mode selection is performed using sub-networks of different sizes and neighborhood prediction probabilities, thereby reducing the prediction mode selection time of coding units.

Benefits of technology

Without compromising subjective quality, it saves coding time, reduces the error rate of network prediction, and achieves a good trade-off between computational complexity and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639979B_ABST
    Figure CN120639979B_ABST
Patent Text Reader

Abstract

This invention discloses a VVC-SCC coding method and apparatus based on a prediction mode combining multiple sizes and neighborhood, relating to the field of video coding. The method includes: inputting the size and pixel value of the current coding unit into a trained multi-size intra-coding unit prediction mode classification model corresponding to specified quantization parameters; if the minimum of the height and width of the current coding unit is greater than a size threshold, then inputting the pixel value of the current coding unit into a first sub-network to obtain a first pixel feature; otherwise, inputting it into a second sub-network to obtain a second pixel feature; then passing it through a classification module to obtain the prediction probability corresponding to each prediction mode; calculating the neighborhood prediction probability corresponding to each prediction mode based on the usage of prediction modes in adjacent reconstructed coding units located in the current coding unit; calculating the correction probability corresponding to each prediction mode and making a prediction mode decision. This invention aims to solve the problem of high encoder computational complexity leading to long coding time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video coding, and specifically to a VVC-SCC coding method and apparatus based on a prediction pattern combining multiple sizes and neighborhood. Background Technology

[0002] With the widespread adoption of mobile devices, screen content video, as a new form of video representation, has received increasing attention. Compared to natural content video, screen content video features sharp edges, contains numerous flat areas, and repeating identical patterns and text. Using traditional video coding standards to process screen content video often results in poor compression, leading to text distortion and image blurring. Therefore, the Joint Video Coding Unit (JCD) developed a new generation of screen content coding standards based on the High Efficiency Video Coding - Screen Content Coding (HEVC-SCC): Versatile Video Coding - Screen Content Coding (VVC-SCC). This standard improves the compression performance of screen content video by adopting the intra-frame block copy mode and color palette prediction modes of HEVC-SCC, but its computational complexity also increases dramatically.

[0003] To reduce the computational complexity of the VVC-SCC encoder, previous fast prediction algorithms for intra-frame prediction modes in VVC-SCC primarily categorized the current encoding into screen content coding units (SMCs) and natural content coding units (NMCs). SMCs skipped intra-frame modes, while NMCs skipped intra-frame block copying modes, palette modes, and other similar modes, thus accelerating VVC-SCC encoding. However, this simplistic content classification ignored the relationship between modes and pixels. With breakthroughs in deep learning technology, it's possible to directly predict the optimal prediction mode for the current coding unit, thereby improving video encoding speed. Therefore, considering these issues, reducing the computational complexity of prediction modes for screen content coding units without affecting subjective quality is a key challenge for VVC-SCC acceleration algorithms. Summary of the Invention

[0004] The purpose of this application is to propose a VVC-SCC coding method and apparatus based on prediction modes combining multiple sizes and neighborhoods to address the aforementioned technical problems, thereby saving the selection time of prediction modes for VVC-SCC coding units.

[0005] In a first aspect, the present invention provides a VVC-SCC coding method based on prediction patterns combining multiple sizes and neighborhoods, comprising the following steps:

[0006] A multi-size intra-coding unit prediction pattern classification model is constructed and trained for different quantization parameters to obtain a trained multi-size intra-coding unit prediction pattern classification model corresponding to each quantization parameter. The multi-size intra-coding unit prediction pattern classification model includes a first sub-network, a second sub-network, and a classification module. The first sub-network and the second sub-network are respectively connected to the classification module.

[0007] The system acquires a video sequence of screen content and specified quantization parameters. Each video frame in the sequence is encoded using a VVC-SCC encoder. During encoding, the size, position, and pixel value of the current coding unit (CMU) are obtained from the video frame division. The size and pixel value of the current CMU are input into a trained multi-size intra-coding unit prediction mode classification model corresponding to the specified quantization parameters. If the minimum of the height and width of the current CMU is greater than a size threshold, the pixel value of the current CMU is input into the first sub-network to extract the first pixel feature. If the minimum of the height and width of the current CMU is less than or equal to the size threshold, the pixel value of the current CMU is input into the second sub-network to extract the second pixel feature. The first or second pixel feature is then processed by a classification module to obtain the prediction probability of the current CMU in each prediction mode.

[0008] The neighborhood prediction probability corresponding to each prediction mode is calculated based on the prediction modes used by the adjacent reconstructed coding units located to the left, above, and top left of the current coding unit; the prediction probability of the current coding unit in each prediction mode is added to its corresponding neighborhood prediction probability to obtain the correction probability corresponding to each prediction mode.

[0009] The prediction mode decision is made based on the correction probability corresponding to each prediction mode, determining whether to place the prediction mode in the rate-distortion optimization list.

[0010] Preferably, the first sub-network includes a first sampling layer, a first convolutional layer and a second convolutional layer connected in sequence, the second sub-network includes a second sampling layer and a third convolutional layer connected in sequence, and the classification module includes a first fully connected layer and a second fully connected layer connected in sequence.

[0011] Preferably, the kernel size of the first sampling layer is 4sw×4sh, and the stride is 4sw×4sh; the kernel size of the second sampling layer is 2sw×2sh, and the stride is 2sw×2sh; where w represents the width of the coding unit, h represents the height of the coding unit, sw represents the width unit, and sh represents the height unit. When w > h, sw = w / h, sh = 1; when w ≤ h, sw = 1, sh = h / w; when w = h, both sw and sh are 1; the kernel size and stride of the first, second, and third convolutional layers are all 2×2; the ReLU activation function is used in the first sampling layer, the first convolutional layer, the second sampling layer, the second convolutional layer, the third convolutional layer, and the first fully connected layer, and the Softmax function is used in the second fully connected layer.

[0012] Preferably, the neighborhood prediction probability corresponding to each prediction mode is calculated based on the prediction modes used by the adjacent reconstructed coding units located to the left, above, and above the current coding unit. Specifically, this includes:

[0013] The neighborhood prediction probability corresponding to each prediction mode is calculated using the following formula:

[0014] ;

[0015] in, , and These represent the usage parameters of the adjacent reconstructed coding units at the left, top, and top-left positions of the current coding unit, respectively. They determine whether the adjacent reconstructed coding units use the prediction mode with index i for prediction. If so, the value of the corresponding usage parameter is 1; otherwise, the value of the corresponding usage parameter is 0. This represents the neighborhood prediction probability corresponding to the prediction mode with index i. The value of i is 0, 1 or 2, where 0 represents Intra mode, 1 represents PLT mode and 2 represents IBC mode.

[0016] Preferably, prediction mode decisions are made based on the correction probability corresponding to each prediction mode, determining whether to include the prediction mode in the rate-distortion optimization list, specifically including:

[0017] The maximum value among all the corrected probabilities corresponding to the prediction modes is taken as the maximum corrected probability.

[0018] Iterate through the correction probabilities corresponding to each prediction mode, and determine whether the difference between the maximum correction probability and the correction probability corresponding to one of the prediction modes is less than or equal to the probability threshold. If so, add one of the prediction modes to the rate distortion optimization list; otherwise, remove one of the prediction modes from the rate distortion optimization list.

[0019] As a preferred method, during the training of the multi-size intra-frame coding unit prediction mode classification model, the screen content video dataset is encoded on the VVC-SCC standard test platform in full intra-frame configuration mode, and the true labels of the prediction modes with different quantization parameters are obtained.

[0020] For each quantization parameter, the corresponding true label of the prediction mode is selected to train the multi-size intra-coding unit prediction mode classification model. The loss function used in the training process is... for:

[0021] ;

[0022] Where N represents the batch training size. and These represent the actual label and the predicted label of the prediction mode with index i for the nth coding unit under the same quantization parameters. The predicted label is the index value corresponding to the maximum value of the prediction probabilities of all prediction modes output by the multi-size intra-coding unit prediction mode classification model. The value of i is 0, 1 or 2, where 0 represents Intra mode, 1 represents PLT mode and 2 represents IBC mode.

[0023] Secondly, the present invention provides a VVC-SCC coding device based on a prediction pattern combining multiple sizes and neighborhoods, comprising:

[0024] The model building module is configured to build a multi-size intra-coding unit prediction pattern classification model and train it for different quantization parameters to obtain a trained multi-size intra-coding unit prediction pattern classification model corresponding to each quantization parameter. The multi-size intra-coding unit prediction pattern classification model includes a first sub-network, a second sub-network, and a classification module. The first sub-network and the second sub-network are respectively connected to the classification module.

[0025] The multi-scale prediction module is configured to acquire the screen content video sequence and specified quantization parameters, and encode each video frame in the screen content video sequence using a VVC-SCC encoder. During the encoding process, the size, position, and pixel value of the current coding unit obtained from the video frame division are obtained. The size and pixel value of the current coding unit are input into the trained multi-scale intra-frame coding unit prediction mode classification model corresponding to the specified quantization parameters. In response to determining that the minimum value of the height and width of the current coding unit is greater than a size threshold, the pixel value of the current coding unit is input into the first sub-network to extract the first pixel feature; in response to determining that the minimum value of the height and width of the current coding unit is less than or equal to the size threshold, the pixel value of the current coding unit is input into the second sub-network to extract the second pixel feature. The first or second pixel feature is processed by the classification module to obtain the prediction probability of the current coding unit in each prediction mode.

[0026] The correction module is configured to calculate the neighborhood prediction probability corresponding to each prediction mode based on the prediction modes used by the adjacent reconstructed coding units located to the left, above, and upper left of the current coding unit; and to add the prediction probability of the current coding unit in each prediction mode to its corresponding neighborhood prediction probability to obtain the correction probability corresponding to each prediction mode.

[0027] The decision module is configured to make prediction mode decisions based on the correction probability corresponding to each prediction mode, and determine whether to place the prediction mode in the rate distortion optimization list.

[0028] Thirdly, the present invention provides an electronic device including one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.

[0029] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the implementations of the first aspect.

[0030] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in any of the implementations in the first aspect.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] (1) The VVC-SCC coding method based on multi-size and neighborhood combination proposed in this invention uses sub-networks with different structures in the multi-size intra-coding unit prediction mode classification model to classify the prediction modes of the current coding units of different sizes, and obtains the prediction probability corresponding to each prediction mode, thereby guiding the selection of prediction modes of intra-coding units of VVC encoder. Without affecting subjective quality, it saves coding time and accelerates the coding process of VVC-SCC.

[0033] (2) The VVC-SCC coding method based on multi-size and neighborhood combination prediction mode proposed in this invention calculates the neighborhood prediction probability corresponding to each prediction mode by using the usage of the neighboring coding units of the current coding unit; and uses the neighborhood prediction probability corresponding to each prediction mode to correct the prediction probability of the current coding unit in each prediction mode, thereby reducing the error rate of network prediction and achieving a good balance between computational complexity and time saving.

[0034] (3) The VVC-SCC coding method based on multi-size and neighborhood combination proposed in this invention trains a multi-size intra-coding unit prediction mode classification model for different quantization parameters, and uses the full intra-frame configuration mode to encode video frames to obtain the true labels of the prediction modes of intra-coding units with different quantization parameters. This can improve the diversity of the training dataset content and conform to the features contained in the screen content video test sequence as much as possible in a wide range of aspects, fields and angles. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a flowchart illustrating the VVC-SCC coding method based on multi-size and neighborhood combination prediction patterns, which is an embodiment of this application.

[0037] Figure 2 This is a schematic diagram of the multi-size intra-frame coding unit prediction mode classification model of the VVC-SCC coding method based on multi-size and neighborhood combination prediction modes, which is an embodiment of this application.

[0038] Figure 3 This is a schematic diagram showing the relative positions between the current coding unit and the adjacent reconstructed coding units in the VVC-SCC coding method based on multi-size and neighborhood combination prediction modes, which is an embodiment of this application.

[0039] Figure 4 A schematic diagram of a VVC-SCC coding device based on a prediction pattern combining multiple sizes and neighborhoods, according to an embodiment of this application;

[0040] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0042] Figure 1 This application illustrates an embodiment of a VVC-SCC coding method based on a prediction pattern combining multiple sizes and neighborhoods, comprising the following steps:

[0043] S1. Construct a multi-size intra-coding unit prediction pattern classification model and train it for different quantization parameters to obtain a trained multi-size intra-coding unit prediction pattern classification model corresponding to each quantization parameter. The multi-size intra-coding unit prediction pattern classification model includes a first sub-network, a second sub-network, and a classification module. The first sub-network and the second sub-network are respectively connected to the classification module.

[0044] In a specific embodiment, the first sub-network includes a first sampling layer, a first convolutional layer, and a second convolutional layer connected in sequence; the second sub-network includes a second sampling layer and a third convolutional layer connected in sequence; and the classification module includes a first fully connected layer and a second fully connected layer connected in sequence.

[0045] In a specific embodiment, the kernel size of the first sampling layer is 4sw×4sh, and the stride is 4sw×4sh; the kernel size of the second sampling layer is 2sw×2sh, and the stride is 2sw×2sh; where w represents the width of the coding unit, h represents the height of the coding unit, sw represents the width unit, and sh represents the height unit. When w > h, sw = w / h, sh = 1; when w ≤ h, sw = 1, sh = h / w; when w = h, both sw and sh are 1; the kernel size and stride of the first, second, and third convolutional layers are all 2×2; the ReLU activation function is used in the first sampling layer, the first convolutional layer, the second sampling layer, the second convolutional layer, the third convolutional layer, and the first fully connected layer, and the Softmax function is used in the second fully connected layer.

[0046] Specifically, a multi-size intra-coding unit prediction pattern classification model is first constructed. This model includes two sub-networks, each handling current coding units of different sizes. First, the size of the current coding unit is determined. The first sub-network handles current coding units where the minimum of the width and height of the input size is greater than a size threshold. The second sub-network handles current coding units where the minimum of the width and height is less than or equal to the size threshold. In one embodiment, the size threshold is 4. Therefore, the first sub-network handles current coding units where the minimum of the width and height is greater than 4, and the second sub-network handles current coding units where the minimum of the width and height is equal to 4. The pixel values ​​of the current coding unit are input to the corresponding sub-network based on its width (w) and height (h). Current coding units where the minimum value Min(w, h) is greater than 4 use the first sub-network; otherwise, the second sub-network is used. Min represents taking the minimum value. In another embodiment, refer to... Figure 2 If the size threshold is 8, then the first sub-network is used to process the current coding unit whose minimum value of the width and height of the input size is greater than 8, and the second sub-network is used to process the current coding unit whose minimum value of the width and height of the input size is less than or equal to 8. The pixel value of the current coding unit is input into the corresponding sub-network according to its width (w) and height (h). The first sub-network is used for current coding units where the minimum value of w and h, Min(w, h), is greater than 8; otherwise, the second sub-network is used. In one embodiment, the pixel value of the current coding unit is selected from the pixel value of the Y channel, i.e., the corresponding luminance component. The first sampling layer in the first sub-network uses an adaptive convolutional layer with a kernel size of 4sw×4sh and a stride of 4sw×4sh; the second sampling layer in the second sub-network uses an adaptive convolutional layer with a kernel size of 2sw×2sh and a stride of 2sw×2sh. When w > h, sw = w / h, sh = 1; otherwise, sw = 1, sh = h / w. If w = h, then both sw and sh are 1. The remaining first, second, and third convolutional layers are ordinary convolutional layers, and the kernel and stride of all ordinary convolutional layers are 2×2. Then, two fully connected layers output the prediction probabilities of the current coding unit in various prediction modes. In the embodiments of this application, there are three prediction modes for the intra coding unit: Intra mode, Intra Block Copy (IBC) mode, and Palette (PLT) mode. The embodiments of this application need to consider the size of the current coding unit, but do not limit the method and process of dividing the current coding unit.

[0047] In a specific embodiment, during the training process of the multi-size intra-frame coding unit prediction mode classification model, the screen content video dataset is encoded on the VVC-SCC standard test platform in full intra-frame configuration mode, and the true labels of prediction modes with different quantization parameters are obtained.

[0048] For each quantization parameter, the corresponding true label of the prediction mode is selected to train the multi-size intra-coding unit prediction mode classification model. The loss function used in the training process is... for:

[0049] ;

[0050] Where N represents the batch training size. and These represent the actual label and the predicted label of the prediction mode with index i for the nth coding unit under the same quantization parameters. The predicted label is the index value corresponding to the maximum value of the prediction probabilities of all prediction modes output by the multi-size intra-coding unit prediction mode classification model. The value of i is 0, 1 or 2, where 0 represents Intra mode, 1 represents PLT mode and 2 represents IBC mode.

[0051] Specifically, we first establish a prediction dataset for VVC-SCC. This prediction dataset includes a training set, a validation set, and a test set. The test set consists of standard screen content test sequences and is used for final testing. The training and validation sets are used to construct the database, which contains two parts: the first part is a video sequence of screen content with a resolution of 1792×1024, and the second part is a video set consisting of screen content images, which has three subsets: the first subset has a resolution of 1024×576, the second subset has a resolution of 1792×1024, and the third subset has a resolution of 2304×1280. The training and validation sets are encoded as follows: On the VVC-SCC standard test platform VTM19.0, video frames are encoded in full intra-frame configuration mode to obtain the true labels of the predicted modes of intra-coding units with quantization parameters of 22, 27, 32, and 37, thus constructing the VVC-SCC predicted mode dataset. For each quantization parameter, a multi-size intra-coding unit predicted mode classification model is trained. During the training process, the true labels of the predicted modes of intra-coding units with the corresponding quantization parameters are used in the calculation of the loss function. Therefore, a trained multi-size intra-coding unit predicted mode classification model corresponding to each quantization parameter is obtained.

[0052] S2: Obtain the screen content video sequence and specified quantization parameters. Encode each video frame in the screen content video sequence using a VVC-SCC encoder. During the encoding process, obtain the size, position, and pixel value of the current coding unit obtained from the video frame division. Input the size and pixel value of the current coding unit into the trained multi-size intra-coding unit prediction mode classification model corresponding to the specified quantization parameters. In response to determining that the minimum value of the height and width of the current coding unit is greater than the size threshold, input the pixel value of the current coding unit into the first sub-network to extract the first pixel feature. In response to determining that the minimum value of the height and width of the current coding unit is less than or equal to the size threshold, input the pixel value of the current coding unit into the second sub-network to extract the second pixel feature. The first pixel feature or the second pixel feature is processed by the classification module to obtain the prediction probability of the current coding unit in each prediction mode.

[0053] Specifically, a trained multi-size intra-coding unit prediction mode classification model corresponding to each quantization parameter is deployed. During inference, the trained multi-size intra-coding unit prediction mode classification model corresponding to the specified quantization parameter is selected according to the quantization parameter specified by the user. The video sequence of the screen content to be encoded is encoded using a VVC-SCC encoder. During the encoding process, the size and pixel value of the current coding unit are input into the trained multi-size intra-coding unit prediction mode classification model corresponding to the specified quantization parameter. First, the size of the current coding unit is determined. Based on the minimum value of the width (w) and height (h) of the current coding unit, the pixel value of the current coding unit is input into the corresponding sub-network. Then, the prediction probability of the current coding unit in various prediction modes is output through two fully connected layers.

[0054] S3, calculate the neighborhood prediction probability corresponding to each prediction mode based on the prediction modes used by the adjacent reconstructed coding units located to the left, above, and upper left of the current coding unit; add the prediction probability of the current coding unit in each prediction mode to its corresponding neighborhood prediction probability to obtain the correction probability corresponding to each prediction mode.

[0055] In a specific embodiment, the neighborhood prediction probability corresponding to each prediction mode is calculated based on the prediction modes used by the adjacent reconstructed coding units located to the left, above, and above the current coding unit. Specifically, this includes:

[0056] The neighborhood prediction probability corresponding to each prediction mode is calculated using the following formula:

[0057] ;

[0058] in, , and These represent the usage parameters of the adjacent reconstructed coding units at the left, top, and top-left positions of the current coding unit, respectively. They determine whether the adjacent reconstructed coding units use the prediction mode with index i for prediction. If so, the value of the corresponding usage parameter is 1; otherwise, the value of the corresponding usage parameter is 0. This represents the neighborhood prediction probability corresponding to the prediction mode with index i. The value of i is 0, 1 or 2, where 0 represents Intra mode, 1 represents PLT mode and 2 represents IBC mode.

[0059] Specifically, the prediction pattern of the current coding unit is significantly correlated with the prediction patterns of its neighboring reconstructed coding units. Since the prediction patterns of neighboring reconstructed coding units have already been determined, the prediction probability of the current coding unit's prediction pattern can be optimized based on the prediction patterns of its surrounding neighboring reconstructed coding units. (Reference) Figure 3 Here, we select the adjacent reconstructed coding units to the left, above, and top-left of the current coding unit, and assign values ​​to the usage parameters of each adjacent reconstructed coding unit in each prediction mode. We then calculate the average of the usage parameters of all adjacent reconstructed coding units in the same prediction mode, thus obtaining the neighborhood prediction probability corresponding to that prediction mode. If there are no adjacent reconstructed coding units to the left, above, and top-left of the current coding unit, we directly... It is 0.

[0060] S4. Make a prediction mode decision based on the correction probability corresponding to each prediction mode, and determine whether to place the prediction mode in the rate-distortion optimization list.

[0061] In a specific embodiment, step S4 specifically includes:

[0062] The maximum value among all the corrected probabilities corresponding to the prediction modes is taken as the maximum corrected probability.

[0063] Iterate through the correction probabilities corresponding to each prediction mode, and determine whether the difference between the maximum correction probability and the correction probability corresponding to one of the prediction modes is less than or equal to the probability threshold. If so, add one of the prediction modes to the rate distortion optimization list; otherwise, remove one of the prediction modes from the rate distortion optimization list.

[0064] Specifically, the correction probability corresponding to each prediction model is calculated using the following formula:

[0065] ;

[0066] in, This represents the corrected probability corresponding to the prediction pattern at index i. This represents the prediction probability of the current coding unit in the prediction mode corresponding to index i.

[0067] The prediction mode decision is made based on the correction probability corresponding to each prediction mode, and the specific formula is as follows:

[0068] ;

[0069] Where max(Q) i Let be the maximum value among Q0, Q1, and Q2, and Thr be the probability threshold, which in one embodiment can be set to 0.5. If the difference between the maximum corrected probability among all predicted modes and the corrected probability of the predicted mode at index i is less than or equal to the probability threshold, then the predicted mode at index i is added to the Rate Distortion Optimization (RDO) list; otherwise, the predicted mode at index i is removed from the RDO list. If the predicted mode at index i is not in the RDO list, then the RDO search for the predicted mode at index i is skipped during RDO execution, and unnecessary predicted mode iterations are also skipped, allowing the execution of the next coding unit iteration. This saves the predicted mode selection time for the VVC-SCC coding unit and accelerates the VVC-SCC coding process.

[0070] The effects of the present invention will be illustrated below using specific experimental results.

[0071] The test platform used in this application is VTM23.1, and the test configuration is All-Intra with IBC and PLT modes enabled. The test sequence is a CTC sequence, as shown in Table 1. Compared with the test platform, the VVC-SCC coding method based on multi-size and neighborhood combination prediction mode proposed in the embodiments of this application can save an average of 32.37% of the coding time when encoding different sequences, while the BDBR (Bjøntegaard Delta Bit Rate) only increases by 1.19%. This shows that the VVC-SCC coding method based on multi-size and neighborhood combination prediction mode proposed in the embodiments of this application can significantly accelerate the coding speed while ensuring rate-distortion performance.

[0072] Table 1

[0073]

[0074] Further reference Figure 4 As an implementation of the methods shown in the above figures, this application provides an embodiment of a VVC-SCC coding device based on a prediction pattern combining multiple sizes and neighborhoods. This device embodiment is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0075] This application provides a VVC-SCC coding device based on a prediction pattern combining multiple sizes and neighborhood, comprising:

[0076] Model building module 1 is configured to build a multi-size intra-coding unit prediction pattern classification model and train it for different quantization parameters to obtain a trained multi-size intra-coding unit prediction pattern classification model corresponding to each quantization parameter; the multi-size intra-coding unit prediction pattern classification model includes a first sub-network, a second sub-network and a classification module, and the first sub-network and the second sub-network are respectively connected to the classification module;

[0077] The multi-scale prediction module 2 is configured to acquire the screen content video sequence and specified quantization parameters, and encode each video frame in the screen content video sequence using a VVC-SCC encoder. During the encoding process, the size, position, and pixel value of the current coding unit obtained from the video frame division are obtained. The size and pixel value of the current coding unit are input into the trained multi-scale intra-frame coding unit prediction mode classification model corresponding to the specified quantization parameters. In response to determining that the minimum value of the height and width of the current coding unit is greater than a size threshold, the pixel value of the current coding unit is input into the first sub-network to extract the first pixel feature. In response to determining that the minimum value of the height and width of the current coding unit is less than or equal to the size threshold, the pixel value of the current coding unit is input into the second sub-network to extract the second pixel feature. The first pixel feature or the second pixel feature is processed by the classification module to obtain the prediction probability of the current coding unit in each prediction mode.

[0078] The correction module 3 is configured to calculate the neighborhood prediction probability corresponding to each prediction mode based on the prediction modes used by the adjacent reconstructed coding units located to the left, above, and upper left of the current coding unit; and is configured to add the prediction probability of the current coding unit in each prediction mode to its corresponding neighborhood prediction probability to obtain the correction probability corresponding to each prediction mode.

[0079] Decision module 4 makes prediction mode decisions based on the correction probability corresponding to each prediction mode, and determines whether to place the prediction mode in the rate distortion optimization list.

[0080] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. For example... Figure 5 As shown, the electronic device of this embodiment includes a processor 501 and a memory 502; wherein the memory 502 is used to store computer execution instructions; and the processor 501 is used to execute the computer execution instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.

[0081] Alternatively, the memory 502 can be either standalone or integrated with the processor 501.

[0082] When the memory 502 is set up independently, the electronic device also includes a bus 503 for connecting the memory 502 and the processor 501.

[0083] This invention also provides a computer storage medium storing computer execution instructions, which, when executed by processor 501, implement the above method.

[0084] This invention also provides a computer program product, including a computer program that, when executed by a processor 501, implements the above-described method.

[0085] In the embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0086] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.

[0087] Furthermore, the functional modules in the various embodiments of this invention can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit formed by the above modules can be implemented in hardware or in the form of hardware plus software functional units.

[0088] The integrated modules implemented as software functional modules described above can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor 501 to execute some steps of the methods of the various embodiments of this application.

[0089] It should be understood that the processor 501 described above can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor, or the processor 501 can be any conventional processor 501. The steps of the method disclosed in this invention can be directly manifested as the hardware processor 501 executing the steps, or as a combination of hardware and software modules within the processor 501 executing the steps.

[0090] The memory 502 may include high-speed RAM memory, and may also include non-volatile memory NVM, such as at least one disk storage device, and may also be a USB flash drive, portable hard drive, read-only memory, disk or optical disc, etc.

[0091] Bus 503 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 503 can be divided into address bus, data bus, control bus, etc. For ease of illustration, the bus 503 in the accompanying drawings of this application is not limited to only one bus 503 or one type of bus 503.

[0092] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.

[0093] An exemplary storage medium is coupled to processor 501, enabling processor 501 to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of processor 501. Processor 501 and storage medium can reside in application-specific integrated circuits (ASICs). Alternatively, processor 501 and storage medium can exist as discrete components in an electronic device or host device.

[0094] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A VVC-SCC coding method based on prediction patterns combining multiple sizes and neighborhoods, characterized in that, Includes the following steps: A multi-size intra-coding unit prediction pattern classification model is constructed and trained for different quantization parameters to obtain a trained multi-size intra-coding unit prediction pattern classification model corresponding to each quantization parameter; the multi-size intra-coding unit prediction pattern classification model includes a first sub-network, a second sub-network and a classification module, wherein the first sub-network and the second sub-network are respectively connected to the classification module; The screen content video sequence and specified quantization parameters are obtained. Each video frame in the screen content video sequence is encoded using a VVC-SCC encoder. During the encoding process, the size, position, and pixel value of the current coding unit obtained by dividing the video frame are obtained. The size and pixel value of the current coding unit are input into the trained multi-size intra-coding unit prediction mode classification model corresponding to the specified quantization parameters. In response to determining that the minimum value of the height and width of the current coding unit is greater than the size threshold, the pixel value of the current coding unit is input into the first sub-network to extract the first pixel feature. In response to determining that the minimum of the height and width of the current coding unit is less than or equal to a size threshold, the pixel value of the current coding unit is input into the second sub-network to extract the second pixel feature; the first pixel feature or the second pixel feature is processed by the classification module to obtain the prediction probability of the current coding unit in each prediction mode; The neighborhood prediction probability corresponding to each prediction mode is calculated based on the prediction modes used by the adjacent reconstructed coding units located to the left, above, and upper left of the current coding unit; the prediction probability of the current coding unit in each prediction mode is added to its corresponding neighborhood prediction probability to obtain the correction probability corresponding to each prediction mode. Based on the correction probability corresponding to each prediction mode, a prediction mode decision is made to determine whether the prediction mode should be placed in the rate-distortion optimization list, specifically including: The maximum value among all the corrected probabilities corresponding to the prediction modes is taken as the maximum corrected probability. Iterate through the correction probabilities corresponding to each prediction mode, and determine whether the difference between the maximum correction probability and the correction probability corresponding to one of the prediction modes is less than or equal to the probability threshold. If so, add one of the prediction modes to the rate distortion optimization list; otherwise, remove one of the prediction modes from the rate distortion optimization list.

2. The VVC-SCC coding method based on prediction patterns combining multiple sizes and neighborhoods as described in claim 1, characterized in that, The first sub-network includes a first sampling layer, a first convolutional layer, and a second convolutional layer connected in sequence; the second sub-network includes a second sampling layer and a third convolutional layer connected in sequence; and the classification module includes a first fully connected layer and a second fully connected layer connected in sequence.

3. The VVC-SCC coding method based on prediction patterns combining multiple sizes and neighborhoods according to claim 2, characterized in that, The first sampling layer has a convolutional kernel size of 4sw×4sh and a stride of 4sw×4sh; the second sampling layer has a convolutional kernel size of 2sw×2sh and a stride of 2sw×2sh; where w represents the width of the coding unit, h represents the height of the coding unit, sw represents the width unit, and sh represents the height unit. When w > h, sw = w / h, sh = 1; when w ≤ h, sw = 1, sh = h / w; when w = h, both sw and sh are 1; the convolutional kernels and strides of the first, second, and third convolutional layers are all 2×2; the first sampling layer, the first convolutional layer, the second sampling layer, the second convolutional layer, the third convolutional layer, and the first fully connected layer all use the ReLU activation function, and the second fully connected layer uses the Softmax function.

4. The VVC-SCC coding method based on prediction patterns combining multiple sizes and neighborhoods according to claim 1, characterized in that, The neighborhood prediction probability for each prediction mode is calculated based on the prediction modes used by the adjacent reconstructed coding units located to the left, above, and above the current coding unit. Specifically, this includes: The neighborhood prediction probability corresponding to each prediction mode is calculated using the following formula: ; in, , and These represent the usage parameters of the adjacent reconstructed coding units at the left, top, and top-left positions of the current coding unit, respectively. They determine whether the adjacent reconstructed coding units use the prediction mode with index i for prediction. If so, the value of the corresponding usage parameter is 1; otherwise, the value of the corresponding usage parameter is 0. This represents the neighborhood prediction probability corresponding to the prediction mode with index i. The value of i is 0, 1 or 2, where 0 represents Intra mode, 1 represents PLT mode and 2 represents IBC mode.

5. The VVC-SCC coding method based on prediction patterns combining multiple sizes and neighborhoods according to claim 1, characterized in that, During the training process of the multi-size intra-frame coding unit prediction mode classification model, the screen content video dataset is encoded on the VVC-SCC standard test platform in full intra-frame configuration mode, and the true labels of prediction modes with different quantization parameters are obtained. For each quantization parameter, the corresponding true label of the prediction mode is selected to train the multi-size intra-coding unit prediction mode classification model. The loss function used in the training process is... for: ; Where N represents the batch training size. and These represent the actual label and the predicted label of the prediction mode with index i for the nth coding unit under the same quantization parameters. The predicted label is the index value corresponding to the maximum value of the prediction probabilities of all prediction modes output by the multi-size intra-coding unit prediction mode classification model. The value of i is 0, 1 or 2, where 0 represents Intra mode, 1 represents PLT mode and 2 represents IBC mode.

6. A VVC-SCC coding device based on a prediction pattern combining multiple sizes and neighborhood, characterized in that, include: The model building module is configured to build a multi-size intra-coding unit prediction pattern classification model and train it for different quantization parameters to obtain a trained multi-size intra-coding unit prediction pattern classification model corresponding to each quantization parameter; the multi-size intra-coding unit prediction pattern classification model includes a first sub-network, a second sub-network and a classification module, wherein the first sub-network and the second sub-network are respectively connected to the classification module. The multi-scale prediction module is configured to acquire a screen content video sequence and specified quantization parameters, and encode each video frame in the screen content video sequence using a VVC-SCC encoder. During the encoding process, the size, position, and pixel value of the current coding unit obtained by dividing the video frame are obtained. The size and pixel value of the current coding unit are input into the trained multi-scale intra-frame coding unit prediction mode classification model corresponding to the specified quantization parameters. In response to determining that the minimum value of the height and width of the current coding unit is greater than the size threshold, the pixel value of the current coding unit is input into the first sub-network to extract the first pixel feature. In response to determining that the minimum of the height and width of the current coding unit is less than or equal to a size threshold, the pixel value of the current coding unit is input into the second sub-network to extract the second pixel feature; the first pixel feature or the second pixel feature is processed by the classification module to obtain the prediction probability of the current coding unit in each prediction mode; The correction module is configured to calculate the neighborhood prediction probability corresponding to each prediction mode based on the prediction modes used by the adjacent reconstructed coding units located to the left, above, and upper left of the current coding unit; and to add the prediction probability of the current coding unit in each prediction mode to its corresponding neighborhood prediction probability to obtain the correction probability corresponding to each prediction mode. The decision module is configured to make prediction mode decisions based on the correction probability corresponding to each prediction mode, and to determine whether to place the prediction mode in the rate-distortion optimization list, specifically including: The maximum value among all the corrected probabilities corresponding to the prediction modes is taken as the maximum corrected probability. Iterate through the correction probabilities corresponding to each prediction mode, and determine whether the difference between the maximum correction probability and the correction probability corresponding to one of the prediction modes is less than or equal to the probability threshold. If so, add one of the prediction modes to the rate distortion optimization list; otherwise, remove one of the prediction modes from the rate distortion optimization list.

7. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-model fusion VVC intra-frame coding rapid CU division method and storage medium

    CN118784835A

  • VVC-SCC intra-frame coding method and device based on multi-stage irregular coding unit division

    CN119299671A